Search 227 Remote CKAD Jobs

227 remote jobs

Platform Engineer

Hybrid

About PayPay Card PayPay Card Corporation was established in 2021 to provide users a FinTech service that is more accessible and convenient compared to previous credit cards and credit services by integrating with the PayPay payment platform which has surpassed 75 million users since its launch (as of August 2026). We are looking for people who are passionate about refining our products at an overwhelming speed that other companies cannot match as well as professionals who are interested in promoting the spread of cashless payments in Japan and the use of these payments as a financial life platform. Let us work together to create new value for users. ※ Please note that you cannot apply or be selected in parallel with PayPay Corporation PayPay Card Corporation and PayPay Securities Corporation. Job Description PayPay Card is looking for a Cloud infrastructure and platform engineer for PayPay Card's cloud system. This role will be responsible for the engineering of Cloud Infrastructure Platforms to provide robust reliable and cost-efficient Platform so that we can deliver the best experience for our growing customers base. Main Responsibilities Architect and build a robust scalable and stable infrastructure Architect and build Infrastructure that is easily maintainable during patching and upgrades Architect and build Infrastructure with AWS regional level failure tolerance Work together with our Security Engineers to provision a secure Infrastructure Automations for AWS Deployments to ensure fast delivery of Cloud Services to our developers Provide through Automations means for developers to easily deploy resources on AWS. Tech Stack AWS VPC EC2 ECS EKS Lambda Cloudfront MWAA RDS ElastiCache DynamoDB Opensearch S3 CloudWatch Cognito SQS KMS Secret Manager KMS MSK CodeCommit CodeBuild CodeDeploy CodePipeline CloudFormation and other services Terraform Github Actions Prometheus Grafana Dynatrace Atlantis ArgoCD OpenTelemetry Required Qualifications More than 5 years of technical experience in Cloud based Infrastructure Platform Ability to demonstrate high degree of ownership in a Production environment Good understanding of Cloud security best practices and payment industry compliance standards Expertise in designing building and troubleshooting large-scale streaming and distributed systems using Kafka. Expertise in building Robust platforms based on NoSQL/KVS datastores like DynamoDB Opensearch/ Elasticsearch and Redis. Experience in Cloud infrastructure and platform systems availability performance and cost management Extensive technical hands-on experience of Computing Storage and Analytical Services on Cloud(AWS) Experience with IAC tools in AWS such as Terraform CloudFormation CDK Experience with Cloud Services Monitoring Detection and Response Experience in Cloud Services Performance Tuning cost controls and management Experience in Cloud Infrastructure Service patching and upgrades PayPay DevOps emphasize automation. Demonstrated skill with the following are required Python ShellScripting Go/ Rust (optional) Have excellent oral written verbal and interpersonal communication skills. Preferred Qualifications Bachelor’s degree and above in a technology related field Experience with other cloud service providers (e.g GCP Azure) Experience with Kubernetes (CKA CKAD or CKS) Experience with Event-Driven Architecture (Kafka preferred) Experience using and contributing to Open Source tools Experience in managing IT compliance and security risk Published papers / blogs / articles - Relevant and verifiable certifications Bilingual in English and Japanese is nice to have but not required. Proficiency in either language is fine. Working Conditions Employment Status Full Time Office Location Hybrid Workstyle (flexible working style including Remote and office) ※ You will be expected to work both in the office and remotely in alignment with organizational guidelines and team objectives. LIFE in JAPAN FACTBOOK Work Hours Full Flex Time (No Core Time) In principle 900am ~ 545pm (actual working hours 7h45m + 1h break) Holidays Every Sat/Sun/National holidays (In Japan)/New Year's break/Company-designated Special days Paid leave Annual leave (up to 14 days in the first year granted proportionally according to the month of employment. Can be used from the date of hire) Personal leave (5 days each year granted proportionally according to the month of employment) *PayPay Group's own special paid leave system which can be used to attend to illnesses injuries hospital visits etc. of the employee family members pets etc. Salary Annual salary paid in 12 installments (monthly) Reviewed once a year Overtime allowance Late overtime allowance Commuting and transportation expenses Benefits Social Insurance (health insurance employee pension employment insurance and compensation insurance) 401K Other Information PayPay Inside-Out (Corporate Blog) ENG Recruiting FACTBOOK for PayPay Card

Senior Software Engineer – Platform & Data Infrastructure

McLean, Virginia Mountain View, California, United States

Company Overview ID.me is the next-generation digital identity wallet that simplifies how individuals securely prove their identity online. Consumers can verify their identity with ID.me once and seamlessly login across websites without having to create a new login and verify their identity again. Over 152 million users experience streamlined login and identity verification with ID.me at 20 federal agencies 45 state government agencies and 70+ healthcare organizations. More than 600+ consumer brands use ID.me to verify communities and user segments to honor service and build more authentic relationships. ID.me’s technology meets the federal standards for consumer authentication set by the Commerce Department and is approved as a NIST 800-63-3 IAL2 / AAL2 credential service provider by the Kantara Initiative. ID.me is committed to “No Identity Left Behind” to enable all people to have a secure digital identity. To learn more visit https//network.id.me/ ID.me is a full-time in-office culture. Unless a specific job description explicitly states otherwise all roles are on-site five days per week at one of our offices in McLean VA Mountain View CA New York City NY or Tampa FL. Certain roles — such as field-based sales or other remote-by-design positions — may have different work arrangements as noted in their individual postings. At ID.me we embrace the thoughtful use of AI tools in our daily work and there are even occasions where we leverage AI in our hiring process. However during the interview process we want to understand your individual skills and experiences. Therefore we have guidelines on how AI can be appropriately used during your application and interviews which can be found here Role Overview ID.me is seeking a Senior Software Engineer (Platform & Data Infrastructure) to architect scale and maintain the core cloud-native platform and data infrastructure that powers our digital identity platform. In this role you will take ownership of multi-region Kubernetes fleets automated GitOps pipelines policy-as-code governance and enterprise data workflow orchestrations. Working closely with Platform Security SRE and Product Engineering teams you will turn complex infrastructure into self-service developer capabilities enforce security guardrails and ensure high availability across critical production services. This is a hands-on technical role for a senior engineer with deep expertise in cloud-native distributed systems Infrastructure as Code (IaC) GitOps continuous delivery and data platform operations. This role is based out of our Mountain View CA or McLean VA offices and requires full-time in-office attendance. Responsibilities Operate Cloud-Native Kubernetes Fleets Own the lifecycle network topology node-pool migrations and multi-region deployment of production Kubernetes clusters (GKE/EKS) supporting microservices at scale. Drive GitOps & CI/CD Delivery Maintain automated GitOps reconciliation pipelines using FluxCD and Helm driving cluster bootstrapping progressive rollouts and secret integrations. Scale Infrastructure as Code & Policy Governance Architect modular Terraform codebases manage enterprise workspaces (Terraform Enterprise/Terragrunt) and enforce compliance guardrails using policy-as-code frameworks (Sentinel/OPA). Maintain Data Platform Infrastructure Build and optimize data pipelines and workflow orchestrations using Apache Airflow Kafka and enterprise data warehouses (e.g. BigQuery). Enforce Zero-Trust Security & Secrets Management Integrate enterprise secret management tools (HashiCorp Vault External Secrets Operator Google Secret Manager/KMS) and enforce Workload Identity and network perimeter controls. Lead Complex Infrastructure Migrations Direct zero-downtime infrastructure initiatives including IP-addressing migrations cluster rebuilds and database cutovers. Drive Reliability DR & On-Call Excellence Active participation in on-call rotations designing observability dashboards (Splunk GCP Cloud Logging/Stackdriver) writing runbooks and executing regional disaster recovery (DR) failover drills. Technical Mentorship Guide mid-level and junior engineers through code/design reviews establishing self-service infrastructure standards across the organization. Minimum Qualifications Bachelor’s or Graduate degree in Computer Science Software Engineering or a related technical field. 4+ years of professional experience in platform engineering Site Reliability Engineering (SRE) or cloud data infrastructure. Strong programming skills in modern languages/scripts such as Python Bash SQL or Java. Deep experience operating enterprise-scale Kubernetes fleets (GKE or EKS) and container ecosystems. Extensive practical experience writing modular Terraform and managing Infrastructure as Code at scale. Demonstrated experience operating event-streaming and workflow systems such as Apache Airflow Kafka or BigQuery. Preferred Qualifications Certifications Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD). Hands-on experience with Terraform Enterprise (TFE) and writing Sentinel or Open Policy Agent (OPA) policy-as-code guardrails. Proven experience with GitOps workflows utilizing FluxCD or ArgoCD across multi-tenant production fleets. Advanced secrets management expertise using HashiCorp Vault and Kubernetes External Secrets Operator. Track record of executing large-scale cloud migrations data warehouse transitions (e.g. Vertica to BigQuery) or networking range overhauls with zero customer impact. The annual base salary listed does not include a company bonus incentive for sales roles equity and benefits which will be determined based on experience skills education relevant training geographic location and role. ID.me offers comprehensive medical dental vision health savings account flexible spending accounts (medical limited purpose dependent care commuter benefit accounts) basic and voluntary life and AD&D insurance 401(k) with company match parental leave ability to participate in unlimited paid time off subject to the terms and conditions of the PTO policy including 8 company wide holidays short and long-term disability insurance accident and critical illness insurance referral bonus policy employee assistance program pet insurance travel assistant program wellbeing and childcare discounts benefit advocates and a learning and development benefit. Final offers may vary from the amount listed based on qualifications professional experiences skills education relevant training geographic location and other job related factors. Mountain View CA Pay Range $190978 $214052 USD ID.me maintains a work environment free from discrimination where employees are treated with dignity and respect. All ID.me employees share in the responsibility for fulfilling our commitment to equal employment opportunity. ID.me does not discriminate against any employee or applicant on the basis of age ancestry color family or medical care leave gender identity or expression genetic information marital status medical condition national origin physical or mental disability political affiliation protected veteran status race religion sex (including pregnancy) sexual orientation or any other characteristic protected by applicable laws regulations and ordinances. ID.me adheres to these principles in all aspects of employment including recruitment hiring training compensation promotion benefits social and recreational programs and discipline. In addition ID.me's policy is to provide reasonable accommodation to qualified employees who have protected disabilities to the extent required by applicable laws regulations and ordinances where a particular employee works. Upon request we will provide you with more information about such accommodations. Please review our Privacy Policy including our CCPA policy at id.me/privacy . If you provide ID.me with any personally identifiable information you confirm that you have read and agree to be bound by the terms and conditions set out in our Privacy Policy. ID.me participates in E-Verify.

AI Platform Engineer

AI Platform Engineer L3 (GPUaaS – AI Neocloud) 📍 EMEA Remote-first About Sharon AI Sharon AI is building the infrastructure powering the next generation of artificial intelligence. Operating across AI infrastructure high-performance compute cloud platforms and large-scale technology environments Sharon AI delivers scalable secure and reliable infrastructure for demanding AI ML and HPC workloads. The Role As an AI Platform Engineer L3 you'll design build and operate the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering across the EMEA region. You'll own the architecture automation and reliability of the platform services sitting above Sharon AI's GPU and network fabric spanning Kubernetes Slurm GPU scheduling MLOps tooling model serving and platform observability across multiple EMEA sites. Reporting to the Head of Operations you'll work closely with Network Engineering Infrastructure and customer-facing teams to solve complex cross-team challenges and ensure Sharon AI's platform can scale reliably and efficiently across the region. This is a hands-on senior individual contributor role suited to an engineer with strong platform engineering DevOps or MLOps experience who can operate independently in a fast-paced AI-native environment and provide technical guidance to less experienced engineers. Key Responsibilities Design and own the AI platform architecture across Sharon AI's EMEA GPU clusters including Kubernetes Slurm and container orchestration Lead the development of CI/CD pipelines and MLOps tooling supporting training fine-tuning and inference workloads across multiple sites Define and implement multi-tenant GPU resource scheduling quota management and workload isolation strategies at scale Own the design of model serving infrastructure balancing high availability performance and cost efficiency Build and evolve platform-wide observability across monitoring logging and alerting covering platform health GPU utilisation and workload performance Drive Infrastructure-as-Code adoption and platform automation using Terraform and Ansible Partner closely with Network Engineering to integrate the platform layer with Sharon AI's underlying InfiniBand/RDMA fabric Act as a senior escalation point for complex platform issues impacting customer AI/ML workloads across EMEA Lead incident response and post-incident reviews contributing to operational runbooks and platform best practice Partner with the Head of Operations on platform capacity planning scaling strategy and cost optimisation across EMEA Mentor and provide technical guidance to less experienced platform engineers Support enterprise GPUaaS customer onboarding and technical escalations across the region Skills & Experience 6–10+ years' experience in platform engineering DevOps MLOps or SRE ideally within HPC cloud or AI/ML infrastructure environments Bachelor's degree in Computer Science Electrical Engineering or a related field Hands-on experience operating Kubernetes and GPU scheduling at production scale Proven experience designing CI/CD and Infrastructure-as-Code practices for platform teams Proven experience supporting GPU or AI/ML infrastructure at scale ideally within a GPUaaS or neocloud environment Deep expertise in Kubernetes and GPU scheduling frameworks including Slurm Kubernetes device plugins and NVIDIA GPU Operator Strong experience designing and operating MLOps tooling and ML pipeline orchestration in production Advanced proficiency in Python Bash and Infrastructure-as-Code tools such as Terraform and Ansible Strong understanding of GPU infrastructure and distributed training concepts including NCCL data/model parallelism and RDMA-aware scheduling Experience with platform-wide observability tooling such as Prometheus Grafana and telemetry stacks Strong Linux systems and networking fundamentals with the judgement to independently solve ambiguous cross-team problems A security-first mindset when operating within multi-tenant environments Awareness of EMEA regulatory and data residency considerations including GDPR as they relate to platform operations Strong communication and collaboration skills across distributed multi-region teams with the ability to mentor other engineers Experience with distributed training frameworks such as PyTorch and TensorFlow MLOps platforms including MLflow Kubeflow or Ray and InfiniBand/RDMA or RoCEv2 networking concepts is advantageous. Kubernetes certifications such as CKA CKAD or CKS and cloud certifications across AWS GCP or Azure are also advantageous. The role requires the right to work in an EMEA jurisdiction with existing eligibility to work across the EU/EEA or UK advantageous given the multi-country remit. Why Join Sharon AI Own the platform architecture powering a growing GPU-as-a-Service and AI neocloud business across EMEA Work hands-on with large-scale GPU infrastructure Kubernetes Slurm and AI-native platform technologies Shape how Sharon AI's platform scales across multiple jurisdictions sites and customer workloads Solve complex technical challenges across GPU infrastructure MLOps networking and distributed AI workloads Influence platform reliability automation capacity and cost optimisation across the region Work closely with Network Engineering and Infrastructure teams on high-performance InfiniBand/RDMA environments Provide technical leadership and mentorship while remaining hands-on as a senior individual contributor Help enterprise customers reliably train fine-tune and run AI/ML workloads at scale Join a highly technical and ambitious team operating at the forefront of AI infrastructure Our Values Integrity Innovation Collaboration Wellbeing Inclusion

AWS Cloud Infrastructure Engineer – RPA Operations (Remote)

Arlington, VA

Tuknik Government Services a Koniag Government Services company is seeking an AWS Cloud Infrastructure Engineer – RPA Operations with a Secret security clearance to support TGS and our government customer. The position is remote. We offer competitive compensation and an extraordinary benefits package including health dental and vision insurance 401K with company matching flexible spending accounts paid holidays three weeks paid time off and more. The AWS Cloud Infrastructure Engineer - RPA Operations will be responsible for designing implementing and maintaining the organization's Amazon Web Services (AWS) cloud infrastructure that supports Robotic Process Automation (RPA) operations specifically in UiPath RPA software. This role serves as the technical expert for all AWS-related services utilized by RPA customers ensuring secure scalable and highly available virtual infrastructure for automation workloads. The position requires deep expertise in AWS services virtualization technologies security management and container orchestration. The ideal candidate will manage complex cloud environments including virtual machines AWS Workspaces Hardware Security Modules (HSM) and certificate management systems while maintaining strict security protocols for Non-Person Entity (NPE) credentials and unattended automation infrastructure. This role combines cloud architecture systems administration security management and DevOps practices with specialization in System Center Configuration Manager (SCCM) and Kubernetes orchestration. The engineer will work closely with RPA development teams security personnel and end-users to ensure optimal performance security and accessibility of cloud resources that enable enterprise-wide automation capabilities. Essential Functions Responsibilities & Duties may include but are not limited to AWS Infrastructure Management & Administration Design deploy and maintain AWS cloud infrastructure supporting RPA/UiPath customer operations Create and manage virtual machines (EC2 instances) optimized for unattended automation workloads Configure and maintain AWS Workspaces for RPA developers and automation operators Implement and manage AWS services including but not limited to EC2 VPC S3 IAM CloudWatch CloudFormation Systems Manager and CloudTrail Monitor cloud resource utilization performance metrics and costs optimize infrastructure for efficiency Implement and maintain backup disaster recovery and business continuity solutions for RPA infrastructure Manage cloud networking configurations including VPCs subnets security groups routing tables and network ACLs Ensure high availability and fault tolerance of critical RPA infrastructure components Plan and execute infrastructure scaling activities to meet growing automation demands Perform system maintenance patching and updates across cloud environments Security & Certificate Management Manage AWS Hardware Security Module (HSM) or similar infrastructure for secure operations Implement and maintain digital certificate lifecycle management for RPA operations Configure and manage Non-Person Entity (NPE) credentials for unattended automation processes Establish secure credential storage rotation and access procedures using AWS HSM or similar and Secrets Manager Implement and enforce security best practices for cloud infrastructure and automation workloads Manage Identity and Access Management (IAM) policies roles and permissions for RPA services Configure and monitor security controls including encryption at rest and in transit Conduct security assessments and implement remediation measures for identified vulnerabilities Maintain compliance with organizational cybersecurity policies and federal security standards Manage PKI (Public Key Infrastructure) components and certificate authorities for automation environments Coordinate with security teams on audits compliance requirements and security incidents Systems Management & Configuration Administer System Center Configuration Manager (SCCM) for endpoint management and software distribution Deploy and manage applications updates and configurations across virtual infrastructure using SCCM Maintain SCCM infrastructure including site servers distribution points and management points Create and maintain SCCM task sequences packages and deployment collections Monitor system compliance and remediate configuration drift Integrate SCCM with cloud services for hybrid environment management Troubleshoot SCCM-related issues and optimize performance Container Orchestration & Kubernetes Management Design deploy and manage Kubernetes clusters Configure and maintain Amazon Elastic Kubernetes Service (EKS) or equivalent container platforms Develop and manage container images pods services and deployments Implement container security best practices and vulnerability scanning Monitor cluster health resource utilization and application performance Troubleshoot container and orchestration issues Implement CI/CD pipelines for container deployments User Support & Documentation Provide technical support to RPA customers for AWS infrastructure access and usage issues Troubleshoot and resolve virtual infrastructure problems including connectivity performance and configuration issues Create and maintain comprehensive documentation for AWS infrastructure procedures and configurations Document step-by-step procedures for NPE credential acquisition and setup in AWS environments Develop user guides for accessing and utilizing AWS Workspaces and virtual machines Maintain runbooks for common administrative tasks and troubleshooting scenarios Conduct knowledge transfer sessions with team members and end-users Respond to user inquiries regarding cloud infrastructure capabilities and limitations Education Required Qualifications Bachelor's degree in Computer Science Information Technology Computer Engineering or related technical field OR Equivalent combination of education and professional experience in cloud infrastructure systems engineering or DevOps Experience Minimum 4-5 years of hands-on experience with AWS cloud services and infrastructure management Minimum 3 years of experience with virtual machine administration and cloud computing environments Demonstrated experience with System Center Configuration Manager (SCCM) administration Minimum 2 years of experience with Kubernetes or container orchestration platforms Experience with security management certificate authorities and cryptographic systems Previous experience supporting enterprise automation or RPA infrastructure (preferred) Experience with Infrastructure as Code (IaC) and DevOps practices Required Technical Skills AWS Services Expert-level knowledge of AWS core services (EC2 VPC S3 IAM CloudWatch) Advanced experience with AWS Workspaces and virtual desktop infrastructure (VDI) Proficiency with AWS Hardware Security Module (CloudHSM or AWS Managed HSM) Strong understanding of AWS security services (Secrets Manager KMS Certificate Manager IAM) Experience with AWS networking security groups and VPC configurations Familiarity with AWS automation tools (CloudFormation Systems Manager CLI SDKs) Security & Certificate Management Deep knowledge of Public Key Infrastructure (PKI) and certificate lifecycle management Experience with digital certificate creation distribution and revocation Understanding of cryptographic concepts and secure key management Experience managing Non-Person Entity (NPE) credentials and service accounts Knowledge of security compliance frameworks and best practices Systems Management Advanced proficiency with Microsoft System Center Configuration Manager (SCCM) Experience with endpoint management patch management and software distribution Knowledge of Windows Server administration and Active Directory integration Familiarity with PowerShell scripting for automation and system management Container & Orchestration Strong experience with Kubernetes architecture deployment and management Proficiency with Amazon EKS or other managed Kubernetes services Knowledge of Docker and container technologies Experience with container security scanning and vulnerability management Understanding of microservices architecture and cloud-native applications Infrastructure As Code & Automation Proficiency with Infrastructure as Code tools (Terraform CloudFormation ARM templates) Experience with configuration management tools (Ansible Puppet Chef) Strong scripting skills (PowerShell Python Bash) Familiarity with CI/CD pipelines and DevOps tools (Jenkins GitLab Azure DevOps) Monitoring & Troubleshooting Experience with cloud monitoring and logging solutions (CloudWatch CloudTrail ELK stack) Strong analytical and troubleshooting skills for complex cloud environments Performance tuning and optimization expertise Preferred AWS certifications (Solutions Architect SysOps Administrator Security Specialty DevOps Engineer) Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) Microsoft certifications (Azure Administrator SCCM/MECM certification) UiPath infrastructure knowledge or RPA platform administration experience Experience with HashiCorp tools (Vault Consul Terraform) Knowledge of networking protocols load balancers and CDN services Experience with disaster recovery and business continuity planning Familiarity with compliance frameworks (FedRAMP NIST FISMA) Professional Competencies Strong analytical and problem-solving abilities for complex technical challenges Excellent documentation skills with attention to detail and accuracy Effective communication skills to explain technical concepts to diverse audiences Customer-service oriented with a focus on user satisfaction Ability to work independently and manage multiple concurrent projects Strong organizational and time management skills Commitment to continuous learning and staying current with cloud technologies Collaborative team player with ability to work across functional groups Security-conscious mindset with understanding of risk management Proactive approach to identifying and resolving potential issues Security & Clearance Requirements Must be able to obtain and maintain Secret clearance Ability to handle sensitive information and credentials with appropriate discretion Understanding of and adherence to federal cybersecurity requirements Other Requirements Availability for on-call support rotation for critical infrastructure issues Willingness to work outside normal business hours for maintenance windows when necessary Ability to support geographically dispersed user base Commitment to maintaining professional certifications and technical skills Strong ethical standards regarding system access and data security U.S. Citizen. Our Equal Employment Opportunity Policy The company is an equal opportunity employer. The company shall not discriminate against any employee or applicant because of race color religion creed ethnicity sex sexual orientation gender or gender identity (except where gender is a bona fide occupational qualification) national origin or ancestry age disability citizenship military/veteran status marital status genetic information or any other characteristic protected by applicable federal state or local law. We are committed to equal employment opportunity in all decisions related to employment promotion wages benefits and all other privileges terms and conditions of employment. The company is dedicated to seeking all qualified applicants. If you require an accommodation to navigate or apply for a position on our website please get in touch with Heaven Wood via e-mail at accommodations@koniag-gs.com or by calling 703-488-9377 to request accommodations. Koniag Government Services (KGS) is an Alaska Native Owned corporation supporting the values and traditions of our native communities through an agile employee and corporate culture that delivers Enterprise Solutions Professional Services and Operational Management to Federal Government Agencies. As a wholly owned subsidiary of Koniag we apply our proven commercial solutions to a deep knowledge of Defense and Civilian missions to provide forward-leaning technical professional and operational solutions. KGS enables successful mission outcomes for our customers through solution-oriented business partnerships and a commitment to exceptional service delivery. We ensure long-term success with a continuous improvement approach while balancing the collective interests of our customers employees and native communities. For more information please visit www.koniag-gs.com. Equal Opportunity Employer/Veterans/Disabled. Shareholder Preference in accordance with Public Law 88-352

AI Platform Engineer

AI Platform Engineer L3 (GPUaaS – AI Neocloud) 📍 EMEA Remote-first About Sharon AI Sharon AI is building the infrastructure powering the next generation of artificial intelligence. Operating across AI infrastructure high-performance compute cloud platforms and large-scale technology environments Sharon AI delivers scalable secure and reliable infrastructure for demanding AI ML and HPC workloads. The Role As an AI Platform Engineer L3 you'll design build and operate the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering across the EMEA region. You'll own the architecture automation and reliability of the platform services sitting above Sharon AI's GPU and network fabric spanning Kubernetes Slurm GPU scheduling MLOps tooling model serving and platform observability across multiple EMEA sites. Reporting to the Head of Operations you'll work closely with Network Engineering Infrastructure and customer-facing teams to solve complex cross-team challenges and ensure Sharon AI's platform can scale reliably and efficiently across the region. This is a hands-on senior individual contributor role suited to an engineer with strong platform engineering DevOps or MLOps experience who can operate independently in a fast-paced AI-native environment and provide technical guidance to less experienced engineers. Key Responsibilities Design and own the AI platform architecture across Sharon AI's EMEA GPU clusters including Kubernetes Slurm and container orchestration Lead the development of CI/CD pipelines and MLOps tooling supporting training fine-tuning and inference workloads across multiple sites Define and implement multi-tenant GPU resource scheduling quota management and workload isolation strategies at scale Own the design of model serving infrastructure balancing high availability performance and cost efficiency Build and evolve platform-wide observability across monitoring logging and alerting covering platform health GPU utilisation and workload performance Drive Infrastructure-as-Code adoption and platform automation using Terraform and Ansible Partner closely with Network Engineering to integrate the platform layer with Sharon AI's underlying InfiniBand/RDMA fabric Act as a senior escalation point for complex platform issues impacting customer AI/ML workloads across EMEA Lead incident response and post-incident reviews contributing to operational runbooks and platform best practice Partner with the Head of Operations on platform capacity planning scaling strategy and cost optimisation across EMEA Mentor and provide technical guidance to less experienced platform engineers Support enterprise GPUaaS customer onboarding and technical escalations across the region Skills & Experience 6–10+ years' experience in platform engineering DevOps MLOps or SRE ideally within HPC cloud or AI/ML infrastructure environments Bachelor's degree in Computer Science Electrical Engineering or a related field Hands-on experience operating Kubernetes and GPU scheduling at production scale Proven experience designing CI/CD and Infrastructure-as-Code practices for platform teams Proven experience supporting GPU or AI/ML infrastructure at scale ideally within a GPUaaS or neocloud environment Deep expertise in Kubernetes and GPU scheduling frameworks including Slurm Kubernetes device plugins and NVIDIA GPU Operator Strong experience designing and operating MLOps tooling and ML pipeline orchestration in production Advanced proficiency in Python Bash and Infrastructure-as-Code tools such as Terraform and Ansible Strong understanding of GPU infrastructure and distributed training concepts including NCCL data/model parallelism and RDMA-aware scheduling Experience with platform-wide observability tooling such as Prometheus Grafana and telemetry stacks Strong Linux systems and networking fundamentals with the judgement to independently solve ambiguous cross-team problems A security-first mindset when operating within multi-tenant environments Awareness of EMEA regulatory and data residency considerations including GDPR as they relate to platform operations Strong communication and collaboration skills across distributed multi-region teams with the ability to mentor other engineers Experience with distributed training frameworks such as PyTorch and TensorFlow MLOps platforms including MLflow Kubeflow or Ray and InfiniBand/RDMA or RoCEv2 networking concepts is advantageous. Kubernetes certifications such as CKA CKAD or CKS and cloud certifications across AWS GCP or Azure are also advantageous. The role requires the right to work in an EMEA jurisdiction with existing eligibility to work across the EU/EEA or UK advantageous given the multi-country remit. Why Join Sharon AI Own the platform architecture powering a growing GPU-as-a-Service and AI neocloud business across EMEA Work hands-on with large-scale GPU infrastructure Kubernetes Slurm and AI-native platform technologies Shape how Sharon AI's platform scales across multiple jurisdictions sites and customer workloads Solve complex technical challenges across GPU infrastructure MLOps networking and distributed AI workloads Influence platform reliability automation capacity and cost optimisation across the region Work closely with Network Engineering and Infrastructure teams on high-performance InfiniBand/RDMA environments Provide technical leadership and mentorship while remaining hands-on as a senior individual contributor Help enterprise customers reliably train fine-tune and run AI/ML workloads at scale Join a highly technical and ambitious team operating at the forefront of AI infrastructure Our Values Integrity Innovation Collaboration Wellbeing Inclusion

Staff Platform Engineer | Service Mesh

Brazil Remote

Your wellbeing our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness mindfulness therapy nutrition and sleep—all in one simple subscription. Headquartered in NYC with team members in 11 countries we’re on a mission to make every company a wellness company. We believe work should be fulfilling inspiring and balanced. Here you’ll find a team that values wellbeing collaboration and different perspectives where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Staff Platform Engineer in our Platform Engineering team Brazil ! This is a Remote – Brazil position meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. The mission is to help us build a global secure recoverable and cost-efficient infrastructure that enables engineering teams to scale our product autonomously. Platform Engineering builds tooling and automation to eliminate operational processes or drive them to as near zero as possible. We want to enable the most reliable real-time logistics platform possible. We want the Wellhub technology stack to be completely autonomous but resilient and reliable. We are continuously evaluating infrastructure-related operational processes looking for opportunities to make it as frictionless as possible and delivering intuitive and reliable tools for other teams to use in their daily activities. We strive to positively impact a large number of engineers in our organization throughout our deliveries. You will have contact with the most popular and modern technologies in the cloud native industry such as Kubernetes Crossplane Kafka AWS Github Actions ArgoCD Hashicorp Vault Istio Prometheus and Grafana. You will also be able to code in Golang Ruby or Python to build and maintain our products and tools. YOUR IMPACT Help to build a global secure scalable and cost-effective Cloud platform using Kubernetes in AWS. Develop and evolve Kubernetes operators and other cloud-native automation in Kubernetes. Build products and tools enabling engineering teams to create and maintain their cloud resources autonomously. Build from zero to hero the Istio Service Mesh multi-cluster global platform Improve observability reliability and cost awareness. Support engineering teams in the products and tools adoption. Participate in the definition of standards RFCs (Request for Comments) guidelines and best practices. Live the mission inspire and empower others by genuinely caring for your own wellbeing and your colleagues. Bring wellbeing to the forefront of work and create a supportive environment where everyone feels comfortable taking care of themselves taking time off and finding work-life balance. WHO YOU ARE Proven technical experience with AWS cloud services Kubernetes and software engineering. Proven technical experience with Istio Service Mesh Deep knowledge of Kubernetes and its ecosystem Solid knowledge of observability systems Experience with operator-managed Infrastructure as Code preferably Crossplane or Kubernetes Operators. Excellent analytical and problem-solving skills and proven experience in identifying solutions for complex problems. Collaboration and learning-driven mindset CNCF Kubernetes Certifications (e.g. CKA CKS or CKAD) Excellent communication skills in both English and Portuguese both verbally and in writing We recognize that individuals approach job applications differently. We strongly encourage all aspiring applicants to go for it even if they don't match the job description 100%. We welcome your application and will be delighted to explore if you could be a great fit for our team. For this specific role please note that prior experience in AWS Kubernetes and Software Engineering are mandatory requirements. WHAT WE OFFER YOU With thoughtful benefits emotional wellbeing resources and a culture that empowers you to take ownership of your role and your wellbeing we create an environment where you can thrive in all dimensions of your life. Our flexible benefits program allows you to customize some of the benefits according to your needs! Our benefits include WELLHUB Free Gold+ membership with access to onsite gyms and studios digital fitness programs and online wellness resources for meditation nutrition mental wellbeing support and more! Add up to three family members to your plan ensuring access to wellness for those who matter most to you. WELLZ A complete emotional wellbeing program with a unique approach. It offers personalized journeys that combine individual therapy sessions (52 per year) and on-demand content. HEALTHCARE Health dental and life insurance. FLEXIBLE WORK As a Flexible First company we offer hybrid and remote options to give you the freedom to work in a way that suits you. The model for this specific role can be discussed with your recruiter and hiring manager. When you join use our home office reimbursement to set up your home office. FLEXIBLE SCHEDULE Flexibility for us isn’t just about where we work—it also means being able to shape how and when we get things done. Together with their leaders employees define schedules that align with their time zones team needs and personal routines. PAID TIME OFF It’s important to take time away from work to recharge.Employees receive vacations after 6 months and additional 3 days off per year + 1 day off for each year of tenure (up to 5 additional days) + an extra holiday for your birthday! PAID PARENTAL LEAVE Welcoming a new child is one of the most special moments in your life. Take the time to be present and enjoy your growing family. We offer 100% paid parental leave to all new parents. Parents giving birth are eligible for an extended leave and a ramp-back period to return part-time while they get settled. CAREER GROWTH Access world-class platforms participate in interactive sessions build your personalized development roadmap and explore internal opportunities. We focus on continuous learning and feedback to support your journey toward personal and professional success. CULTURE You’ll join a team of passionate people who come together to break boundaries support each other and create a meaningful impact in workplace wellness. We win together building trust through open communication and a culture where every perspective matters. Learn more about our shared culture and values here. And to get a glimpse of life at Wellhub… Follow us on Instagram @lifeatwellhub and LinkedIn Diversity Equity and Belonging at Wellhub We aim to create a collaborative supportive and inclusive space where everyone knows they belong. Wellhub is committed to creating a diverse work environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race religion color sex gender identity or expression sexual orientation age non-disqualifying physical or mental disability national origin veteran status or any other basis covered by appropriate law. Questions on how we treat your personal data? See our Aviso de Privacidade para Candidatos. LI-REMOTE LI-CM1

Staff Platform Engineer | Observability

Brazil Remote

"Your wellbeing our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness mindfulness therapy nutrition and sleep—all in one simple subscription. Headquartered in NYC with team members in 11 countries we’re on a mission to make every company a wellness company. We believe work should be fulfilling inspiring and balanced. Here you’ll find a team that values wellbeing collaboration and different perspectives where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Staff Platform Engineer with a dedicated focus on our Observability ecosystem for our Platform area in Brazil ! This is a Remote – Brazil position meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. In an environment of rapid growth and high-scale distributed architecture your mission is to transform Observability from a passive toolset into a strategic asset using open source standards. You will act as an architect of efficiency and reliability building a global platform that empowers engineering teams to ""own what they build"" with confidence. We are moving beyond basic monitoring to build a comprehensive ""Observability as a Service"" ecosystem. You will be responsible for evolving a self-service platform that balances performance with cost-effectiveness solving complex challenges related to high-cardinality metrics log retention strategies and distributed tracing. We strive to eliminate friction. You will design the ""Golden Paths"" that allow developers to instrument their code instantly and gain high-fidelity signals without operational overhead. You will have contact with the most popular and modern technologies in the cloud native industry such as Grafana OpenTelemetry Prometheus Thanos Loki Tempo Pyroscope Kafka Kubernetes AWS Github Actions ArgoCD Crossplane and Hashicorp Vault. You will also be able to code in Golang Ruby or Python to build and maintain our products and tools. YOUR IMPACT The team focuses on enabling engineering teams to own and run their decisions end-to-end rather than taking over support for a set of services. We are responsible for building the necessary tooling and abstractions that simplify service lifecycles. By fulfilling this role you can expect to Take full end-to-end ownership of our observability stack ensuring it remains resilient scalable and capable of supporting our rapid growth Evolve and maintain a cloud-native infrastructure based on Kubernetes ensuring that our foundation is always at the state-of-the-art Bridge the gap between complex infrastructure and developer experience building the necessary abstractions so teams can manage their own observability lifecycle seamlessly Contribute to establishing and maintaining standards guidelines and best practices and formalize them through processes like RFCs (Request for Comments). Live the mission inspire and empower others by genuinely caring for your own well-being and your colleagues. Bring wellbeing to the forefront of work and create a supportive environment where everyone feels comfortable taking care of themselves taking time off and finding work-life balance. WHO YOU ARE We are continuously evaluating infrastructure-related operational processes looking for opportunities to make it as frictionless as possible and delivering intuitive and reliable tools for other teams to use in their daily activities. We strive to positively impact a large number of engineers in our organization throughout our deliveries. This role may fit if you have Proven technical experience with observability practices and tooling (metrics logs traces profiling) Proficiency with a major cloud provider (AWS GCP or Azure) and its ecosystem of services for building cloud-native applications Deep knowledge of Kubernetes and its ecosystem Hands-on experience of tools like Prometheus Thanos Grafana Loki Tempo or similar open source solutions Hands-on experience with OpenTelemetry infrastructure and instrumentation Excellent analytical and problem-solving skills and proven experience in identifying solutions for complex problems. The ability to translate technical complexities into clear actionable guidance for different stakeholders fostering a culture of collaboration and knowledge sharing Bonus points if you have Experience building internal tools CLIs or Kubernetes operators to extend platform capabilities Familiarity with Crossplane to manage cloud resources through Kubernetes-native APIs A high sense of end-to-end ownership and find satisfaction in building tools that act as a catalyst for the success of other developers Service Mesh Knowledge of Istio Linkerd or similar technologies to enhance network observability and security FinOps Mindset Experience in cost-optimization strategies for high-volume telemetry and cloud spend. CNCF Kubernetes Certifications (e.g. CKA CKS or CKAD) CNCF Prometheus/OpenTelemetry Certifications(e.g PCAOTCA) We recognize that individuals approach job applications differently. We strongly encourage all aspiring applicants to go for it even if they don't match the job description 100%. We welcome your application and will be delighted to explore if you could be a great fit for our team. For this specific role please note that prior experience in AWS Kubernetes Software Engineering and Observability Practices with open source solutions are mandatory requirement. WHAT WE OFFER YOU With thoughtful benefits emotional wellbeing resources and a culture that empowers you to take ownership of your role and your wellbeing we create an environment where you can thrive in all dimensions of your life. Our flexible benefits program allows you to customize some of the benefits according to your needs! Our benefits include WELLHUB Free Gold+ membership with access to onsite gyms and studios digital fitness programs and online wellness resources for meditation nutrition mental wellbeing support and more! Add up to three family members to your plan ensuring access to wellness for those who matter most to you. WELLZ A complete emotional wellbeing program with a unique approach. It offers personalized journeys that combine individual therapy sessions (52 per year) and on-demand content. HEALTHCARE Health dental and life insurance. FLEXIBLE WORK As a Flexible First company we offer hybrid and remote options to give you the freedom to work in a way that suits you. The model for this specific role can be discussed with your recruiter and hiring manager. When you join use our home office reimbursement to set up your home office. PAID TIME OFF It’s important to take time away from work to recharge.Employees receive vacations after 6 months and additional 3 days off per year + 1 day off for each year of tenure (up to 5 additional days) + an extra holiday for your birthday! PAID PARENTAL LEAVE Welcoming a new child is one of the most special moments in your life. Take the time to be present and enjoy your growing family. We offer 100% paid parental leave to all new parents. Parents giving birth are eligible for an extended leave and a ramp-back period to return part-time while they get settled. CAREER GROWTH Access world-class platforms participate in interactive sessions build your personalized development roadmap and explore internal opportunities. We focus on continuous learning and feedback to support your journey toward personal and professional success. CULTURE You’ll join a team of passionate people who come together to break boundaries support each other and create a meaningful impact in workplace wellness. We win together building trust through open communication and a culture where every perspective matters. Learn more about our shared culture and values here. And to get a glimpse of life at Wellhub… Follow us on Instagram @lifeatwellhub and LinkedIn Diversity Equity and Belonging at Wellhub We aim to create a collaborative supportive and inclusive space where everyone knows they belong. Wellhub is committed to creating a diverse work environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race religion color sex gender identity or expression sexual orientation age non-disqualifying physical or mental disability national origin veteran status or any other basis covered by appropriate law. Our commitment to inclusion also extends to how we recognize and reward our people. We’re proud to be Syndio Fair Pay Certified reflecting our ongoing dedication to equitable and fair pay practices across our global team. Read more about it here. Questions on how we treat your personal data? See our Aviso de Privacidade para Candidatos. LI-REMOTE LI-CM1"

Staff Platform Engineer

Brazil Remote

Your wellbeing our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness mindfulness therapy nutrition and sleep—all in one simple subscription. Headquartered in NYC with team members in 11 countries we’re on a mission to make every company a wellness company. We believe work should be fulfilling inspiring and balanced. Here you’ll find a team that values wellbeing collaboration and different perspectives where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Staff Platform Engineer in our Platform Engineering team Brazil ! This is a Remote – Brazil position meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. The mission is to help us build a global secure recoverable and cost-efficient infrastructure that enables engineering teams to scale our product autonomously. Platform Engineering builds tooling and automation to eliminate operational processes or drive them to as near zero as possible. We want to enable the most reliable real-time logistics platform possible. We want the Wellhub technology stack to be completely autonomous but resilient and reliable. We are continuously evaluating infrastructure-related operational processes looking for opportunities to make it as frictionless as possible and delivering intuitive and reliable tools for other teams to use in their daily activities. We strive to positively impact a large number of engineers in our organization throughout our deliveries. You will have contact with the most popular and modern technologies in the cloud native industry such as Kubernetes Crossplane Kafka AWS Github Actions ArgoCD Hashicorp Vault Istio Prometheus and Grafana. You will also be able to code in Golang or Python to build and maintain our products and tools. YOUR IMPACT Help to build a global secure scalable and cost-effective Cloud platform using Kubernetes in AWS. Develop and evolve Kubernetes operators and other cloud-native automation in Kubernetes. Build products and tools enabling engineering teams to create and maintain their cloud resources autonomously. Help to ensure security and compliance by delivering secure products and implementing DevSecOps integrations. Improve observability reliability and cost awareness. Support engineering teams in the products and tools usage. Build and maintain a modern CI/CD set of tools and services. Keep all the Kubernetes clusters highly available and reliable. Contribute to our product documentation (e.g. user guide configurations operations and troubleshooting procedures) Participate in the definition of standards RFCs (Request for Comments) guidelines and best practices. Live the mission inspire and empower others by genuinely caring for your own wellbeing and your colleagues. Bring wellbeing to the forefront of work and create a supportive environment where everyone feels comfortable taking care of themselves taking time off and finding work-life balance. WHO YOU ARE Proven technical expertise in AWS cloud services Kubernetes and software engineering. In-depth knowledge of Kubernetes and its ecosystem. CNCF Kubernetes Certifications (e.g. CKA CKS or CKAD) will be considered a plus. Comprehensive understanding of observability systems. Experience with operator-managed Infrastructure as Code preferably cross-plane or Kubernetes Operators. Ability to write software for production environments. Excellent analytical and problem-solving skills and proven experience in identifying solutions for complex problems. Collaboration and learning-driven mindset Deep expertise in the AWS ecosystem and if you have AWS certifications it will be considered a plus. Excellent communication skills in both English and Portuguese both verbally and in writing Might suit you very well if you are interested in Changing and improving things. Solving complex problems gracefully. Reading and writing good documentation. Staying current on industry-leading practices and technologies. Building something that you are proud of and would like to work with. We recognize that individuals approach job applications differently. We strongly encourage all aspiring applicants to go for it even if they don't match the job description 100%. We welcome your application and will be delighted to explore if you could be a great fit for our team. For this specific role please note that prior experience in AWS Kubernetes and containers are mandatory requirements (change this last sentence according to the core requirements of the role). WHAT WE OFFER YOU With thoughtful benefits emotional wellbeing resources and a culture that empowers you to take ownership of your role and your wellbeing we create an environment where you can thrive in all dimensions of your life. Our flexible benefits program allows you to customize some of the benefits according to your needs! Our benefits include WELLHUB Free Gold+ membership with access to onsite gyms and studios digital fitness programs and online wellness resources for meditation nutrition mental wellbeing support and more! Add up to three family members to your plan ensuring access to wellness for those who matter most to you. WELLZ A complete emotional wellbeing program with a unique approach. It offers personalized journeys that combine individual therapy sessions (52 per year) and on-demand content. HEALTHCARE Health dental and life insurance. FLEXIBLE WORK As a Flexible First company we offer hybrid and remote options to give you the freedom to work in a way that suits you. The model for this specific role can be discussed with your recruiter and hiring manager. When you join use our home office reimbursement to set up your home office. FLEXIBLE SCHEDULE Flexibility for us isn’t just about where we work—it also means being able to shape how and when we get things done. Together with their leaders employees define schedules that align with their time zones team needs and personal routines. PAID TIME OFF It’s important to take time away from work to recharge.Employees receive vacations after 6 months and additional 3 days off per year + 1 day off for each year of tenure (up to 5 additional days) + an extra holiday for your birthday! PAID PARENTAL LEAVE Welcoming a new child is one of the most special moments in your life. Take the time to be present and enjoy your growing family. We offer 100% paid parental leave to all new parents. Parents giving birth are eligible for an extended leave and a ramp-back period to return part-time while they get settled. CAREER GROWTH Access world-class platforms participate in interactive sessions build your personalized development roadmap and explore internal opportunities. We focus on continuous learning and feedback to support your journey toward personal and professional success. CULTURE You’ll join a team of passionate people who come together to break boundaries support each other and create a meaningful impact in workplace wellness. We win together building trust through open communication and a culture where every perspective matters. Learn more about our shared culture and values here. And to get a glimpse of life at Wellhub… Follow us on Instagram @lifeatwellhub and LinkedIn Diversity Equity and Belonging at Wellhub We aim to create a collaborative supportive and inclusive space where everyone knows they belong. Wellhub is committed to creating a diverse work environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race religion color sex gender identity or expression sexual orientation age non-disqualifying physical or mental disability national origin veteran status or any other basis covered by appropriate law. Questions on how we treat your personal data? See our Aviso de Privacidade para Candidatos. LI-REMOTE LI-CM1

Kubernetes Technical Support Engineer (565)

Remote

" About the opportunity Join an enterprise AI vision platform company to troubleshoot and resolve complex Kubernetes infrastructure and deployment challenges for high-impact video analytics software. About our client Our client is an enterprise AI vision platform company specializing in high-impact video analytics software. What you'll do Provide Tier 2 / Tier 3 technical support to enterprise customers. Troubleshoot Kubernetes-based deployments software infrastructure hardware networking and configuration issues. Reproduce customer environments and run tests to diagnose technical problems. Work directly with customers during technical discussions and troubleshooting sessions. Manage and resolve technical support tickets. Collaborate with Engineering Sales and Customer Success teams. Escalate complex issues and product bugs to Engineering and help drive them to resolution. Clearly communicate technical findings workarounds and resolutions to customers. What you bring (must-haves) 2+ years of technical support experience preferably in Tier 2 / Tier 3 support. Strong hands-on Kubernetes experience in production environments. At least one relevant Kubernetes or cloud certification such as CKA CKAD CKS AWS Azure or equivalent. Strong troubleshooting skills across software infrastructure networking and distributed environments. Experience working directly with customers in a technical capacity. Strong written and verbal English communication skills (conversational English is mandatory). Ability to independently investigate complex technical issues and collaborate with engineering teams. LATAM-based candidate available to work U.S. East Coast business hours. Bonus points for (nice-to-haves) Experience with video analytics software AI vision platforms or media streaming/processing infrastructure. Perks & benefits 💻 Equipment provided — none of that ""bring your own device"" stuff here 🛡️ Full back-office support — Legal Accounting HR Business Partner and Delivery team 🧭 Career and cultural mentoring — how to show up stand out and navigate US work culture 🗣️ Free English lessons with a native speaker 🤝 Referral bonus — recommend Ubi to your tech friends and get paid for it 🏖️ Florianópolis HQ always open — 100% remote but the office is there whenever you want it. How the process works Interview with our Tech Recruiter (+ quick AI assessment if needed) Client interview process (varies by company) Offer 🎉 Onboarding with full Ubiminds support Why Ubiminds? With 9 years in the market and GPTW-certified Ubiminds connects Latin American tech professionals with software companies in the US and Canada. Hundreds of professionals in data design product and engineering are already growing in North American teams with our support. We're with you throughout the entire journey — Recruitment Legal Accounting PeopleOps and career guidance — so you can thrive internationally with confidence. ➡ ➡"

AI Platform Engineer

Hybrid

About PayPay Card PayPay Card Corporation was established in 2021 to provide users a FinTech service that is more accessible and convenient compared to previous credit cards and credit services by integrating with the PayPay payment platform which has surpassed 75 million users since its launch (as of August 2026). We are looking for people who are passionate about refining our products at an overwhelming speed that other companies cannot match as well as professionals who are interested in promoting the spread of cashless payments in Japan and the use of these payments as a financial life platform. Let us work together to create new value for users. ※ Please note that you cannot apply or be selected in parallel with PayPay Corporation PayPay Card Corporation and PayPay Securities Corporation. Job Description PayPay Card is looking for an AI Platform Engineer focused on cloud-native GenAI infrastructure and enablement. This role will build and operate the foundation that enables internal teams to deliver and operate GenAI applications agents RAG systems and related AI workloads reliably safely and cost-effectively Responsibilities Architect and build AI platform capabilities for applications agents RAG systems and related AI workloads Architect and build infrastructure that is easy to maintain update and improve Architect and build infrastructure with appropriate reliability and recovery capabilities for internal AI platform services Work together with our Security Engineers to provision secure and governed AI platform infrastructure Build and maintain deployment automation to ensure fast delivery of AI platform services to our developers Provide self-service capabilities and standard deployment patterns for developers to easily deploy and operate AI-powered application infrastructure Build and maintain reusable platform templates deployment patterns and integrations for GenAI applications agents RAG systems MCP-based integrations and agent-to-agent workflows Build and support monitoring and evaluation capabilities for GenAI systems including usage cost reliability agent execution and adoption metrics Continuously research evaluate and prototype emerging AI trends frameworks and open-source tools to ensure the platform remains cutting-edge. Drive R&D initiatives for new AI platform capabilities keeping pace with the rapid evolution of agentic workflows and LLM infrastructure. Tech Stack AWS Bedrock Bedrock Knowledge Bases OpenSearch Neptune S3 ECS EKS Lambda CloudWatch Cognito SQS KMS Secrets Manager MSK CodeCommit CodeBuild CodeDeploy CodePipeline CloudFormation and other services AI platform / GenAI capabilities RAG vector stores graph databases model access patterns MCP-based integrations agent orchestration agent-to-agent workflows evaluation and observability tooling Terraform GitHub Actions Prometheus Grafana Dynatrace Atlantis ArgoCD OpenTelemetry Required Qualifications More than 5 years of technical experience in cloud-based infrastructure platforms Ability to demonstrate high degree of ownership in a Production environment Good understanding of cloud security best practices and payment industry compliance standards Experience designing building and operating cloud platform capabilities for internal developers Experience with cloud infrastructure and platform systems availability performance and cost management Extensive technical hands-on experience with compute storage and analytics services on cloud platforms Experience with IaC tools such as Terraform CloudFormation CDK Experience with cloud services monitoring detection and response Experience with cloud services performance tuning cost controls and management Experience in cloud infrastructure service patching and upgrades Familiarity with AI platform concepts such as GenAI applications agents RAG systems vector stores model access patterns and evaluation/observability capabilities PayPay DevOps emphasize automation. Demonstrated skill with the following are required Have excellent oral written verbal and interpersonal communication skills Preferred Qualifications Bachelor’s degree and above in a technology related field Experience with other cloud service providers (e.g. GCP Azure) Experience with Kubernetes (CKA CKAD or CKS) Experience with AWS AI services such as Bedrock Bedrock Knowledge Bases Bedrock AgentCore Bedrock Prompt Management or similar services Experience with RAG systems vector stores graph databases semantic search or knowledge management platforms Experience with MCP agent orchestration agent-to-agent workflows or related AI integration patterns Experience with agent frameworks or orchestration tools such as OpenAI Agents SDK Google ADK Strands Agents LangGraph CrewAI LlamaIndex or similar Experience with monitoring evaluation or observability tooling for AI-powered systems Experience with Event-Driven Architecture (Kafka preferred) Experience using and contributing to Open Source tools Experience in managing IT compliance and security risk Demonstrated track record of self-driven learning and a passion for continuously catching up with the rapidly evolving AI ecosystem. Experience conducting R&D or building proofs-of-concept (PoCs) for emerging AI technologies. Active engagement with the AI community—evidenced by published papers technical blogs open-source contributions or personal AI hobby projects. Bilingual in English and Japanese is nice to have but not required. Proficiency in either language is fine. Working Conditions Employment Status Full Time Office Location Hybrid Workstyle (flexible working style including Remote and office) ※ You will be expected to work both in the office and remotely in alignment with organizational guidelines and team objectives. LIFE in JAPAN FACTBOOK Work Hours Full Flex Time (No Core Time) In principle 900am ~ 545pm (actual working hours 7h45m + 1h break) Holidays Every Sat/Sun/National holidays (In Japan)/New Year's break/Company-designated Special days Paid leave Annual leave (up to 14 days in the first year granted proportionally according to the month of employment. Can be used from the date of hire) Personal leave (5 days each year granted proportionally according to the month of employment) *PayPay Group's own special paid leave system which can be used to attend to illnesses injuries hospital visits etc. of the employee family members pets etc. Salary Annual salary paid in 12 installments (monthly) Reviewed once a year Overtime allowance Late overtime allowance Commuting and transportation expenses Benefits Social Insurance (health insurance employee pension employment insurance and compensation insurance) 401K Other Information PayPay Inside-Out (Corporate Blog) ENG Recruiting FACTBOOK for PayPay Card

unlock: sign-up for free / login and use the searches from your home page
🔥 job listings updated in real time

For 94 similar CKAD position(s) we've listed in the previous 30 days we've processed the salary ranges data as posted by employers in job descriptions and the resulting overall range is: 82K - 252.2K USD.

View highest-paying job

Note: If any discrepancies or 'wild' numbers appear this could be because of: typos in job postings, data source errors, ghost jobs, etc. We do not modify salary data inserted in JDs by employers nor use estimates.


Login & search by other job titles, a specific location or any keyword.
Additional custom search filters are available once you login.