AI Platform Engineer

Sharon AI, Inc All jobs
12 hour(s) ago
Job Overview
Company Sharon AI, Inc
Job Type Full-time
Job Level entry
Category Engineering and Information Technology
Posted 2026-08-28
Last Seen 12 hour(s) ago

Job Description

AI Platform Engineer L3 (GPUaaS – AI Neocloud) 📍 EMEA Remote-first About Sharon AI Sharon AI is building the infrastructure powering the next generation of artificial intelligence. Operating across AI infrastructure high-performance compute cloud platforms and large-scale technology environments Sharon AI delivers scalable secure and reliable infrastructure for demanding AI ML and HPC workloads.

The Role

As an AI Platform Engineer L3 you'll design build and operate the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering across the EMEA region. You'll own the architecture automation and reliability of the platform services sitting above Sharon AI's GPU and network fabric spanning Kubernetes Slurm GPU scheduling MLOps tooling model serving and platform observability across multiple EMEA sites. Reporting to the Head of Operations you'll work closely with Network Engineering Infrastructure and customer-facing teams to solve complex cross-team challenges and ensure Sharon AI's platform can scale reliably and efficiently across the region. This is a hands-on senior individual contributor role suited to an engineer with strong platform engineering DevOps or MLOps experience who can operate independently in a fast-paced AI-native environment and provide technical guidance to less experienced engineers.

Key Responsibilities

Design and own the AI platform architecture across Sharon AI's EMEA GPU clusters including Kubernetes Slurm and container orchestration Lead the development of CI/CD pipelines and MLOps tooling supporting training fine-tuning and inference workloads across multiple sites Define and implement multi-tenant GPU resource scheduling quota management and workload isolation strategies at scale Own the design of model serving infrastructure balancing high availability performance and cost efficiency Build and evolve platform-wide observability across monitoring logging and alerting covering platform health GPU utilisation and workload performance Drive Infrastructure-as-Code adoption and platform automation using Terraform and Ansible Partner closely with Network Engineering to integrate the platform layer with Sharon AI's underlying InfiniBand/RDMA fabric Act as a senior escalation point for complex platform issues impacting customer AI/ML workloads across EMEA Lead incident response and post-incident reviews contributing to operational runbooks and platform best practice Partner with the Head of Operations on platform capacity planning scaling strategy and cost optimisation across EMEA Mentor and provide technical guidance to less experienced platform engineers Support enterprise GPUaaS customer onboarding and technical escalations across the region

Skills

& Experience 6–10+ years' experience in platform engineering DevOps MLOps or SRE ideally within HPC cloud or AI/ML infrastructure environments Bachelor's degree in Computer Science Electrical Engineering or a related field Hands-on experience operating Kubernetes and GPU scheduling at production scale Proven experience designing CI/CD and Infrastructure-as-Code practices for platform teams Proven experience supporting GPU or AI/ML infrastructure at scale ideally within a GPUaaS or neocloud environment Deep expertise in Kubernetes and GPU scheduling frameworks including Slurm Kubernetes device plugins and NVIDIA GPU Operator Strong experience designing and operating MLOps tooling and ML pipeline orchestration in production Advanced proficiency in Python Bash and Infrastructure-as-Code tools such as Terraform and Ansible Strong understanding of GPU infrastructure and distributed training concepts including NCCL data/model parallelism and RDMA-aware scheduling Experience with platform-wide observability tooling such as Prometheus Grafana and telemetry stacks Strong Linux systems and networking fundamentals with the judgement to independently solve ambiguous cross-team problems A security-first mindset when operating within multi-tenant environments Awareness of EMEA regulatory and data residency considerations including GDPR as they relate to platform operations Strong communication and collaboration skills across distributed multi-region teams with the ability to mentor other engineers Experience with distributed training frameworks such as PyTorch and TensorFlow MLOps platforms including MLflow Kubeflow or Ray and InfiniBand/RDMA or RoCEv2 networking concepts is advantageous. Kubernetes certifications such as CKA CKAD or CKS and cloud certifications across AWS GCP or Azure are also advantageous. The role requires the right to work in an EMEA jurisdiction with existing eligibility to work across the EU/EEA or UK advantageous given the multi-country remit. Why Join Sharon AI Own the platform architecture powering a growing GPU-as-a-Service and AI neocloud business across EMEA Work hands-on with large-scale GPU infrastructure Kubernetes Slurm and AI-native platform technologies Shape how Sharon AI's platform scales across multiple jurisdictions sites and customer workloads Solve complex technical challenges across GPU infrastructure MLOps networking and distributed AI workloads Influence platform reliability automation capacity and cost optimisation across the region Work closely with Network Engineering and Infrastructure teams on high-performance InfiniBand/RDMA environments Provide technical leadership and mentorship while remaining hands-on as a senior individual contributor Help enterprise customers reliably train fine-tune and run AI/ML workloads at scale Join a highly technical and ambitious team operating at the forefront of AI infrastructure

Our Values

Integrity Innovation Collaboration Wellbeing Inclusion

Unlock: Sign Up for free / Sign In and use the searches from your home page or the links in the footer.