Machine Learning Engineer – Distributed ML Systems
Pluralis Research is seeking Senior/Staff Machine Learning Engineers with strong experience in distributed systems and large-scale machine learning training. The team is developing a novel substrate for training distributed ML models across consumer-grade internet connections as part of its Protocol Learning research. In this role you will design implement and optimize distributed machine learning systems operating across heterogeneous hardware and challenging network conditions. You will work on model parallelism GPU and memory optimization decentralized networking resilience checkpointing synchronization and monitoring for large-scale training workloads.
Design and implement large-scale distributed training systems optimized for heterogeneous hardware low-bandwidth connections and high-latency environments. Develop and optimize model-parallel training strategies including data tensor and pipeline parallelism. Implement custom sharding techniques to minimize communication overhead. Optimize GPU utilization memory efficiency and compute performance across distributed nodes. Implement checkpointing state synchronization and recovery mechanisms for long-running and fault-prone training jobs. Build monitoring and metrics systems to track training progress model quality and system bottlenecks. Architect resilient training systems that support node failures network partitions and participants dynamically joining or leaving. Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes. Implement NAT traversal peer discovery dynamic routing and connection lifecycle management. Profile and optimize communication patterns to reduce latency and bandwidth overhead in multi-participant environments.
5+ years of experience in distributed systems and large-scale machine learning training. Strong experience building and operating distributed systems in production. Hands-on expertise with distributed training frameworks such as FSDP DeepSpeed Megatron or similar. Deep understanding of model parallelism including data tensor and pipeline parallelism. Expert-level Python experience in production environments including concurrency error handling retry logic and clean architecture. Strong networking fundamentals including P2P systems gRPC routing NAT traversal and distributed coordination. Experience optimizing GPU workloads memory management and large-scale compute efficiency.
The provided job description does not specify additional preferred qualifications.
Distributed machine learning systems Large-scale ML training Distributed systems architecture Model parallelism Data parallelism Tensor parallelism Pipeline parallelism FSDP DeepSpeed or Megatron Python Concurrency Error handling and retry logic GPU optimization Memory management Compute performance optimization Peer-to-peer networking gRPC Routing NAT traversal Peer discovery Distributed coordination State synchronization Checkpointing and recovery Monitoring and metrics Communication optimization System resilience Technical problem-solving Education & Experience Education No specific education requirement is stated in the provided job description. Experience 5+ years of experience in distributed systems and large-scale ML training. Production experience building and operating distributed systems. Hands-on experience with distributed ML training frameworks. Experience optimizing GPU workloads memory and large-scale compute systems. Work Arrangement & Schedule Listed
San Francisco California United States Work Arrangement Remote / Remote-first Role Level Senior/Staff Employment Type Not specified in the provided job description. Weekly Hours Not specified. Schedule Flexibility Not specified. Weekend
Not specified. Additional
Information The job description states that the company is remote-first with optional access to a Melbourne hub. It also states that competitive base salary is offered for senior engineering roles in Australia which differs from the listed San Francisco location.
Equity-heavy compensation with meaningful ownership in the company. Competitive base salary for senior engineering roles in Australia. Visa sponsorship available for exceptional candidates. Remote-first work environment with optional access to the Melbourne hub. Opportunity to work with a team whose members have previously worked at Google Amazon Microsoft and leading startups. Compliance /
Pluralis Research conducts foundational research on Protocol Learning described as multi-participant training of foundation models where no single participant has or can ever obtain a full copy of the model. The company states that its goal is to facilitate the creation of community-trained and community-owned frontier models with self-sustaining economics.
Unlock: Sign Up for free / Sign In and use the searches from your home page or the links in the footer.