T

Senior/Staff Machine Learning Engineer - Distributed Systems

torentify United State
Remote Visa Sponsorship
Apply
AI Summary

Design and implement large-scale distributed training systems for heterogeneous hardware under low-bandwidth conditions. Develop resilient peer-to-peer networking architectures and optimize model parallelism strategies for decentralized ML. Requires 5+ years of experience in distributed systems, large-scale ML training, and production Python.

Key Highlights
Build novel technical foundation for training distributed ML models over consumer-grade internet connections.
Architect resilient systems supporting dynamic node joining/leaving and operating despite network partitions.
Optimize GPU utilization, memory efficiency, and communication overhead across distributed nodes.
Key Responsibilities
Design and implement large-scale distributed training systems for heterogeneous hardware operating under low-bandwidth and high-latency network conditions.
Develop and optimize model-parallel training strategies, including data, tensor, and pipeline parallelism.
Implement custom sharding techniques designed to minimize communication overhead.
Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.
Build robust checkpointing, state synchronization, and recovery mechanisms for long-running and fault-prone training jobs.
Develop monitoring and metrics systems to track training progress, model quality, and system bottlenecks.
Architect resilient distributed training systems that can operate despite node failures and network partitions.
Support systems where participants can dynamically join or leave the network.
Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes.
Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.
Profile and optimize communication patterns to reduce latency and bandwidth usage across multi-participant environments.
Technical Skills Required
Distributed Systems Python Machine Learning
Benefits & Perks
Equity-heavy compensation
Visa sponsorship available for exceptional candidates
Remote-first work with optional access to the Melbourne hub

Job Description



## About the Company


Pluralis Research conducts foundational research into Protocol Learning, a distributed approach to training foundation models in which no single participant has or can obtain a complete copy of the model.


The company is focused on enabling community-trained and community-owned frontier models with self-sustaining economics. Pluralis Research is backed by Union Square Ventures and other tier-1 investors and brings together ML researchers and engineers with experience from leading technology companies and startups.


## About the Role


Pluralis Research is seeking **Senior/Staff Machine Learning Engineers** with **5+ years of experience in distributed systems and large-scale machine learning training**. In this role, you will help build a novel technical foundation for training distributed ML models over consumer-grade internet connections.


The role is suited to engineers with deep expertise in **distributed systems, distributed ML training, model parallelism, networking, GPU optimization, and production Python**. You will design systems capable of operating across heterogeneous hardware and unreliable networks while maintaining performance, resilience, and efficient communication between participants.


## Key Responsibilities


### Distributed Training Architecture & Optimization


* Design and implement large-scale distributed training systems for heterogeneous hardware operating under low-bandwidth and high-latency network conditions.

* Develop and optimize model-parallel training strategies, including data, tensor, and pipeline parallelism.

* Implement custom sharding techniques designed to minimize communication overhead.

* Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.

* Build robust checkpointing, state synchronization, and recovery mechanisms for long-running and fault-prone training jobs.

* Develop monitoring and metrics systems to track training progress, model quality, and system bottlenecks.


### Decentralized Networking & Resilience


* Architect resilient distributed training systems that can operate despite node failures and network partitions.

* Support systems where participants can dynamically join or leave the network.

* Design and optimize peer-to-peer topologies for decentralized coordination across non-co-located nodes.

* Implement NAT traversal, peer discovery, dynamic routing, and connection lifecycle management.

* Profile and optimize communication patterns to reduce latency and bandwidth usage across multi-participant environments.


## Required Qualifications


* **5+ years of experience in distributed systems and large-scale machine learning training.**

* Strong experience building and operating distributed systems in production.

* Hands-on experience with distributed training frameworks such as **FSDP, DeepSpeed, Megatron, or similar**.

* Deep understanding of **model parallelism**, including data, tensor, and pipeline parallelism.

* Expert-level **Python** experience in production environments.

* Experience with Python concurrency, error handling, retry logic, and clean software architecture.

* Strong networking fundamentals, including **P2P systems, gRPC, routing, NAT traversal, and distributed coordination**.

* Experience optimizing GPU workloads, memory management, and large-scale compute efficiency.


## Preferred Qualifications


The original job description does not specify separate preferred or nice-to-have qualifications. All qualifications listed above are presented as requirements in the source description.


## Skills & Competencies


* Distributed Systems

* Distributed Machine Learning

* Large-Scale ML Training

* Machine Learning Engineering

* Model Parallelism

* Data Parallelism

* Tensor Parallelism

* Pipeline Parallelism

* FSDP

* DeepSpeed

* Megatron

* Python

* GPU Optimization

* GPU Utilization

* Memory Management

* Large-Scale Computing

* Custom Model Sharding

* Checkpointing

* State Synchronization

* Fault Recovery

* Monitoring and Metrics

* Peer-to-Peer (P2P) Networking

* gRPC

* Routing

* NAT Traversal

* Peer Discovery

* Dynamic Routing

* Distributed Coordination

* Connection Lifecycle Management

* Network Optimization

* Latency and Bandwidth Optimization

* Production Systems

* Resilient Infrastructure


## Education & Experience


**Education:**

No specific education or degree requirement is stated in the original job description.


**Experience:**


* 5+ years of experience in distributed systems and large-scale ML training.

* Production experience building and operating distributed systems.

* Hands-on experience with distributed ML training frameworks and model-parallel training.

* Production-level Python experience.

* Experience with distributed networking and GPU workload optimization.


## Work Arrangement & Schedule


* **Location:** California, Missouri, United States

* **Work Arrangement:** Remote-first

* **Employment Type:** Senior/Staff engineering role

* **Remote Work:** Remote-first

* **Optional Office Access:** Melbourne hub

* No specific weekly hours or weekend requirements are stated in the original job description.


**Location note:** The source listing identifies the job location as **California, MO, US**, while the company description states that the role is remote-first with optional access to a **Melbourne hub**. The source also mentions visa sponsorship for exceptional candidates and a competitive base salary for senior engineering roles in Australia. These location-related details appear inconsistent, so the listed job location and remote-first arrangement have been preserved without inventing a correction.


## Compensation & Benefits


The original job description does not provide a specific salary range.


It states that the company offers:


* Equity-heavy compensation with meaningful ownership in a mission-driven company

* Competitive base salary for senior engineering roles in Australia

* Visa sponsorship available for exceptional candidates

* Remote-first work with optional access to the Melbourne hub

* Opportunity to work with a technical team whose members have experience at Google, Amazon, Microsoft, and leading startups


## Compliance / Additional Information


Pluralis Research describes Protocol Learning as an approach intended to support community-trained and community-owned frontier models and reduce concentration of model development, access, and economic value among a small number of large corporations.


The company is backed by **Union Square Ventures and other tier-1 investors**.



Similar Jobs

Explore other opportunities that match your interests

Mid-Level PHP Developer

Programming
12h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

online jobs philippines - work...

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Medtronic

United State

Senior Lead AI Engineer

Programming
16h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State

Subscribe our newsletter

New Things Will Always Update Regularly