The role focuses on co-designing and optimizing the communication stack for large‑scale distributed AI training. You will design high‑performance collective operations, improve fault tolerance, and drive performance across thousands of GPUs. Requires deep expertise in RDMA/InfiniBand, NCCL/UCX, systems programming (C/C++, Rust/Go) and experience with PyTorch at massive scale.
Key Highlights
Co‑design communication stack for massive GPU clusters
Optimize hierarchical collectives for Mixture‑of‑Experts workloads
Build fault‑tolerant, low‑latency distributed execution
Key Responsibilities
Design and optimize expert‑parallel and hybrid‑parallel communication patterns.
Drive high‑performance hierarchical collectives for Mixture‑of‑Experts workloads.
Co‑design runtime orchestration with communication topology awareness.
Reduce tail latency and improve determinism across thousands of GPUs.
Architect fault‑tolerant distributed execution under real‑world cluster failures.
Perform deep debugging of NCCL, RDMA, and custom communication layers.
Analyze congestion and optimize routing across InfiniBand/RoCE fabrics.
Microbenchmark and model performance for communication‑heavy workloads.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Technical Skills Required
RDMA/InfiniBand networking
Systems programming (C/C++, Rust, Go)
PyTorch
Benefits & Perks
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off, sick leave and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability
Job Description
The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.