I

Senior Systems Engineer – Communication Stack for Large-Scale AI Training

Visa Sponsorship
Apply Now

The role focuses on co-designing and optimizing the communication stack for large‑scale distributed AI training. You will design high‑performance collective operations, improve fault tolerance, and drive performance across thousands of GPUs. Requires deep expertise in RDMA/InfiniBand, NCCL/UCX, systems programming (C/C++, Rust/Go) and experience with PyTorch at massive scale.

Key Highlights
Co‑design communication stack for massive GPU clusters
Optimize hierarchical collectives for Mixture‑of‑Experts workloads
Build fault‑tolerant, low‑latency distributed execution
Key Responsibilities
Design and optimize expert‑parallel and hybrid‑parallel communication patterns.
Drive high‑performance hierarchical collectives for Mixture‑of‑Experts workloads.
Co‑design runtime orchestration with communication topology awareness.
Reduce tail latency and improve determinism across thousands of GPUs.
Architect fault‑tolerant distributed execution under real‑world cluster failures.
Perform deep debugging of NCCL, RDMA, and custom communication layers.
Analyze congestion and optimize routing across InfiniBand/RoCE fabrics.
Microbenchmark and model performance for communication‑heavy workloads.
Technical Skills Required
RDMA/InfiniBand networking Systems programming (C/C++, Rust, Go) PyTorch
Benefits & Perks
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off, sick leave and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability

Job Description

The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.
Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

omiz staffing solutions (oss)

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

oddity labs

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

institute of foundation models

United State

Subscribe our newsletter

New Things Will Always Update Regularly