I

Senior Distributed ML Infrastructure Engineer

Visa Sponsorship
Apply Now

Extend and scale distributed training frameworks for large-scale foundation model pre-training. Implement distributed optimizers, robust launch systems, and monitoring tools for multi-node GPU clusters. Requires 5+ years of experience in ML systems, distributed training, and strong software engineering fundamentals.

Key Highlights
Work on core cutting-edge foundation model training infrastructure
Extend distributed frameworks like DeepSpeed and FSDP
Implement distributed optimizers from mathematical specs
Own experiment tracking and job monitoring for multi-node clusters
Key Responsibilities
Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures
Translate mathematical optimizer specs into distributed implementations
Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets
Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers
Write production-quality code and tests for ML infra in PyTorch or JAX
Build and maintain data-loading and checkpoint/restart workflows
Technical Skills Required
Distributed Machine Learning Python PyTorch
Benefits & Perks
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off, sick leave, and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability
Visa sponsorship
Nice to Have
Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
Familiarity with performance profiling, kernel fusion, or memory optimization
Open-source contributions or published research (MLSys, ICML, NeurIPS)
CUDA or Triton kernel experience
Experience building custom training pipelines at scale
Deep familiarity with training infrastructure and performance tuning

Job Description

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

wayve

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

wayve

United State

RL Infrastructure Engineer

Machine Learning
•
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

institute of foundation models

United State

Subscribe our newsletter

New Things Will Always Update Regularly