Extend and scale distributed training frameworks for large-scale foundation model pre-training. Implement distributed optimizers, robust launch systems, and monitoring tools for multi-node GPU clusters. Requires 5+ years of experience in ML systems, distributed training, and strong software engineering fundamentals.
Key Highlights
Work on core cutting-edge foundation model training infrastructure
Extend distributed frameworks like DeepSpeed and FSDP
Implement distributed optimizers from mathematical specs
Own experiment tracking and job monitoring for multi-node clusters
Key Responsibilities
Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures
Translate mathematical optimizer specs into distributed implementations
Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets
Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers
Write production-quality code and tests for ML infra in PyTorch or JAX
Build and maintain data-loading and checkpoint/restart workflows
Technical Skills Required
Distributed Machine Learning
Python
PyTorch
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Benefits & Perks
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off, sick leave, and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability
Visa sponsorship
Nice to Have
Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
Familiarity with performance profiling, kernel fusion, or memory optimization
Open-source contributions or published research (MLSys, ICML, NeurIPS)
CUDA or Triton kernel experience
Experience building custom training pipelines at scale
Deep familiarity with training infrastructure and performance tuning
Job Description
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.