Build and scale distributed pre-training frameworks for foundation models using PyTorch, JAX, and DeepSpeed/FSDP. Optimize training efficiency through mixed-precision, kernel fusion, and custom CUDA/Triton implementations. Requires 5+ years of experience in large-scale deep learning and distributed training on 100+ GPU clusters.
Key Highlights
Lead large-scale transformer pre-training runs on multi-node GPU clusters (100+ GPUs).
Implement custom optimizers and attention methods, converting them into efficient CUDA/Triton kernels.
Own mixed-precision training paths (bf16, fp8) and optimize for state-of-the-art throughput and stability.
Key Responsibilities
Build and scale distributed pre-training frameworks using DeepSpeed, FSDP, or Megatron-LM across multi-node GPU clusters.
Create robust launch scripts, resilient checkpoints, and job monitoring systems for NCCL/GLOO/GPU.
Prototype new optimizers or attention methods and convert them into efficient CUDA/Triton kernels with custom gradients.
Lead mixed-precision training efforts, tracking accuracy-vs-speed gains and analyzing numeric stability.
Apply kernel fusion, communication tuning, and memory optimization to reach state-of-the-art throughput.
Build logging, metrics, and experiment-tracking tools to accelerate research velocity.
Design ablation studies and statistical tests to validate new ideas.
Mentor interns and junior engineers through clear async design docs and code reviews.
Technical Skills Required
PyTorch
Distributed Training
CUDA
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Benefits & Perks
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off, sick leave and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability
Visa sponsorship
Nice to Have
NeurIPS / ICML / ICLR papers or open-source contributions to major ML frameworks
Experience implementing optimization algorithms (e.g., SGD variants, Adam, second-order methods)
Background in numerical computing
Ability to translate math and build high-perf CUDA/Triton kernels
Job Description
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.