Senior Distributed ML Infrastructure Engineer
Own the distributed training and inference backbone for a foundation model trained from scratch. Design, deploy, and maintain large-scale ML clusters and pipelines at petabyte scale. Requires 2-10 years of hands-on experience with foundation model infrastructure and deep GPU optimization expertise.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
San Francisco, CA
- On-site (5 days/week)
- Full-time Compensation: $200K–$400K + competitive early-stage equity
Our client is a Series A AI research lab building large-scale foundation models for scientific and physical-AI domains. Backed by top-tier investors, they are pursuing a deliberately non-consensus technical thesis and are among the best-funded teams in their space. The founding team comes from self-driving, robotics, and scientific research, and they are scaling their research and engineering org significantly this year.
Founded 2024
- Small, fast-growing team
- Industry: AI / foundation models / physical AI
Looking to advance your Machine Learning & AI career with relocation support? Explore Machine Learning & AI Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
What You'll Be Doing
- Design, deploy, and maintain large distributed ML training and inference clusters
- Build efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycle
- Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales
- Profile and debug low-level GPU operations to optimize performance
- Track new research and bring fresh ideas into the work
Requirements
- 2–10 years building large-scale ML infrastructure for core foundation models
- Hands-on experience building infrastructure for foundation models trained from scratch, rather than fine-tuning existing models
- A background at a science-focused or physical-AI company (for example self-driving, robotics, or biology)
- Deep, demonstrable expertise optimizing large-scale training and inference workloads
- Working proficiency with distributed training frameworks such as FSDP or DeepSpeed
- A clear pattern of intentional, mission-driven career decisions
- Able to work on-site 5 days/week in San Francisco (relocation supported)
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Generalist experience spanning the full ML lifecycle
- Low-level GPU performance optimization and debugging (CUDA, JAX)
Interested in relocating to United State? Check out our comprehensive Relocation Jobs in United State page with detailed relocation packages and benefits.
- Take a bet on a distinctive, non-consensus approach to building intelligence
- Join early, with real ownership of the training and inference backbone
- Work in a domain with fast, objective ground-truth feedback and data at a scale beyond typical LLM training
- Well-funded and building a strong, senior research and engineering team
- Location: San Francisco, CA
- Work policy: In-person 5 days/week (relocation supported)
- Compensation: $200K–$400K + competitive early-stage equity
- Visa sponsorship: Open to supporting work authorization for the right candidate
- Employment type: Full-time
Similar Jobs
Explore other opportunities that match your interests