D

Senior Distributed ML Infrastructure Engineer

david joseph & company United State
Visa Sponsorship Relocation
Apply
AI Summary

Own the distributed training and inference backbone for a foundation model trained from scratch. Design, deploy, and maintain large-scale ML clusters and pipelines at petabyte scale. Requires 2-10 years of hands-on experience with foundation model infrastructure and deep GPU optimization expertise.

Key Highlights
Own distributed training and inference backbone for foundation models
Build petabyte-scale data and training pipelines
Optimize low-level GPU operations across model scales
Key Responsibilities
Design, deploy, and maintain large distributed ML training and inference clusters
Build efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycle
Profile and debug low-level GPU operations to optimize performance
Technical Skills Required
Distributed training frameworks GPU optimization Kubernetes/Docker Python/C++
Benefits & Perks
Competitive early-stage equity
Relocation supported
Visa sponsorship
Nice to Have
Generalist experience spanning the full ML lifecycle
Low-level GPU performance optimization and debugging (CUDA, JAX)

Job Description


San Francisco, CA

  • On-site (5 days/week)
  • Full-time Compensation: $200K–$400K + competitive early-stage equity


About The Company

Our client is a Series A AI research lab building large-scale foundation models for scientific and physical-AI domains. Backed by top-tier investors, they are pursuing a deliberately non-consensus technical thesis and are among the best-funded teams in their space. The founding team comes from self-driving, robotics, and scientific research, and they are scaling their research and engineering org significantly this year.

Founded 2024

  • Small, fast-growing team
  • Industry: AI / foundation models / physical AI


The Role

You would own the distributed training and inference backbone for a foundation model trained from scratch — standing up clusters, building data and training pipelines at petabyte scale, and squeezing performance out of GPUs at a low level across model scales.

What You'll Be Doing

  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Build efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycle
  • Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales
  • Profile and debug low-level GPU operations to optimize performance
  • Track new research and bring fresh ideas into the work


Tech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).

Requirements

  • 2–10 years building large-scale ML infrastructure for core foundation models
  • Hands-on experience building infrastructure for foundation models trained from scratch, rather than fine-tuning existing models
  • A background at a science-focused or physical-AI company (for example self-driving, robotics, or biology)
  • Deep, demonstrable expertise optimizing large-scale training and inference workloads
  • Working proficiency with distributed training frameworks such as FSDP or DeepSpeed
  • A clear pattern of intentional, mission-driven career decisions
  • Able to work on-site 5 days/week in San Francisco (relocation supported)


Nice to Haves

  • Generalist experience spanning the full ML lifecycle
  • Low-level GPU performance optimization and debugging (CUDA, JAX)


Why Join

  • Take a bet on a distinctive, non-consensus approach to building intelligence
  • Join early, with real ownership of the training and inference backbone
  • Work in a domain with fast, objective ground-truth feedback and data at a scale beyond typical LLM training
  • Well-funded and building a strong, senior research and engineering team


Details

  • Location: San Francisco, CA
  • Work policy: In-person 5 days/week (relocation supported)
  • Compensation: $200K–$400K + competitive early-stage equity
  • Visa sponsorship: Open to supporting work authorization for the right candidate
  • Employment type: Full-time

Similar Jobs

Explore other opportunities that match your interests

Senior Staff AI Engineer – Virtual Agent Platform Architecture & Leadership

Machine Learning
2h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

GEICO

United State

Senior Distributed Training Systems Engineer (Foundation Model Pre-Training)

Machine Learning
3h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

reflection

United State

Senior Machine Learning Engineer

Machine Learning
3h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

GEICO

United State

Subscribe our newsletter

New Things Will Always Update Regularly