I

Senior ML Infrastructure Engineer

inventure United State
Visa Sponsorship
Apply
AI Summary

Architect and scale distributed training infrastructure for large-scale multimodal robotic AI models. Build high-throughput data pipelines and low-latency inference systems to support real-time robot control. Requires 5+ years of experience in ML infrastructure, PyTorch, and distributed training frameworks.

Key Highlights
Own end-to-end training infrastructure for a robotics company with $140M+ funding.
Work on-site in Redwood City, CA, with visa sponsorship available.
Competitive base salary of $220K–$350K plus equity.
Key Responsibilities
Architect and scale distributed training across large GPU clusters using sharding, activation checkpointing, and memory optimization.
Build researcher-friendly tooling and job scheduling on Kubernetes and SLURM with automated retries and failure recovery.
Design high-throughput data pipelines to ingest and transform terabytes of multimodal robot data.
Build low-latency inference pipelines for real-time robot control using quantization, distillation, and model compilation.
Profile deep into the stack to optimize GPU utilization, I/O bottlenecks, and memory fragmentation.
Technical Skills Required
PyTorch Distributed Training Kubernetes
Benefits & Perks
Competitive equity
Visa sponsorship (OPT and H1B transfers)
Nice to Have
Robotics experience
Experience building multimodal systems for video, audio, or other rich media models

Job Description


ML Infrastructure Engineer | Redwood City, CA (Onsite) | $220K–$350K base


I've partnered with a robotics company building general-purpose robotic arms powered by their own embodied AI foundation model, and unlike most of the field they are already out of the lab: their robots are deployed at real customer sites doing commercial-grade work, starting in the hospitality and restaurant world. They have raised over $140M, including a $120M Series A, from a top-tier investor group spanning NVIDIA's venture arm, Samsung NEXT, Salesforce Ventures, First Round Capital and CRV. The founders are repeat founders who previously built and sold their last company to a major consumer marketplace, and the team is stacked with people from Google DeepMind, Meta and Cruise. The near-term focus is shipping the next generation of robots with whole-body control; the long-term vision is nothing less than physical AGI.


This is a rare opportunity to own training infrastructure end to end and become the connective tissue between researchers and compute. There is no infra layer above you: you'll turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models, and drive the architecture yourself. It is a genuinely high-ownership seat where every optimisation you ship shortens the path from model to deployed robot. If you want your systems work to move real machines in the physical world rather than sit behind an API, this is it.


The Role:


As an ML Infrastructure Engineer, you will:

  • Architect and scale distributed training across large GPU clusters, implementing sharding, activation checkpointing and memory optimisation such as ZeRO and FSDP for multimodal models.
  • Build researcher-friendly tooling and job scheduling on Kubernetes and SLURM, with fast iteration, automated retries and seamless failure recovery.
  • Design high-throughput data pipelines that ingest and transform terabytes of multimodal robot data, from video to proprioception to 3D signals, so the dataloaders never starve the GPUs.
  • Build low-latency inference pipelines for real-time robot control, applying quantisation, distillation and model compilation with tools like TensorRT and Triton.
  • Profile deep into the stack, chasing GPU utilisation, I/O bottlenecks and memory fragmentation to squeeze maximum performance out of an expanding compute fleet.


About You:


  • You have at least 5 years of infrastructure engineering experience, ideally 7 or more, and you've built and maintained ML or data infrastructure on a team with a high talent bar.
  • You have deep, hands-on PyTorch experience and real command of distributed training frameworks such as DeepSpeed or Accelerate.
  • You are an expert on GPU bottlenecks, model serving optimisation and monitoring, and you enjoy the systems-profiling side of the work.
  • You are genuinely excited about robotics and physical AI, and you thrive in a fast-paced, high-ownership startup environment.
  • One of these backgrounds fits you: an ML or HPC infrastructure engineer from a top AI lab or research team, an early or founding infrastructure hire at a startup, or an infrastructure engineer coming from robotics or autonomous vehicles.
  • Nice to have: robotics experience, or having built multimodal systems for video, audio or other rich media models.
  • You are based in or willing to relocate to the Bay Area and happy to work onsite five days a week in Redwood City.
  • Visa sponsorship is available, including OPT and H1B transfers.


The role comes with competitive equity on top of base. The team is moving quickly and hiring several engineers across ML and data infrastructure. If you're interested in owning the training engine behind robots already working in the real world, apply now or send your CV directly to [email protected].


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

Palo Alto Networks

United State

Senior Staff AI Engineer - Enterprise AI Architecture & Platforms

Devops
1d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

American Express

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

clera

United State

Subscribe our newsletter

New Things Will Always Update Regularly