A

ML Infrastructure/Systems Engineer

Aurora Greater Seattle Area
Visa Sponsorship Relocation
Apply
AI Summary

Design and operate scalable ML infrastructure, manage GPU clusters, and ensure real-time multimodal AI stability.

Key Highlights
End-to-end ownership of ML infrastructure and systems
Design and operate long-lived WebRTC connections for stable audio and video
Manage Kubernetes- and Terraform-based GPU clusters for capacity, reliability, and deployment ergonomics
Key Responsibilities
Design the system, ship it, operate it, and improve it when the model, traffic, or hardware profile changes
Own the production path for multimodal inference, with explicit attention to latency, throughput, and cost
Manage Kubernetes- and Terraform-based GPU clusters, including capacity, reliability, and deployment ergonomics
Technical Skills Required
Python Kubernetes Terraform
Benefits & Perks
Salary: $250K–$450K
Location: Seattle, WA
Relocation: available
Visa support: available, including new H-1B applications
Benefits: healthcare, PTO, commuter benefits, office lunch/snacks

Job Description


ML Infrastructure/Systems Engineer


Seattle, WA · On-site · Full-time

$250K–$450K salary



The company


This is an early-stage company building visual conversational AI: a product where users can jump on a video call with an AI and have a real-time, multimodal interaction.


The engineering problem is not a standard app layer. It spans media transport, model serving, GPU scheduling, evaluation, and deployment safety, all under tight latency constraints.


The company was founded in 2024, has 14 employees, and has raised $60M.


The team is intentionally small and built around people who have worked on robotics, graphics, avatar systems, and ML research at Apple, Google, Microsoft, Meta, DJI, Niantic Labs, Bosch Research, and Adobe Research.


The leadership team includes PhDs from MIT, Oxford, and the University of Washington, with publications in top AI conferences and journals.


The team values integrity, transparency, and in-person collaboration. The work depends on tight iteration between research and infrastructure, which is why this role is full-time on-site in Seattle.



The role


This is an ML infrastructure and systems role for someone who wants broad ownership across serving, realtime media, GPU infrastructure, and release automation.


You will work directly with researchers and product engineers to turn model requirements into systems that are measurable, deployable, and cost-controlled.


The expectation is end-to-end ownership: design the system, ship it, operate it, and improve it when the model, traffic, or hardware profile changes.



The technical problem


Real-time multimodal AI creates a systems problem that spans long-lived sessions, live audio/video, inference costs, and model iteration.


A bad choice in one layer shows up as latency, degraded media quality, GPU waste, or slow releases elsewhere.


The hard part is building a platform where serving, WebRTC, GPU scheduling, data pipelines, and release automation can all move independently without breaking the product.



What you'll own


• Serving runtime: own the production path for multimodal inference, with explicit attention to latency, throughput, and cost.

• Realtime media: design and operate long-lived WebRTC connections so audio and video remain stable under real-world network conditions.

• Offline pipelines: build orchestration for batch processing, evaluation, and training workflows using Dagster, Ray, and Airflow where they fit.

• GPU infrastructure: manage Kubernetes- and Terraform-based GPU clusters, including capacity, reliability, and deployment ergonomics.

• Release safety: build CI/CD, model versioning, and evaluation gates that support zero-downtime deployment and fast rollback.

• System debugging: trace failures across infra, media, and inference layers when production behavior diverges from expectations.

• Research interface: translate new model requirements into systems constraints that researchers and product engineers can use immediately.



Who this is for


You are likely a strong fit if you have:


• Shipped production ML infrastructure or distributed systems end to end.

• Owned latency, throughput, or cost improvements and can explain the tradeoffs behind your decisions.

• Debugged systems across network, compute, orchestration, and application layers.

• Experience with GPU-backed serving, inference runtimes, or other performance-sensitive systems.

• Comfort working in Python and at least one systems language such as Rust or Go.

• Built data pipelines or evaluation systems that researchers and engineers actually depended on.

• The ability to make good decisions with incomplete requirements.

• Preference for direct ownership over narrow ticket work.

• Comfort working in person in Seattle with a small team.



Tech stack


• Infrastructure: Kubernetes, Terraform

• Languages: Python, Rust, Go

• Pipelines: Dagster, Ray, Airflow

• Realtime: WebRTC

• Serving: vLLM, Triton Inference Server, TensorRT


Experience across all of these is not required. What matters is whether you can reason about the failure modes of the stack and improve them under production constraints.



Why now


The company is early enough that major infrastructure choices are still being made, but real enough that those choices already matter.


The next bottleneck is not a lack of ideas; it is whether the team can keep a real-time AI product stable while model behavior, traffic, and compute profiles change.


This hire will set defaults for how the system serves models, moves media, schedules GPUs, and ships changes.



This role is not for you if


• You want a narrowly defined role with one subsystem and little cross-functional work.

• You need fully specified tickets before you can start.

• You want remote work.

• You do not want to work on production systems that affect latency, media quality, or compute spend.

• You are not comfortable operating close to research and product change.



Compensation and logistics


• Salary: $250K–$450K

• Location: Seattle, WA

• Work model: full-time on-site, 5 days per week

• Relocation: available

• Visa support: available, including new H-1B applications

• Benefits: healthcare, PTO, commuter benefits, office lunch/snacks



Interview process


Typical process:


• Intro call

• Technical phone screen 1

• Technical phone screen 2

• On-site interview


The process is designed to evaluate systems judgment, performance intuition, and how you work through real infrastructure tradeoffs.



About Aurora


Aurora helps exceptional engineers find the right role at some of the most ambitious startups worldwide.


We work with teams that value high ownership, strong technical standards, and clear scope.


Similar Jobs

Explore other opportunities that match your interests

Manufacturing Engineer

Programming
3d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Blue Origin

Greater Seattle Area

Principal Systems Integrator, TeraWave Constellation

Programming
3d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Blue Origin

Greater Seattle Area

Senior Embedded Software Engineer - Team Lead (National Security Programs)

Programming
3w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Blue Origin

Greater Seattle Area

Subscribe our newsletter

New Things Will Always Update Regularly