We are building a high‑performance AI inference platform from the ground up. The role designs and implements a Rust‑based inference runtime, handling batching, scheduling, routing, and the full serving stack. Requires 2‑10 years of systems engineering experience with deep knowledge of LLM inference internals.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Member of Technical Staff, Inference Systems (On-site)
Location: On-site (5 days/week in-office - Palo Alto, CA)
Salary: $230K-$350K + competitive equity
Sponsorship: open to visa transfers (OPT, H1B) and new sponsorships (H1B, TN)
We're hiring a Member of Technical Staff, Inference Systems at a well-funded, founding-stage team building a high-performance AI inference platform from the ground up. You'll build a new inference runtime in Rust, owning batching, scheduling, request routing, and the full serving stack.
Searching for Development & Programming roles that provide visa sponsorship? Connect with international employers through Development & Programming Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
This is a from-scratch build, not a wrapper around existing tools. The team is architecting the entire runtime with latency, throughput, and cost per token as first-order concerns. Every core architectural decision is still open, and you'll be one of the people making them.
The problem you'd help solve:
Serving LLMs at scale is a systems problem, not a model problem. Throughput and cost per token are decided by scheduling, batching, KV cache management, and how well the runtime uses GPUs across nodes. Most teams inherit these decisions from a general-purpose engine. This team is building the runtime itself, with no legacy constraints, for engineers who want to work on inference internals rather than around them.
You'll build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token, scale serving across multi-GPU and multi-node setups, profile and benchmark the full inference pipeline, and work directly with the founding team on the architecture that defines the platform.
You'll likely be a fit if you have:
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
* 2-10 years of experience as a backend, systems, or distributed systems engineer
* Deep familiarity with inference internals: attention, KV cache, batching, scheduling
* Hands-on time inside engines like vLLM, SGLang, or TensorRT-LLM, or experience building LLM serving infrastructure
* Strong systems programming skills in Rust, C++, or similar, and a genuine willingness to work in Rust day to day
* Experience profiling and optimizing performance-critical systems
* Comfort owning an entire stack rather than a narrow slice
Nice to have:
Interested in opportunities specifically in United State? Discover our dedicated Visa Sponsorship Jobs in United State page featuring roles from top employers in this location.
* Multi-GPU and multi-node serving experience
* CUDA, Triton, or NCCL experience
* Production Rust
* Early-stage startup experience
What you won't find here:
This role won't suit you if you want remote or hybrid work, or if you prefer a defined scope with clear boundaries. The team is small, the pace is high, and you'll be shaping architecture rather than picking up well-specified tickets. You'll be on-site five days a week in Palo Alto.
Similar Jobs
Explore other opportunities that match your interests