C

Member of Technical Staff, Inference Systems

Confidential • United State
Visa Sponsorship
Apply Now
AI Summary

We are building a high‑performance AI inference platform from the ground up. The role designs and implements a Rust‑based inference runtime, handling batching, scheduling, routing, and the full serving stack. Requires 2‑10 years of systems engineering experience with deep knowledge of LLM inference internals.

Key Highlights
Build a new AI inference runtime in Rust from scratch
Own end‑to‑end serving stack including batching, scheduling, and request routing
On‑site role in Palo Alto with competitive salary and equity
Key Responsibilities
Build a new inference runtime in Rust, owning batching, scheduling, request routing, and the full serving stack.
Design KV cache management and prefix caching to reduce latency and cost per token.
Scale serving across multi‑GPU and multi‑node configurations.
Profile and benchmark the complete inference pipeline for performance optimization.
Collaborate directly with the founding team to define the platform architecture.
Technical Skills Required
Rust Distributed Systems LLM Inference
Benefits & Perks
$230K-$350K annual salary
Competitive equity
Visa sponsorship (OPT, H1B, TN)
Nice to Have
Multi‑GPU and multi‑node serving experience
CUDA, Triton, or NCCL expertise
Production‑grade Rust experience
Early‑stage startup background

Job Description


Member of Technical Staff, Inference Systems (On-site)


Location: On-site (5 days/week in-office - Palo Alto, CA)


Salary: $230K-$350K + competitive equity


Sponsorship: open to visa transfers (OPT, H1B) and new sponsorships (H1B, TN)


We're hiring a Member of Technical Staff, Inference Systems at a well-funded, founding-stage team building a high-performance AI inference platform from the ground up. You'll build a new inference runtime in Rust, owning batching, scheduling, request routing, and the full serving stack.


This is a from-scratch build, not a wrapper around existing tools. The team is architecting the entire runtime with latency, throughput, and cost per token as first-order concerns. Every core architectural decision is still open, and you'll be one of the people making them.


The problem you'd help solve:


Serving LLMs at scale is a systems problem, not a model problem. Throughput and cost per token are decided by scheduling, batching, KV cache management, and how well the runtime uses GPUs across nodes. Most teams inherit these decisions from a general-purpose engine. This team is building the runtime itself, with no legacy constraints, for engineers who want to work on inference internals rather than around them.


You'll build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token, scale serving across multi-GPU and multi-node setups, profile and benchmark the full inference pipeline, and work directly with the founding team on the architecture that defines the platform.


You'll likely be a fit if you have:


* 2-10 years of experience as a backend, systems, or distributed systems engineer

* Deep familiarity with inference internals: attention, KV cache, batching, scheduling

* Hands-on time inside engines like vLLM, SGLang, or TensorRT-LLM, or experience building LLM serving infrastructure

* Strong systems programming skills in Rust, C++, or similar, and a genuine willingness to work in Rust day to day

* Experience profiling and optimizing performance-critical systems

* Comfort owning an entire stack rather than a narrow slice


Nice to have:


* Multi-GPU and multi-node serving experience

* CUDA, Triton, or NCCL experience

* Production Rust

* Early-stage startup experience


What you won't find here:


This role won't suit you if you want remote or hybrid work, or if you prefer a defined scope with clear boundaries. The team is small, the pace is high, and you'll be shaping architecture rather than picking up well-specified tickets. You'll be on-site five days a week in Palo Alto.


Similar Jobs

Explore other opportunities that match your interests

Senior Full-Stack/Backend Engineer

Programming
•
7h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Agave

United State

Senior Salesforce Developer

Programming
•
7h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

intellibee inc

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

international recruiting llc

United State

Subscribe our newsletter

New Things Will Always Update Regularly