R

Senior Inference Systems Engineer (AI Model Serving & Performance Optimization)

river ai United State
Visa Sponsorship Relocation
Apply
AI Summary

Build high-performance inference engines for River AI’s proprietary AI models, optimizing GPU compute, latency, and reliability. Own serving runtime, distributed execution, and RL sampling workloads while collaborating with researchers and kernel engineers. Focus on dense/mixture-of-experts models, batching, caching, and fault tolerance under production traffic.

Key Highlights
Own end-to-end inference serving stack, including scheduling, batching, KV-cache, and multi-GPU execution
Optimize performance for dense/mixture-of-experts models with strict attention to latency, throughput, and cost
Collaborate with GPU kernel engineers, researchers, and infrastructure teams to scale model serving
Key Responsibilities
Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters
Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness
Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency
Build reliable streaming, cancellation, and recovery mechanisms under failures and overload conditions
Profile bottlenecks and validate improvements through reproducible performance and correctness tests
Technical Skills Required
GPU Memory Management Distributed Systems Transformer Inference
Benefits & Perks
Comprehensive health, dental, and vision insurance
Unlimited paid time off (PTO)
Relocation assistance as needed
Nice to Have
Experience with SGLang, vLLM, or TensorRT-LLM frameworks
Familiarity with expert parallelism, GPU collectives, and quantized inference
Experience with multi-adapter serving, dynamic checkpoint loading, or RL sampling

Job Description


At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.


Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.


About The Role

We are looking for exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.


You will take ownership of the serving runtime, from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Your work will support both customer-facing inference and the sampling workloads that power reinforcement learning.


Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will bring new models into production and improve their performance across realistic workloads. You will measure success through latency, throughput, reliability, and cost, with careful attention to numerical correctness and model behavior.


What You’ll Do

  • Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters.
  • Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness.
  • Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency.
  • Improve RL sampling throughput while keeping samples and log probabilities tied to the correct model version.
  • Build reliable streaming, cancellation, and recovery under failures and overload.
  • Profile bottlenecks and validate improvements through reproducible performance and correctness tests.


Skills & Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong understanding of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • A collaborative mindset and strong ownership of engineering outcomes.


Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Experience extending SGLang, vLLM, TensorRT-LLM, or similar frameworks.
  • Work on advanced serving techniques, such as speculative decoding or disaggregated prefill and decode.
  • Familiarity with expert parallelism, GPU collectives, and quantized inference.
  • Experience with multi-adapter serving, dynamic checkpoint loading, or RL sampling.
  • Familiarity with CUDA graphs, custom kernels, and NVIDIA profiling tools.
  • Experience operating model-serving systems under production traffic.


Logistics & Benefits

  • Location: Palo Alto, California.
  • Compensation: $200,000–$420,000 USD annual base pay, depending on experience and skills.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and support the process for the right candidate.

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

pluralis research

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

redirecruit, llc

United State

Protective Intelligence Analyst (Global Safety, Intelligence, and Security)

Programming
1h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

anthropic

United State

Subscribe our newsletter

New Things Will Always Update Regularly