T

Senior Systems Engineer - GPU Virtualization

thunder compute โ€ข United State
Visa Sponsorship Relocation
Apply
AI Summary

Build the core C++ systems for GPU virtualization, focusing on low-latency performance optimization, distributed systems debugging, and production reliability. Design and implement a systems layer that abstracts GPUs across TCP networking to improve cluster utilization. Requires exceptional low-level engineering talent with strong work ethic and attention to detail.

Key Highlights
Develop invisible GPU virtualization layer compatible with all workloads
Optimize remote CUDA operations for low latency and high performance
Debug complex distributed systems without clear reproduction steps
Own projects from prototype through 100% production deployment
Work directly with founders on category-defining infrastructure
Key Responsibilities
Profile and reduce latency across remote CUDA operations
Build high-performance networking and data-transfer paths
Debug failures across customer processes, userspace runtime, network, and remote GPU servers
Design systems for GPU allocation, scheduling, failure recovery, and observability
Research and productionize new GPU virtualization and oversubscription techniques
Expand compatibility across CUDA applications, frameworks, and GPU architectures
Technical Skills Required
C++ CUDA Linux systems programming Distributed systems
Benefits & Perks
Competitive salary
Meaningful equity
Daily lunch, snacks, and coffee
Team dinners and events
401(k)
Health, dental, and vision insurance
Nice to Have
Experience with CUDA, GPU systems, compilers, runtime interception, dynamic linking, high-performance networking, or distributed computing
Experience at a trading firm such as Citadel Securities or Jane Street
Experience at a hardware or AI infrastructure company such as NVIDIA or SambaNova
Experience in a systems research group
Strong computer science fundamentals demonstrated through academic work, competitive programming, open-source contributions, or exceptional professional experience

Job Description


Company

Thunder Compute is building the VMware for GPUs. We have raised over $17M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic.

Deployed GPU fleets are currently only 5โ€“20% utilized. Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer.

We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer.

Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center.

This enables something like โ€œCeph for GPUsโ€: GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.

Role

Your work will focus on building the core C++ systems behind our virtualization layer. This includes low-latency performance optimization, distributed systems debugging, production reliability, and research into new techniques for improving GPU utilization.

You will take ownership of complex systems from early experimentation through production deployment. Example projects may include:

  • Profiling and reducing latency across remote CUDA operations
  • Building high-performance networking and data-transfer paths
  • Debugging failures across customer processes, our userspace runtime, the network, and remote GPU servers
  • Improving support for process forking, signals, multithreading, dynamic linking, and unusual application behavior
  • Designing systems for GPU allocation, scheduling, failure recovery, and observability
  • Researching and productionizing new GPU virtualization and oversubscription techniques
  • Expanding compatibility across CUDA applications, frameworks, and GPU architectures

You will spend your days bouncing between the weeds of complex, performance-critical systems that are live in production. One week, you may be tracing a synchronization bug across a distributed CUDA workload; the next, you may be redesigning a hot data path to remove microseconds of overhead.

This work is not easy. It blends the hardest parts of systems research and production engineering.

We look for exceptional low-level engineering talent, strong work ethic, and extreme attention to detail. We must move quickly while shipping high-quality, reliable systems code.

Core Technical Skills

  • Exceptional modern C++ ability, including memory management, concurrency, performance optimization, and systems-level abstraction design
  • Deep understanding of operating systems, low-level networking, compilers, distributed systems, or computer architecture
  • Experience building and operating performance-critical C++ systems in production
  • Strong Linux systems programming and debugging ability
  • Ability to reason through unfamiliar systems across multiple layers of the stack

Must Haves

  • Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines
  • Attention to detail and the ability to deliver production-ready, thoroughly tested code without significant oversight
  • Strong ownership over correctness, reliability, performance, and operational outcomes
  • Ability to debug ambiguous problems without a clear reproduction, existing playbook, or obvious owner
  • Willingness to work directly with customers and investigate difficult production failures

Preferred

  • Experience with CUDA, GPU systems, compilers, runtime interception, dynamic linking, high-performance networking, or distributed computing
  • Experience at a trading firm such as Citadel Securities or Jane Street; a hardware or AI infrastructure company such as NVIDIA or SambaNova; a systems research group; or a similarly demanding engineering environment
  • Strong computer science fundamentals demonstrated through academic work, systems research, competitive programming, open-source contributions, or exceptional professional experience
  • Experience taking new systems research from a paper or prototype into a reliable production system

Why Join

You will join early enough to meaningfully shape the architecture, engineering standards, and technical direction of the company.

You will work directly with the founders on a category-defining systems problem, with a short path between writing code and seeing it run in production. The systems you build will form the foundation of a new infrastructure layer for GPU computing.

Logistics

  • You will report to co-founder and CTO Brian Model, formerly a Quantitative Developer at Citadel Securities
  • This role is full-time and in person, five days per week, at our office in downtown San Francisco
  • Relocation support and visa sponsorship are available

Benefits

  • Competitive salary and meaningful equity
  • Daily lunch, snacks, and coffee
  • Team dinners and events
  • 401(k)
  • Health, dental, and vision insurance

Compensation Range: $200K - $300K


Similar Jobs

Explore other opportunities that match your interests

Controls Systems Engineer

Programming
โ€ข
54m ago
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Not Applicable

Actalent

United State

Senior Accounting Manager - Autonomous Trucking

Programming
โ€ข
1h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

โ€ขโ€ขโ€ขโ€ขโ€ขโ€ข โ€ขโ€ขโ€ขโ€ขโ€ขโ€ข โ€ขโ€ขโ€ขโ€ขโ€ขโ€ข
Job Type โ€ขโ€ขโ€ขโ€ขโ€ขโ€ข
Experience Level โ€ขโ€ขโ€ขโ€ขโ€ขโ€ข

kodiak

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State

Subscribe our newsletter

New Things Will Always Update Regularly