I

On-site LLM Systems Engineer (Inference Systems and Hardware Performance)

Visa Sponsorship Relocation
Apply Now
AI Summary

Build and optimize emerging LLM inference systems, including algorithm and engine prototyping. Develop memory models for expandable vRAM and write efficient GPU kernels for data movement. Contribute to design-space exploration, benchmarking, and guide hardware team product definition requirements.

Key Highlights
Prototyping and optimizing emerging ML/LLM inference systems on real hardware and simulations
Developing memory models for expandable vRAM and implementing efficient GPU kernels for data movement
On-site role in Santa Clara, CA or Boston, MA with responsibility for hardware-informed product definition
Key Responsibilities
Prototype and optimize emerging ML inference systems.
Develop novel memory models for expandable vRAM.
Write efficient GPU kernels for data movement.
Perform design-space exploration, implementation, and benchmarking of inference engines in simulations and on real hardware.
Guide the hardware team on product definition.
Technical Skills Required
LLM inference GPU kernel programming Accelerator programming (CUDA, JAX/Pallas, ROCm)
Benefits & Perks
Comprehensive benefits including health, dental, vision, and life insurance
Relocation assistance and visa sponsorship
401k match and daily lunch stipend
Nice to Have
Advanced computer architectures and performance engineering skills

Job Description


About the role

We’re looking for a motivated LLM Systems Engineer willing to explore new and

unconventional inference systems based on emerging hardware.

This role is part engineering, part research – you’ll be responsible for searching and

prototyping various algorithms suitable for our inference hardware, as well as guiding our

hardware team on the product definition. The ideal candidate has a proven track of record

of pursuing ML systems research, and is very familiar with industry-standard LLM inference

systems.


This role will be performed on-site from one of our offices in Santa Clara, CA or Boston, MA.


Essential Duties and Responsibilities

• Prototype and optimize emerging ML inference systems.

• Develop novel memory models for expandable vRAM.

• Write efficient GPU kernels for data movement.

• Perform design-space exploration, implementation, and benchmarking of inference

engines, both in simulations and on real hardware.


Qualifications

• MS or PhD in computer systems, ideally with a focus on LLM inference and/or

distributed systems.

• Prior experience contributing to the core LLM inference infrastructures (vLLM,

SGLang, TensorRT, etc.).

• Prior experience in accelerator programming (e.g. CUDA, JAX/Pallas, ROCm).

• Advanced computer architectures and performance engineering skills is a big plus.


Compensation & Benefits

• Competitive salary commensurate with experience including base salary, incentive-

based bonus, and early stage equity grant.

• Comprehensive benefits including health, dental, vision, and life insurance.

• Well-equipped, sunny offices in Santa Clara, CA and Boston, MA.

• Relocation assistance and visa sponsorship.

• Perks include a daily lunch stipend, 401k match, and more.• A collaborative, continuous-learning work environment with smart, dedicated

colleagues engaged in developing the next generation of architecture for high-

performance computing.


The Opportunity

• Impact: We are tackling a fundamental challenge at the infrastructure layer:

unlocking greater AI capability while dramatically improving efficiency. The work we

do here compounds across state-of-the-art AI models, systems, and real-world

applications.

• Timing: Joining now means real ownership of the company and meaningful

influence over product direction and execution. You’ll work from first principles,

move quickly from insight to execution, and see your contributions directly reflected

in what we build.

• Culture: You’ll work alongside a group of people who care deeply about rigor, clarity,

and impact. We value thoughtful disagreement, fast learning, and intellectual

fearlessness. This is a place where strong ideas shine, curiosity is encouraged, and

growth is a daily practice.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Agave

United State

Senior Salesforce Developer

Programming
•
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

intellibee inc

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

poseidon aerospace

United State

Subscribe our newsletter

New Things Will Always Update Regularly