A

Member of Technical Staff - Kernel Engineer

Acceler8 Talent • United State
Visa Sponsorship
Apply Now

Design and optimize high-performance GPU and accelerator kernels for a high-throughput AI inference platform. Identify performance bottlenecks in production LLM workloads and implement hardware-aware optimizations. Requires deep expertise in CUDA, C++, and modern GPU architectures to improve model latency and throughput.

Key Highlights
Focus on performance-critical layers of the inference stack including GEMM, attention, and MoE workloads
Optimize end-to-end LLM inference performance across heterogeneous compute fleets
Significant technical ownership of the hardware-performance roadmap and runtime execution layer
Key Responsibilities
Design, implement, and optimize high-performance GPU kernels for AI operations such as GEMM, attention, and MoE workloads.
Develop and optimize kernels using CUDA, Triton, C++, PTX, and CUTLASS.
Profile production inference workloads to identify bottlenecks across compute, memory bandwidth, and synchronization.
Optimize GPU execution through memory coalescing, shared-memory utilization, tiling, and Tensor Core utilization.
Implement and evaluate lower-precision execution and quantization strategies including FP8, FP4, and INT8.
Integrate optimized kernels into inference frameworks like PyTorch, vLLM, SGLang, or TensorRT-LLM.
Build reliable benchmarks and correctness tests to ensure kernel improvements are production-ready.
Optimize workloads across multi-GPU and distributed environments.
Collaborate with hardware vendors to translate new accelerator architectures into production performance improvements.
Technical Skills Required
C++ CUDA GPU Programming
Benefits & Perks
Equity / stock options
Medical, dental, and vision coverage
Unlimited PTO
H-1B and other work visa sponsorship available
Nice to Have
Experience with PTX/SASS, CUTLASS, CuTe, CUB, or Thrust
Experience with inference frameworks such as vLLM, SGLang, TensorRT-LLM, or FlashInfer
Experience implementing FlashAttention or fused attention kernels
Experience with ROCm / HIP and AMD GPU architectures
Experience with NCCL and collective communication in multi-GPU systems
Knowledge of ML compilers such as torch.compile, XLA, MLIR, or TVM
Experience with tensor, pipeline, or expert parallelism
Contributions to open-source projects in GPU kernels or ML systems

Job Description

Member of Technical Staff - Kernel Engineer

Location: Santa Clara, CA

About the Company

This company builds and operates a high-performance AI inference platform that helps enterprises run large language models faster, more efficiently, and at scale.

Its infrastructure sits at the foundation of the AI application stack, with a focus on improving the speed, reliability, utilization, and economics of production model serving.

About the Role

A fast-growing AI infrastructure company operates an enterprise inference platform providing high-throughput, low-latency access to large language models across a distributed, heterogeneous compute fleet.

Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

systems technology group, inc....

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

remotehunter

United State

Senior Backend Engineer, Observability Platform

Programming
•
41m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United State

Subscribe our newsletter

New Things Will Always Update Regularly