N

Senior Kernel Engineer (GPU/Accelerator Performance Optimization)

neurospark ai • United State
Remote Visa Sponsorship
Apply Now

Design and optimize high-performance GPU kernels for NeuroSpark’s AI inference platform, improving latency, throughput, and memory efficiency. Work at the intersection of model architecture, GPU execution, and distributed systems to push performance boundaries across emerging hardware. Hands-on role with full technical ownership and cross-team collaboration.

Key Highlights
Hands-on kernel engineering for GPU/accelerator performance optimization in AI inference
Full ownership of performance-critical layers (GEMM, attention, quantization, MoE, etc.)
Cross-disciplinary collaboration with inference, distributed systems, and hardware teams
Key Responsibilities
Design, implement, and optimize high-performance GPU kernels for performance-critical AI operations (GEMM, attention, normalization, quantization, KV-cache, MoE).
Profile production inference workloads to identify bottlenecks and develop hardware-aware optimizations across compute, memory, and synchronization layers.
Improve end-to-end LLM inference performance by optimizing kernels for latency, throughput, memory utilization, and hardware efficiency.
Develop and integrate optimized kernels into inference frameworks (PyTorch, Triton, vLLM, SGLang, TensorRT-LLM).
Collaborate with hardware vendors and engineering teams to evaluate new accelerator architectures and translate capabilities into production performance improvements.
Technical Skills Required
CUDA C++ GPU Architecture Optimization
Benefits & Perks
Equity / stock options
Medical, dental, and vision coverage
Unlimited PTO
Nice to Have
Experience with PTX/SASS, CUTLASS, or CuTe
Optimization of FlashAttention or fused attention kernels
ROCm/HIP and AMD GPU architectures
ML compiler/runtime systems (torch.compile, MLIR, TVM)

Job Description

About the Role

NeuroSpark Inc operates an enterprise AI inference platform providing high-throughput, low-latency access to large language models across a distributed, heterogeneous compute fleet.

As a Member of Technical Staff, Kernel Engineer, you will work at the lowest performance-critical layers of NeuroSpark's inference stack, designing and optimizing GPU and accelerator kernels that directly determine model latency, throughput, memory efficiency, and hardware utilization.

Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Senior Front-End Platform Engineer

Programming
•
29m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Verkada

United State

Founding Member of Technical Staff (MTS)

Programming
•
47m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

glen

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Progressive Insurance

United State

Subscribe our newsletter

New Things Will Always Update Regularly