N

Senior Kernel Engineer (GPU & AI Inference Optimization)

neurospark ai • United State
Visa Sponsorship
Apply Now
AI Summary

Design and optimize high-performance GPU kernels for NeuroSpark’s AI inference platform, improving latency, throughput, and hardware efficiency. Work at the intersection of model architecture, GPU execution, and distributed systems to push performance boundaries across emerging hardware. Hands-on role with full technical ownership, collaborating with cross-functional teams to scale LLM inference at enterprise scale.

Key Highlights
Optimize performance-critical GPU kernels for AI workloads (GEMM, attention, quantization, MoE, etc.) using CUDA, Triton, and low-level GPU techniques
Profile and diagnose bottlenecks in production inference workloads, implementing hardware-aware optimizations for latency, throughput, and memory efficiency
Collaborate with hardware vendors and engineering teams to extend NeuroSpark’s inference stack across heterogeneous accelerators (NVIDIA, AMD, and beyond)
Key Responsibilities
Design, implement, and optimize high-performance GPU kernels for AI operations (GEMM, attention, normalization, quantization, KV-cache, and MoE workloads)
Profile production inference workloads to identify bottlenecks in compute, memory bandwidth, kernel launch overhead, and synchronization, then optimize using techniques like memory coalescing, shared-memory utilization, tiling, and warp-level programming
Develop and integrate optimized kernels into inference frameworks (PyTorch, Triton, vLLM, SGLang, TensorRT-LLM) while ensuring numerical correctness and production readiness
Extend NeuroSpark’s inference stack across heterogeneous hardware (NVIDIA/AMD GPUs) and emerging AI accelerators, collaborating with hardware vendors and engineering teams
Build and maintain benchmarks, correctness tests, and performance regression tests to validate kernel improvements
Optimize workloads for multi-GPU and distributed environments, improving concurrency, latency, and hardware utilization
Technical Skills Required
CUDA and GPU Programming C++ Performance Optimization (Profiling, Benchmarking, Low-Level Hardware Analysis)
Benefits & Perks
Equity / stock options
Medical, dental, and vision coverage
Unlimited PTO
Nice to Have
PTX/SASS, CUTLASS, or CuTe for low-level GPU programming
Experience with inference frameworks (vLLM, SGLang, TensorRT-LLM, FlashInfer)
FlashAttention or fused attention kernel implementation
ROCm/HIP for AMD GPU optimization
ML compiler/runtime systems (torch.compile, XLA, MLIR, TVM)
Distributed training/inference (tensor/pipeline/expert parallelism)
Open-source contributions in GPU kernels, ML systems, or inference engines

Job Description


About NeuroSpark

NeuroSpark builds and operates a high-performance AI inference platform that helps enterprises run large language models faster, cheaper, and at scale. Inference infrastructure is the foundation the entire AI application layer runs on — every AI product ultimately depends on how fast, how reliably, and how affordably models can serve their users. Our vision is to make that layer so efficient that compute is never the reason a good AI product fails.

About The Role

NeuroSpark Inc operates an enterprise AI inference platform providing high-throughput, low-latency access to large language models across a distributed, heterogeneous compute fleet.

As a Member of Technical Staff, Kernel Engineer, you will work at the lowest performance-critical layers of NeuroSpark's inference stack, designing and optimizing GPU and accelerator kernels that directly determine model latency, throughput, memory efficiency, and hardware utilization.

You will work across the boundary between model architecture, GPU execution, inference runtimes, and distributed serving systems. The role involves identifying performance bottlenecks in real production workloads, developing hardware-aware optimizations, and integrating those improvements into the systems that serve models at scale.

This is a hands-on individual-contributor role with significant technical ownership. You will work closely with engineers across inference, distributed systems, and infrastructure to push model performance across current and emerging accelerator platforms.

Responsibilities

  • Design, implement, and optimize high-performance GPU kernels for performance-critical AI operations, including GEMM, attention, normalization, quantization, KV-cache operations, and Mixture-of-Experts (MoE) workloads.
  • Develop and optimize kernels using technologies such as CUDA, Triton, C++, PTX, CUTLASS, and related GPU programming frameworks.
  • Profile production inference workloads to identify bottlenecks across compute, memory bandwidth, memory hierarchy, kernel launch overhead, synchronization, and data movement.
  • Optimize GPU execution through techniques including memory coalescing, shared-memory utilization, tiling, warp-level programming, Tensor Core utilization, operator fusion, latency hiding, and compute/communication overlap.
  • Improve end-to-end LLM inference performance across latency, throughput, memory utilization, concurrency, and hardware efficiency, rather than optimizing kernels in isolation.
  • Develop and optimize kernels for modern model architectures, including Transformer-based LLMs, attention variants, MoE models, and emerging model architectures.
  • Implement and evaluate lower-precision execution and quantization strategies, including FP16, BF16, FP8, FP4, INT8, and other hardware-supported formats.
  • Integrate optimized kernels and operators into inference frameworks and internal runtimes built around technologies such as PyTorch, Triton, vLLM, SGLang, TensorRT-LLM, or equivalent systems.
  • Use profiling and performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, roofline analysis, and internal benchmarking infrastructure to diagnose and resolve performance regressions.
  • Build reliable benchmarks, correctness tests, and performance regression tests to ensure kernel improvements remain numerically correct and production-ready.
  • Optimize workloads across multi-GPU and distributed environments, working with the broader infrastructure team on communication, parallelism, and compute efficiency.
  • Help extend NeuroSpark's inference stack across heterogeneous hardware, including NVIDIA GPUs, AMD GPUs, and other current and emerging AI accelerators.
  • Work closely with hardware vendors, inference engineers, and distributed-systems engineers to evaluate new accelerator architectures and translate hardware capabilities into production performance improvements.
  • Contribute to architectural decisions affecting the Company's inference runtime, model execution layer, and hardware-performance roadmap.

Qualifications

  • Strong experience in GPU programming, kernel development, high-performance computing, ML systems, or performance engineering.
  • Proficiency in C++ and hands-on experience with CUDA, Triton, or another accelerator programming model.
  • Strong understanding of modern GPU architecture, including:
    • GPU memory hierarchy
    • Threads, warps, blocks, and grids
    • Shared memory and register usage
    • Memory bandwidth and access patterns
    • Tensor Cores
    • Synchronization and parallel execution
    • Occupancy and instruction-level parallelism
  • Experience profiling and optimizing GPU workloads for latency, throughput, memory usage, and hardware utilization.
  • Experience with performance-critical machine-learning operations such as attention, GEMM, quantization, KV cache, MoE, or other Transformer operators.
  • Familiarity with PyTorch and modern ML inference or training execution stacks.
  • Strong understanding of numerical correctness, floating-point behavior, and mixed-precision computation.
  • Ability to reason from first principles about performance bottlenecks across hardware and software layers.
  • Strong debugging skills and the ability to take performance work from profiling and hypothesis through implementation, benchmarking, and production deployment.
  • Strong written and verbal communication skills and the ability to work effectively in a highly collaborative engineering environment.
Nice to Have

  • Experience with PTX/SASS, CUTLASS, CuTe, CUB, Thrust, or other low-level GPU libraries and programming abstractions.
  • Experience with inference frameworks such as vLLM, SGLang, TensorRT-LLM, FlashInfer, or similar systems.
  • Experience implementing or optimizing FlashAttention or other fused attention kernels.
  • Experience with ROCm / HIP and AMD GPU architectures.
  • Experience optimizing workloads across multi-GPU or multi-node systems, including NCCL and collective communication.
  • Knowledge of ML compiler and runtime systems such as torch.compile, XLA, MLIR, TVM, or related compiler stacks.
  • Experience with distributed training and inference, tensor parallelism, pipeline parallelism, or expert parallelism.
  • Experience optimizing workloads on multiple accelerator architectures or developing hardware-portable kernels.
  • Contributions to open-source projects in GPU kernels, ML systems, inference engines, compilers, or high-performance computing.
  • Experience bringing new model architectures or accelerator platforms into production.

Compensation & Benefits

The expected base salary range for this position is:

$170,000 – $350,000 USD per year

Actual compensation will depend on experience, technical depth, level, and role scope.

This Position Also Includes

  • Equity / stock options
  • Medical, dental, and vision coverage
  • Unlimited PTO
  • Opportunities for significant technical ownership and impact
  • H-1B and other work visa sponsorship available

Similar Jobs

Explore other opportunities that match your interests

Staff Engineer - CI/CD Platform Modernization

Programming
•
9m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State

Senior Front-End Software Engineer

Programming
•
11m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State

Senior Manager, Software Engineering

Programming
•
12m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State

Subscribe our newsletter

New Things Will Always Update Regularly