D

Senior Compute Kernel Engineer

densityai • United State
Visa Sponsorship
Apply Now

Develop and optimize specialized compute kernels for a custom AI accelerator, serving as the critical interface between ML workloads and silicon. Build profiling infrastructure and define kernel programming models to drive performance validation and architectural design decisions. Requires production-grade C/C++, deep CUDA/accelerator programming experience, and strong computer architecture knowledge.

Key Highlights
Write and optimize performance-critical compute kernels for a custom AI accelerator
Develop profiling infrastructure to measure kernel performance against architectural targets
Collaborate with architecture and compiler teams on kernel DSL and MLIR dialect design
Key Responsibilities
Write and optimize compute kernels for a custom AI accelerator, including tensor operations, data movement patterns, and memory hierarchy exploitation
Develop and maintain profiling infrastructure to measure kernel performance against architectural targets
Define and document shuffle patterns for ML kernel primitives across CPU-like control, tensor cores, and CUTLASS-style operations
Drive kernel DSL design decisions regarding thread spawn mechanisms, register passing conventions, and memory management strategies
Enable end-to-end kernel execution on the architectural simulator
Collaborate with the compiler team on the MLIR dialect to validate kernel implementations
Create onboarding documentation and kernel writing guides for the broader team
Technical Skills Required
C/C++ CUDA Computer Architecture
Benefits & Perks
Equity grant
Medical / dental / vision insurance
401(k)
Standard PTO
Nice to Have
RISC-V, x86, or ARM64 ISA experience
MLIR or LLVM compiler infrastructure
HPC or scientific computing background
FPGA or Verilog/SystemVerilog
Familiarity with CUTLASS, Triton, or similar kernel libraries

Job Description

You will write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. This is the critical interface between high-level ML workloads and silicon — your code directly determines how effectively the hardware performs. You'll work closely with the architecture and compiler teams to define the kernel programming model, implement core tensor operations, and drive the performance profiling workflow that validates silicon design decisions.What you'll do
Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Founding Engineer - AI Security

Programming
•
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

RemoteStar

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

cvector

United State

Principal Digital Architect for Physical AI Data Platform

Programming
•
3h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Caterpillar Inc.

United State

Subscribe our newsletter

New Things Will Always Update Regularly