S

Founding GPU Infrastructure Engineer (HPC & AI Systems)

sentiro partners San Francisco Bay Area
Relocation
Apply
AI Summary

Build and scale a GPU neocloud infrastructure from bare metal to production, designing high-performance HPC and AI systems. Own GPU cluster commissioning, Kubernetes/Slurm orchestration, distributed training, and observability for enterprise-grade AI workloads. Founding role with meaningful equity, competitive compensation, and relocation support in San Francisco.

Key Highlights
Founding role in a seed-stage GPU neocloud company with VC backing
End-to-end responsibility for GPU infrastructure from hardware deployment to customer-ready production
Equity, competitive salary, benefits, and relocation support in San Francisco
Key Responsibilities
Commission and burn-in NVIDIA GPU clusters for production deployment
Design and operate Kubernetes and Slurm infrastructure for high-performance workloads
Develop orchestration layers for distributed GPU training, MPI, and batch workloads
Architect distributed object storage, NVMe systems, and high-bandwidth networking (InfiniBand, RDMA)
Build observability, telemetry, and reliability tooling for cluster lifecycle management
Establish low-latency, high-throughput inference infrastructure for AI workloads
Define customer-facing interfaces and operational patterns for scalable deployments
Technical Skills Required
Kubernetes Slurm Distributed Systems
Benefits & Perks
Meaningful equity
Competitive San Francisco salary
Performance bonus
Relocation support
Nice to Have
Experience with GPU power capping or energy telemetry
Background in AI labs, GPU neoclouds, or quantitative trading firms
Hyperscale infrastructure or serious HPC environments

Job Description


FOUNDING GPU INFRASTRUCTURE ENGINEER [HPC & AI SYSTEMS] - MEANINGFUL EQUITY

REPEAT FOUNDERS WITH SERIOUSLY IMPRESSIVE EXITS.


Build a GPU neocloud from bare metal to inference  

San Francisco | On-site (5 days with some flex) | Relocation support


Most infrastructure roles ask you to keep somebody else’s platform alive.

This one asks you to build the platform.


Sentiro Partners has been retained to find the Founding GPU Infrastructure Engineer for my client, a seed-stage GPU neocloud operating in stealth in San Francisco.


The company is building managed GPU infrastructure and high-performance inference for demanding AI workloads. It has secured backing from well-known VCs and is pursuing deployments at hundreds-of-GPUs scale.


THE FOUNDERS HAVE DONE THIS BEFORE


One founder built an AI infrastructure company operating at the HW layer, bringing low-latency AI computation to resource-constrained devices. Acquired.


The other founded & scaled an enterprise technology company to 400+ customers.


Their previous companies were backed by major VCs. Ivy league & top tier finance backgrounds. Seriously down to earth. Seriously hard workers. They are in the SoMa office right now, as you read this. Now they’re building again.


WHAT SUCCESS LOOKS LIKE


Success means taking new GPU capacity from delivered racks to reliable, customer-ready infrastructure and building the systems and team required to repeat that process at increasing scale.


You will be the company’s founding infrastructure engineer and one of its first technical hires.


When the racks arrive, you will help turn them into a production platform:


- Commission and burn in NVIDIA GPU clusters.

- Build and operate Kubernetes and Slurm infrastructure.

- Develop the orchestration layer above the underlying clusters.

- Enable distributed training, MPI and batch workloads.

- Architect distributed object storage and high-performance NVMe systems.

- Build around InfiniBand, RDMA and other high-bandwidth networking.

- Own telemetry, observability, reliability and cluster lifecycle tooling.

- Help develop low-latency, high-throughput inference infrastructure.

- Create the interfaces through which customers access and operate the platform.

- Establish the patterns that make each deployment faster and more reliable than the last.


THIS IS GPU INFRASTRUCTURE BUILT WITH HPC DISCIPLINE


You should understand what happens between a rack arriving at a facility and a customer successfully running a distributed GPU workload.


You don’t need to be a data-centre real-estate or cooling specialist. You do need enough hardware and systems depth to reason about GPU topology, storage, networking, workload behaviour and production reliability as one connected system.


WE SHOULD TALK IF YOU HAVE


- c. 3-7 years in GPU infrastructure, HPC, AI systems or distributed systems (we're not limited by exp. but you need to be able to go deep).

- Meaningful production experience with both Kubernetes and Slurm (either tbh)

- Built infrastructure rather than only deployed platforms created by other teams.

- Strong knowledge of distributed object storage, NVMe and high-performance networking.

- Experience with distributed training, MPI, InfiniBand, RDMA or similar technologies.

- Worked across hardware, systems software and production operations.

- The ability to move quickly and make sound decisions with incomplete information.

- The appetite to own outcomes rather than defend a narrow area of responsibility.


Particularly relevant backgrounds include AI labs, GPU neoclouds, quantitative trading firms, hyperscale infrastructure teams and serious HPC environments.


Experience with GPU power capping, power-aware scheduling, energy telemetry or grid-flexible compute would be unusually valuable, but it isn’t essential.


WHY NOW?


Because the upside is real.


You will work directly with the founders who have already raised capital, scaled teams and delivered successful exits. You will influence the architecture before the boundaries harden, help deliver the company’s first major GPU clusters and have the opportunity to build the infrastructure team around you.


This is a founding role with a competitive San Francisco salary, performance bonus, benefits and meaningful equity.


The team is building in person in San Francisco. Relocation support is available for exceptional candidates ready to move.


If your ideal role begins where the racks arrive—and ends with customers running serious AI workload... contact me directly.


Adrian Clarke  

Founder, MD, Executive Search Partner Sentiro Partners  


Frontier AI Search & Advisory

(https://sentiropartners.com)


ABOUT SENTIRO PARTNERS

  • Sentiro Partners is a global executive search firm specialising in frontier AI, deep technology and high-performance technical leadership.
  • We partner with ambitious founders, AI labs and technology companies to identify the engineers, researchers and leaders building the next generation of intelligent infrastructure.
  • Headquartered in Dublin, Sentiro Partners conducts searches across North America, Europe and Asia-Pacific.

Similar Jobs

Explore other opportunities that match your interests

Software Engineer - Ads Team

Programming
11h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Discord

San Francisco Bay Area
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

oneclick smart resume

San Francisco Bay Area

Forward Deployed Team Lead (Enterprise AI Applications)

Programming
1w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Fractal

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly