H

Senior Network Solutions Architect - GPU Compute Infrastructure

Hamilton Barnes 🌳 β€’ San Francisco Bay Area
Relocation
Apply
AI Summary

Design, deploy, and manage GPU compute clusters. Collaborate with founders, lead technical onboarding, and troubleshoot customer issues. Hands-on role in San Francisco.

Key Highlights
Design and deploy GPU compute clusters end-to-end
Collaborate directly with founders and customers
Hands-on troubleshooting and real-time build decisions
Key Responsibilities
Design compute, storage, and networking topology
Specify node configurations and redundancy plans
Make real-time build decisions during deployment
Diagnose and fix network bottlenecks under load
Lead technical onboarding and workload validation for customers
Act as direct escalation point for customer cluster underperformance
Technical Skills Required
GPU cluster design and deployment InfiniBand and/or RoCEv2 CUDA and ROCm ecosystems
Benefits & Perks
Founding-level ownership and visibility
Direct access to founders
Onsite role in San Francisco

Job Description


Job Title: Senior Network Solutions Architect


Role:


Join a Stealth Neocloud building GPU compute infrastructure from the ground up in San Francisco. As one of the founding team, the systems you design and build this year are the systems the company runs on. You'll have direct access to and collaboration with the founders, high ownership, high visibility, and no bureaucracy standing between you and the rack.


As a Senior Network Solutions Architect, you will own the design, deployment, and performance of the company's GPU compute clusters end to end, from the network fabric up through the software stack; combined with a genuinely customer-facing role. This is emphatically not a management position. You will be the person racking, cabling, configuring, benchmarking, and debugging the infrastructure yourself, not delegating it, while also leading technical onboarding and acting as the direct escalation point for customers (Obviously you wouldn't be doing physical work every day, but it's a start-up...).


You'll operate as the technical authority on cluster architecture, fabric engineering, and multi-vendor GPU enablement, reporting directly to the founders and making real-time build decisions on-site during deployments. As the team grows, you may build out a small team under you, but the expectation is that you stay hands-on and technical rather than shifting into a managerial role.


Responsibilities:


Design compute, storage, and networking topology for new cluster deployments

Specify node configurations, redundancy, and scaling plans

Make real-time build decisions on-site during deployment, not just on paper

Bridge datacenter specifications to cluster deployment, including power whips and rack fit-out

Design and personally implement GPU interconnect fabric using InfiniBand and/or RoCEv2

Plan and validate bandwidth, topology, and east-west throughput at scale

Diagnose and fix network bottlenecks under load, hands-on rather than in theory

Get workloads running well across both CUDA and ROCm stacks

Own driver and firmware compatibility, NCCL/RCCL tuning, and performance benchmarking

Troubleshoot low-level issues directly across drivers, firmware, and fabric managers

Lead technical onboarding and workload validation for new customers

Act as the direct escalation point when a customer's cluster underperforms, diagnosing it yourself

Translate customer workload requirements into concrete infrastructure and configuration decisions

Document runbooks and SOPs that reflect how the infrastructure actually gets built and fixed


Skills/Must have:


Direct, hands-on experience standing up GPU clusters at production scale, with specific clusters you built rather than systems you oversaw.

Real InfiniBand or RoCEv2 design and implementation experience.

Comfortable working across both CUDA and ROCm ecosystems, or clearly demonstrated ability to ramp fast on a new GPU vendor stack.

Direct customer-facing technical experience, comfortable being the person a customer talks to when something's wrong.

Genuine preference for staying hands-on over moving into pure management.

Based in or willing to relocate to San Francisco (onsite).


Benefits:


Founding-level ownership and visibility.

Direct access to and collaboration with the founders.

Onsite role in San Francisco and relocation provided if needed.


Salary:


$275,000 to $350,000 Base Salary


Similar Jobs

Explore other opportunities that match your interests

Recruiting Coordinator/Operations

Networking
β€’
5h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

fal

San Francisco Bay Area

Director of Engineering - Safety Infrastructure Engineering

Networking
β€’
1w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

β€’β€’β€’β€’β€’β€’ β€’β€’β€’β€’β€’β€’ β€’β€’β€’β€’β€’β€’
Job Type β€’β€’β€’β€’β€’β€’
Experience Level β€’β€’β€’β€’β€’β€’

Discord

San Francisco Bay Area

Director of Strategic Customer Success

Networking
β€’
1w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

β€’β€’β€’β€’β€’β€’ β€’β€’β€’β€’β€’β€’ β€’β€’β€’β€’β€’β€’
Job Type β€’β€’β€’β€’β€’β€’
Experience Level β€’β€’β€’β€’β€’β€’

tessera data

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly