F

GPU Cluster Infrastructure Engineer

far.ai • United State
Remote Visa Sponsorship
Apply Now

FAR.AI seeks an infrastructure engineer to operate and scale their GPU Kubernetes fleet, focusing on scheduling, storage, and security. Key responsibilities include managing node lifecycles, ensuring fault tolerance for distributed training, and hardening the platform. Requires 3+ years in systems/infrastructure engineering with production Kubernetes for GPU workloads and strong programming skills.

Key Highlights
Operate and scale a large-scale GPU Kubernetes cluster fleet.
Ensure performance, fault tolerance, and security of distributed AI training infrastructure.
Work directly with researchers to address infrastructure challenges and drive platform improvements.
Key Responsibilities
Operate the Kubernetes GPU fleet day to day, handling node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning.
Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams.
Design and run the storage under the fleet, from high-performance shared filesystems for datasets and checkpoints to object storage tiers, quotas, and backups.
Keep multi-node training runs fault-tolerant, owning node health and automated draining, debugging NCCL and fabric problems, tracking down stragglers and flaky GPUs, and building checkpoint and restart patterns.
Harden the platform, covering identity and access, network policy, secrets, workload isolation, and sandboxing for the AI agents that run on the cluster.
Bring new capacity online, acceptance-testing providers on fabric, NCCL, and storage throughput, holding them to their SLAs, and integrating new clusters into the platform with infrastructure as code.
Work directly with research teams on their infrastructure problems and turn the recurring ones into platform fixes.
Share the on-call rotation, runbooks, and postmortems.
Technical Skills Required
Kubernetes GPU Infrastructure Python
Benefits & Perks
Health Insurance
401(k) plan with match
25 days Paid Time Off
WFH Stipend & Equipment
Paid Leave
Visa Sponsorship
Nice to Have
Distributed training infrastructure: multi-node PyTorch and NCCL debugging, the NVIDIA node stack (drivers, GPU Operator, DCGM), InfiniBand or RoCE fabrics, topology-aware placement.
Distributed storage: VAST, Weka, Lustre, Ceph, or object storage at scale; checkpoint I/O.
Cluster security: admission control, RBAC, node and container hardening, sandboxed runtimes (gVisor, Kata, Firecracker), and isolating autonomous agents on shared infrastructure.
Scheduler internals: Kubernetes scheduler plugins or custom controllers, gang scheduling, fair-share and quota, and the utilization, fairness, and latency tradeoffs between them.
Multi-provider platforms: scheduling and storage across clusters at different providers so users see one system, including clusters with no shared network and uneven data locality.

Job Description

FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.
Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Executive

valtrion group

United State

Principal Software Engineer, Threat Data Platform (TDP) - AI/ML Backend & Data Pipeline

Programming
•
1h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United State

Engineering Manager, AI Posture (Backend)

Programming
•
1h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United State

Subscribe our newsletter

New Things Will Always Update Regularly