T

Senior Infrastructure Engineer (AI/ML Systems)

typesafe ai United State
Visa Sponsorship
Apply Now
AI Summary

Build and scale global AI infrastructure at TypeSafe AI, designing Kubernetes clusters, optimizing GPU-based LLM inference, and managing multi-cloud networking for high-traffic production systems. Own observability, autoscaling, and specialized ML workloads while collaborating with a small, fast-moving team. Requires deep Kubernetes expertise, AWS proficiency, and hands-on IaC experience.

Key Highlights
Own global-scale Kubernetes infrastructure across multiple clouds and regions
Design and optimize GPU-based LLM inference systems for production reliability
Lead networking, observability, and autoscaling for high-traffic AI workloads
Key Responsibilities
Design, deploy, and operate Kubernetes clusters across multiple regions and clouds
Build and maintain infrastructure for global LLM inference workloads, including GPU autoscaling
Manage networking (VPCs, peering, load balancing, DNS, service mesh) and observability (monitoring, alerting, tracing, logging)
Technical Skills Required
Kubernetes Amazon Web Services Infrastructure as Code
Benefits & Perks
Base salary of $150,000–$250,000 plus equity
100% covered health insurance
Visa sponsorship
Nice to Have
Experience with large-scale LLM/ML inference infrastructure (e.g., vLLM, KubeRay)
Kubernetes networking depth (Cilium, Istio, Envoy)
Multi-cloud infrastructure expertise

Job Description


About TypeSafe

TypeSafe AI is an AI lab building machine-native intelligence infrastructure for automation, designed to make decisions within software by combining the intelligence of LLMs with the efficiency and reliability of code into a new shape of AI: System One Models. Based in San Francisco, TypeSafe AI recently launched its first public model, Jev. While others chase benchmarks and academic puzzles, we’ve been quietly rethinking the LLM stack from first principles — building a new kind of general frontier model designed for real-world reliability, decision-making, and autonomy in production. We’re a small, fast-moving team from OpenAI, Google Brain, and Meta/FAIR, backed by top-tier investors. Since mid-2024, we’ve been engineering the foundation for what comes after the current “state-of-the-art” — a model that actually gets things done.


About The Role

We're looking for an Infrastructure Engineer to build and operate the infrastructure behind TypeSafe AI's products at global scale. You'll own the systems that serve millions of users across regions — from provisioning Kubernetes clusters across multiple clouds to optimizing networking for low-latency AI inference. This is a high-impact role on a small, fast-moving team. You'll work across the full infrastructure stack: cloud primitives, container orchestration, networking, observability, and the specialized infra that makes large-scale model inference efficient.


What you'll do

  • Design, deploy, and operate Kubernetes clusters across multiple regions and clouds
  • Build and maintain infrastructure for the platform that powers LLM inference workloads globally
  • Own networking, including VPCs, peering, load balancing, DNS, service mesh, CNI
  • Manage GPU infrastructure and autoscaling for ML workloads
  • Write and maintain infrastructure as code (Pulumi / Python)
  • Operate and improve observability: monitoring, alerting, tracing, logging


Requirements

  • Deep experience with Kubernetes in production at scale: networking, storage, scheduling, upgrades
  • Strong background in AWS
  • Hands-on experience with infrastructure as code (Pulumi, Terraform, or similar)
  • Solid understanding of Linux networking
  • Track record with high-traffic production ML systems
  • Programming fluency, Python preferred


Nice to have

  • Experience with large-scale LLM / ML inference infrastructure (GPU scheduling, model serving, vLLM, KubeRay, Kubernetes-native tooling)
  • Kubernetes networking depth with Cilium or other CNI plugins; service mesh (Istio, Envoy)
  • Multi-cloud infrastructure
  • Background in site reliability engineering including SLOs, incident response, capacity planning


Life at TypeSafe

We’re a small, flat, close-knit team working to make intelligence dependable enough to become part of everyday software. We work fully in person from our San Francisco office near Embarcadero station. We love what we do and care deeply about the work. We strive for excellence and craftsmanship and won’t stop until we get there. When the team wins, we all win, and we enjoy collaborating and inspiring each other to grow—as a team and as individuals. We value emotional honesty, kindness, and bringing your whole self to work. We build machines; we don’t try to be machines. We want TypeSafe to be the place where you do the most impactful work of your career and help define our future as a company.


We provide

  • Base salary of $150k–250k plus equity, based on leveling
  • 100% covered health insurance
  • Daily lunch and dinner
  • Visa sponsorships
  • 401K plans

Similar Jobs

Explore other opportunities that match your interests

Cloud Engineer 5

Devops
5h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State

AI Architect (AVP/VP)

Devops
16h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

wissen technology

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

blue river technology

United State

Subscribe our newsletter

New Things Will Always Update Regularly