B

Senior Site Reliability Engineer (SRE) for Distributed AI Training Platform

basis set • United State
Visa Sponsorship Relocation
Apply
AI Summary

Join Thinking Machines to build AI infrastructure that enhances human decision-making. As a Senior SRE, you’ll design and maintain the reliability of Tinker’s distributed AI training platform, ensuring robust observability, incident response, and multi-tenant isolation for scalable GPU workloads. Collaborate with engineering and research teams to optimize performance while balancing speed and reliability.

Key Highlights
Own end-to-end reliability for a distributed AI training platform, from CI/CD to production observability and incident response
Define and enforce SLOs for distributed training systems, balancing job completion reliability and scheduling latency
Collaborate with research and engineering teams to harden multi-tenant isolation and resource scheduling for LoRA-based workloads
Key Responsibilities
Define and own end-to-end reliability, including CI/CD flows, production observability, and incident response
Develop and enforce Service Level Objectives (SLOs) for distributed training systems, balancing reliability and scheduling latency
Design and implement monitoring and observability across the full training pipeline
Lead incident response for Tinker platform issues, ensuring rapid recovery and systematic improvements
Harden multi-tenant isolation and resource scheduling to maximize GPU utilization without compromising reliability or data separation
Collaborate with security teams to address production vulnerabilities
Technical Skills Required
Distributed Systems Kubernetes Incident Response & Postmortems
Benefits & Perks
Generous health, dental, and vision benefits
Unlimited paid time off (PTO)
Paid parental leave
Nice to Have
Deep experience operating production cloud services at scale (e.g., AWS, GCP)
Background in distributed training frameworks and infrastructure failure handling
Track record building checkpoint and recovery systems for long-running distributed jobs

Job Description


Home

Portfolio

Portfolio

Team

Team

Research

Research

Jobs

Jobs

Home

Portfolio

Portfolio

Team

Team

Research

Research

Jobs

Jobs

Let's build together...

Join our exceptional founders forging an intelligent future

Search

jobs

936

Explore

companies

117

Join talent network

Talent

Site Reliability Engineer (SRE)

Thinking Machines Lab

Software Engineering

San Francisco, CA, USA

Posted on Aug 5, 2026

Apply now

The mission of Thinking Machines is to build AI that extends human will and judgment.

About Tinker

Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs. Tinker is rapidly adding new customers, features, and novel use-cases. We’re hiring to grow the platform alongside the Tinker community.

About The Role

We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient.

What You’ll Do

  • Define and own end-to-end reliability, from CI/CD flows to production observability and incident response.
  • Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity.
  • Design and implement monitoring and observability across the full training path.
  • Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence.
  • Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation
  • Collaborate with security teams to address production vulnerabilities

Skills And Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Experience in distributed systems, cloud infrastructure, or site reliability engineering.
  • Proficiency writing software to solve reliability problems, including building tooling and automation.
  • Experience with production incident response, postmortems, and systematic reliability improvement.
  • Strong communication skills and track record of coordination across engineering and research teams.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services)
  • Background in distributed training frameworks and how infrastructure failures surface in training behavior.
  • Track record building checkpoint and recovery systems for long-running distributed jobs.
  • Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads.

Logistics

  • Location: This role is based in San Francisco, California.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 – $475,000 USD.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Apply now

See more open positions at Thinking Machines Lab

Powered by Getro.com

Privacy policyCookie policy

Portfolio

Portfolio

Research

Research

Team

Team

Read our FAQs

© 2025 Basis Set Builders. All rights reserved

Basis Set Ventures®, BSV® and the Basis Set Logo are registered trademarks of Basis Set Builders, LLC

Similar Jobs

Explore other opportunities that match your interests

Head of IT Service Management

Networking
•
3h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

the talent magnet

United State

IT Engineer

Networking
•
8h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

basis set

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

phota labs

United State

Subscribe our newsletter

New Things Will Always Update Regularly