A

Site Reliability Engineer - GPU Orchestration Platform

Aurora United State
Visa Sponsorship
Apply
AI Summary

Seeking a Site Reliability Engineer to ensure the dependability of a multi-cloud GPU orchestration platform. Key responsibilities include owning observability, SLIs/SLOs, automation, and incident response. Requires 3+ years of SRE/Production Engineering experience with strong automation and distributed systems knowledge.

Key Highlights
Own observability, SLIs/SLOs, automation, and incident response for a GPU orchestration platform.
Focus on converting reliability into engineering primitives to manage scaling demand.
Work directly with the founding engineering team to define infrastructure operations.
Key Responsibilities
Build and own dashboards, alerts, and distributed tracing with OpenTelemetry, Prometheus, and Grafana.
Define and implement reliability targets across the API layer and internal orchestration services.
Write Python or Go to remove repetitive operational work, including provider API reconciliation, automated health checks, and capacity rebalancing.
Maintain and extend Terraform or Pulumi modules and Kubernetes configurations.
Participate in on-call, drive root-cause analysis, and convert incidents into durable fixes.
Help shape the standards for rollouts, monitoring, runbooks, and reliability reviews.
Technical Skills Required
Python Go Kubernetes
Benefits & Perks
$170K–$250K base salary
Competitive equity
Transfer-only sponsorship support available
Nice to Have
Exposure to GPU, AI/ML, or other accelerated compute workloads.

Job Description


Site Reliability Engineer — GPU Orchestration Platform


San Francisco or Palo Alto, CA · On-site 4 days/week

$170K–$250K base + competitive equity



The company


This is an AI infrastructure company building software that makes GPU compute more accessible and more affordable for leading enterprises, startups, and researchers.


Founded in 2022, the company has raised $80M from Sequoia and Lightspeed and grew revenue 6x last year.


The product is already serving real workloads, and the next constraint is not whether it works in principle. The constraint is whether the platform can stay observable, reliable, and economical as demand grows.


The team is small, so early engineers get direct exposure to consequential infrastructure decisions across the company.



The role


This is a site reliability role for someone who wants ownership over the systems that keep a multi-cloud GPU orchestration platform dependable in production.


You will own observability, SLOs, automation, infrastructure-as-code, and incident response for services that sit close to the customer experience and the company’s revenue.


The goal is to convert reliability into engineering primitives — SLIs, SLOs, automation, and safe deployment patterns — so operational burden does not scale linearly with demand.


You will work directly with the founding engineering team and help define how infrastructure engineering operates as the company scales.



The technical problem


GPU infrastructure has failure modes that are expensive and messy: cloud-provider drift, partial outages, capacity swings, noisy neighbors, inconsistent health signals, and operational work that grows quickly if it is not automated early.


The hard part is not adding more dashboards.


The hard part is reconciling provider state with internal orchestration logic, keeping visibility high across clouds, and building guardrails that let the team move quickly without accumulating hidden reliability debt.



What you'll own


• Observability: build and own dashboards, alerts, and distributed tracing with OpenTelemetry, Prometheus, and Grafana so the team can see system behavior at the level that matters.

• SLIs and SLOs: define and implement reliability targets across the API layer and internal orchestration services, and make sure new features are designed with operability in mind.

• Automation: write Python or Go to remove repetitive operational work, including provider API reconciliation, automated health checks, and capacity rebalancing.

• Infrastructure-as-code: maintain and extend Terraform or Pulumi modules and Kubernetes configurations as the multi-cloud footprint expands.

• Incident response: participate in on-call, drive root-cause analysis, and convert incidents into durable fixes instead of recurring pages.

• Operational playbook: help shape the standards for rollouts, monitoring, runbooks, and reliability reviews so the infrastructure surface stays manageable as the company grows.



Who this is for


You are likely a strong fit if you have:


• 3+ years in SRE, Production Engineering, or Infrastructure roles with real production ownership.

• Built automation, observability, or platform tooling end to end, not just contributed tickets inside a larger system.

• Worked in environments where distributed systems behavior, failure handling, and operational hygiene were core to the job.

• Comfort across multi-cloud infrastructure and the failure modes that come with provider APIs, orchestration, and partial outages.

• Strong instincts for Kubernetes, Terraform or Pulumi, and operationally safe rollouts.

• Experience with Python or Go for infrastructure automation.

• Ability to define useful SLIs and SLOs and reason about them in production, not just in design docs.

• The judgment to improve reliability without creating a heavier process layer than the team needs.

• Bonus: exposure to GPU, AI/ML, or other accelerated compute workloads.



Tech stack


• Automation: Python, Go

• Infrastructure: Terraform, Pulumi, Kubernetes

• Observability: OpenTelemetry, Prometheus, Grafana

• Environment: multi-cloud GPU infrastructure and orchestration services


You do not need to be deep in every tool above, but you should understand how they fit together in a production system with expensive compute and tight reliability constraints.



Why now


The company has already crossed the line from early validation to real usage: growing revenue, real customers, and enough scale that manual operational work will become a drag unless it is engineered away.


This is the point where reliability work stops being a support function and becomes part of the product architecture.


The engineer in this seat will help set the standard for how the platform is observed, operated, and debugged for the next stage of growth.



This role is not for you if


• You want a narrowly scoped role with fully specified tickets.

• You prefer owning alerts over owning the systems that generate them.

• You are uncomfortable with on-call and incident accountability.

• You want to avoid infrastructure work that includes automation, rollouts, and systems design.

• You need a mature playbook before you can make judgment calls.



Compensation and logistics


• Base salary: $170K–$250K

• Equity: competitive

• Employment: full-time

• Location: San Francisco or Palo Alto, CA

• Work model: on-site 4 days per week

• Cadence: all team members come to the Palo Alto office on Mondays

• Visa support: transfer-only sponsorship support available



About Aurora


Aurora helps exceptional engineers find the right role at some of the most ambitious startups worldwide.


We work with teams that value high ownership, strong technical standards, and clear scope.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Strategic Staffing Solutions

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Mindlance

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Raas Infotek

United State

Subscribe our newsletter

New Things Will Always Update Regularly