Seeking a Site Reliability Engineer to ensure the dependability of a multi-cloud GPU orchestration platform. Key responsibilities include owning observability, SLIs/SLOs, automation, and incident response. Requires 3+ years of SRE/Production Engineering experience with strong automation and distributed systems knowledge.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Site Reliability Engineer — GPU Orchestration Platform
San Francisco or Palo Alto, CA · On-site 4 days/week
$170K–$250K base + competitive equity
The company
This is an AI infrastructure company building software that makes GPU compute more accessible and more affordable for leading enterprises, startups, and researchers.
Founded in 2022, the company has raised $80M from Sequoia and Lightspeed and grew revenue 6x last year.
The product is already serving real workloads, and the next constraint is not whether it works in principle. The constraint is whether the platform can stay observable, reliable, and economical as demand grows.
The team is small, so early engineers get direct exposure to consequential infrastructure decisions across the company.
The role
This is a site reliability role for someone who wants ownership over the systems that keep a multi-cloud GPU orchestration platform dependable in production.
You will own observability, SLOs, automation, infrastructure-as-code, and incident response for services that sit close to the customer experience and the company’s revenue.
The goal is to convert reliability into engineering primitives — SLIs, SLOs, automation, and safe deployment patterns — so operational burden does not scale linearly with demand.
You will work directly with the founding engineering team and help define how infrastructure engineering operates as the company scales.
Searching for Devops roles that provide visa sponsorship? Connect with international employers through Devops Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
The technical problem
GPU infrastructure has failure modes that are expensive and messy: cloud-provider drift, partial outages, capacity swings, noisy neighbors, inconsistent health signals, and operational work that grows quickly if it is not automated early.
The hard part is not adding more dashboards.
The hard part is reconciling provider state with internal orchestration logic, keeping visibility high across clouds, and building guardrails that let the team move quickly without accumulating hidden reliability debt.
What you'll own
• Observability: build and own dashboards, alerts, and distributed tracing with OpenTelemetry, Prometheus, and Grafana so the team can see system behavior at the level that matters.
• SLIs and SLOs: define and implement reliability targets across the API layer and internal orchestration services, and make sure new features are designed with operability in mind.
• Automation: write Python or Go to remove repetitive operational work, including provider API reconciliation, automated health checks, and capacity rebalancing.
• Infrastructure-as-code: maintain and extend Terraform or Pulumi modules and Kubernetes configurations as the multi-cloud footprint expands.
• Incident response: participate in on-call, drive root-cause analysis, and convert incidents into durable fixes instead of recurring pages.
• Operational playbook: help shape the standards for rollouts, monitoring, runbooks, and reliability reviews so the infrastructure surface stays manageable as the company grows.
Who this is for
You are likely a strong fit if you have:
• 3+ years in SRE, Production Engineering, or Infrastructure roles with real production ownership.
• Built automation, observability, or platform tooling end to end, not just contributed tickets inside a larger system.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
• Worked in environments where distributed systems behavior, failure handling, and operational hygiene were core to the job.
• Comfort across multi-cloud infrastructure and the failure modes that come with provider APIs, orchestration, and partial outages.
• Strong instincts for Kubernetes, Terraform or Pulumi, and operationally safe rollouts.
• Experience with Python or Go for infrastructure automation.
• Ability to define useful SLIs and SLOs and reason about them in production, not just in design docs.
• The judgment to improve reliability without creating a heavier process layer than the team needs.
• Bonus: exposure to GPU, AI/ML, or other accelerated compute workloads.
Tech stack
• Automation: Python, Go
• Infrastructure: Terraform, Pulumi, Kubernetes
• Observability: OpenTelemetry, Prometheus, Grafana
• Environment: multi-cloud GPU infrastructure and orchestration services
You do not need to be deep in every tool above, but you should understand how they fit together in a production system with expensive compute and tight reliability constraints.
Why now
The company has already crossed the line from early validation to real usage: growing revenue, real customers, and enough scale that manual operational work will become a drag unless it is engineered away.
This is the point where reliability work stops being a support function and becomes part of the product architecture.
The engineer in this seat will help set the standard for how the platform is observed, operated, and debugged for the next stage of growth.
Interested in opportunities specifically in United State? Discover our dedicated Visa Sponsorship Jobs in United State page featuring roles from top employers in this location.
This role is not for you if
• You want a narrowly scoped role with fully specified tickets.
• You prefer owning alerts over owning the systems that generate them.
• You are uncomfortable with on-call and incident accountability.
• You want to avoid infrastructure work that includes automation, rollouts, and systems design.
• You need a mature playbook before you can make judgment calls.
Compensation and logistics
• Base salary: $170K–$250K
• Equity: competitive
• Employment: full-time
• Location: San Francisco or Palo Alto, CA
• Work model: on-site 4 days per week
• Cadence: all team members come to the Palo Alto office on Mondays
• Visa support: transfer-only sponsorship support available
About Aurora
Aurora helps exceptional engineers find the right role at some of the most ambitious startups worldwide.
We work with teams that value high ownership, strong technical standards, and clear scope.
Similar Jobs
Explore other opportunities that match your interests