A

Senior Site Reliability Engineer

asobbi • United State
Remote
Apply
AI Summary

We are seeking a Senior Site Reliability Engineer to build and lead a US-based operations team. The role involves turning runbooks, alerts, and operational workflows into safe, auditable automation. The ideal candidate will have experience in SRE, Platform Engineering, or production infrastructure operations.

Key Highlights
Build and lead a US-based operations team
Turn runbooks, alerts, and operational workflows into automation
Work on critical infrastructure powering the next generation of AI
Key Responsibilities
Building Python-based automation for incident triage, runbook execution, and routine operational tasks
Integrating observability, ITSM, and infrastructure APIs to enrich alerts and automate workflows
Improving monitoring signal quality through correlation, enrichment, suppression, and deduplication
Technical Skills Required
Python Observability/Monitoring Infrastructure Operations
Benefits & Perks
Fully remote work
High autonomy and visibility with leadership
Competitive salary range: $170,000 - $220,000
Nice to Have
GPU, datacentre, or colocation infrastructure experience
ITSM integrations (ServiceNow, Halo, Jira Service Management, or similar)
ChatOps tooling (Slack or Microsoft Teams bots)

Job Description


SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US)

Location: Fully remote - US-based, working a US timezone

Package: $170,000 - $220,000



Overview

We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites.

Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup.

This is a Platform/SRE role with a strong automation and software-engineering bias — not an AI model-building role. You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform.


Why join?

  • Get in at the ground floor of a brand-new US operations function and shape how it runs at scale
  • Work on critical infrastructure powering the next generation of AI
  • Automation-first culture - reduce toil and build tooling, rather than fire fight
  • Real autonomy and high visibility with leadership
  • Fully remote, on a US timezone (West-coast preferred)


What you'll be doing

  • Building Python-based automation for incident triage, runbook execution and routine operational tasks
  • Integrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflows
  • Improving monitoring signal quality through correlation, enrichment, suppression and deduplication
  • Building internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboards
  • Maintaining version-controlled runbook-as-code and automation libraries
  • Turning post-incident learnings into better tooling, automation and operational standards


We're keen to speak with candidates who have

Essential:

  • Experience in SRE, Platform Engineering or production infrastructure operations
  • Hands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)
  • Exposure to incident management / on-call, and converting manual runbooks into automation
  • Strong Python for automation, APIs and integrations


Nice to have:

  • GPU, datacentre or colocation infrastructure experience
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar)
  • ChatOps tooling (Slack or Microsoft Teams bots)
  • OpenTelemetry, logging or distributed tracing experience
  • DCIM, IPAM or hypervisor-control-plane integrations
  • Experience with LLM-assisted or agent-based operational automation


Next steps

Interested? Apply directly or message me for a confidential discussion.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

Jobgether

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

pacer group

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

pacer group

United State

Subscribe our newsletter

New Things Will Always Update Regularly