We are seeking a Senior Site Reliability Engineer to build and lead a US-based operations team. The role involves turning runbooks, alerts, and operational workflows into safe, auditable automation. The ideal candidate will have experience in SRE, Platform Engineering, or production infrastructure operations.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US)
Location: Fully remote - US-based, working a US timezone
Package: $170,000 - $220,000
Overview
We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites.
Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup.
This is a Platform/SRE role with a strong automation and software-engineering bias — not an AI model-building role. You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform.
Why join?
- Get in at the ground floor of a brand-new US operations function and shape how it runs at scale
- Work on critical infrastructure powering the next generation of AI
- Automation-first culture - reduce toil and build tooling, rather than fire fight
- Real autonomy and high visibility with leadership
- Fully remote, on a US timezone (West-coast preferred)
Interested in remote work opportunities in Devops? Discover Devops Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
What you'll be doing
- Building Python-based automation for incident triage, runbook execution and routine operational tasks
- Integrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflows
- Improving monitoring signal quality through correlation, enrichment, suppression and deduplication
- Building internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboards
- Maintaining version-controlled runbook-as-code and automation libraries
- Turning post-incident learnings into better tooling, automation and operational standards
We're keen to speak with candidates who have
Essential:
- Experience in SRE, Platform Engineering or production infrastructure operations
- Hands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)
- Exposure to incident management / on-call, and converting manual runbooks into automation
- Strong Python for automation, APIs and integrations
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Nice to have:
- GPU, datacentre or colocation infrastructure experience
- ITSM integrations (ServiceNow, Halo, Jira Service Management or similar)
- ChatOps tooling (Slack or Microsoft Teams bots)
- OpenTelemetry, logging or distributed tracing experience
- DCIM, IPAM or hypervisor-control-plane integrations
- Experience with LLM-assisted or agent-based operational automation
Next steps
Interested? Apply directly or message me for a confidential discussion.
Similar Jobs
Explore other opportunities that match your interests