We're seeking a Staff level SRE who will own production reliability end-to-end for our client. You'll define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency. You'll work close to the metal across Kubernetes internals, Linux performance, and network debugging.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
Staff Site Reliability Engineer – Bare Metal Linux – Data Center – Networking
Location: 100% remote within the EU
Our client is building a cloud platform for high-throughput, compute-heavy workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not "handled".
We're seeking a Staff level SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You'll build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise.
Interested in remote work opportunities in Devops? Discover Devops Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
This is a bare-metal environment: think Linux, datacenters, physical fleets, and real hardware constraints, not managed services. You'll work close to the metal across Kubernetes internals (scheduling, autoscaling behavior, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/IO contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it.
Requirements:
• Production Engineering experience running / on-prem / data center infrastructure (not public cloud only)
• Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO, -level behaviors)
• Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting under load)
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
• Strong Kubernetes experience beyond manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane
• Experience with Terraform, Docker, Helm, and modern CI/CD practices
• Strong recent coding skills in Go, and/or Python is a must for this role
• Experience in Low Latency environments.
If you're looking for complexity and a new place to nerd out on infrastructure optimization, we'd love to hear from you!
Similar Jobs
Explore other opportunities that match your interests