D

Staff Site Reliability Engineer

Doghouse Recruitment European Union
Remote
Apply
AI Summary

We're seeking a Staff level SRE who will own production reliability end-to-end for our client. You'll define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency. You'll work close to the metal across Kubernetes internals, Linux performance, and network debugging.

Key Highlights
Own production reliability end-to-end
Define SLIs/SLOs and run error budget conversations
Ship changes that reduce incidents and improve latency
Key Responsibilities
Define SLIs/SLOs
Run error budget conversations
Ship changes that reduce incidents and improve latency
Build automation to kill toil
Improve deployment safety
Turn observability into signal rather than noise
Technical Skills Required
Linux Kubernetes Networking
Benefits & Perks
100% remote within the EU
On-call is part of the job, but success is measured by how much you reduce it

Job Description


Staff Site Reliability Engineer – Bare Metal Linux – Data Center – Networking

Location: 100% remote within the EU


Our client is building a cloud platform for high-throughput, compute-heavy workloads. They operate large-scale infrastructure where failure modes are real, capacity is finite, and reliability needs to be engineered, not "handled".


We're seeking a Staff level SRE who will own production reliability end-to-end for our client: define SLIs/SLOs, run error budget conversations, and ship changes that reduce incidents and improve latency (p95/p99). You'll build automation to kill toil, improve deployment safety (canary/rollback), and turn observability into signal rather than noise.


This is a bare-metal environment: think Linux, datacenters, physical fleets, and real hardware constraints, not managed services. You'll work close to the metal across Kubernetes internals (scheduling, autoscaling behavior, kubelet pressure/evictions, etcd/control plane), Linux performance (CPU/memory/IO contention), and network debugging (DNS/TCP/TLS, packet loss, congestion). On-call is part of the job, but success is measured by how much you reduce it.


Requirements:

• Production Engineering experience running / on-prem / data center infrastructure (not public cloud only)

• Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO, -level behaviors)

• Strong understanding of networking (DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting under load)

• Strong Kubernetes experience beyond manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane

• Experience with Terraform, Docker, Helm, and modern CI/CD practices

• Strong recent coding skills in Go, and/or Python is a must for this role

• Experience in Low Latency environments.


If you're looking for complexity and a new place to nerd out on infrastructure optimization, we'd love to hear from you!


Similar Jobs

Explore other opportunities that match your interests

Senior Camunda 8 / Java Engineer - Workflow Transformation

Devops
3h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Radley James

European Union
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Infinity Quest

European Union
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Edison Smart®

European Union

Subscribe our newsletter

New Things Will Always Update Regularly