I

Senior Site Reliability Engineer (SRE Culture & Developer Platform)

italenters • Spain
Remote Visa Sponsorship Relocation
Apply Now
AI Summary

Evolve a reactive operations model into a preventive, data-driven SRE culture for a fast-growing product organization. Own reliability practices by defining SLIs/SLOs/error budgets, improving observability, and automating resilient operations for many microservices. Build and enable an internal developer platform while delivering hands-on reliability engineering with strong Kubernetes, cloud, and programming skills.

Key Highlights
Transition from reactive DevOps to proactive, scalable SRE with strong operational autonomy
Define and manage SLIs, SLOs, and error budgets; improve production reliability via automation and resilience
Build observability and an internal developer platform to enable safe independent deployments
Key Responsibilities
Drive the transition from reactive DevOps to a preventive, scalable SRE model by embedding SRE practices and increasing operational autonomy.
Define and evolve SLIs, SLOs, and error budgets to support better engineering and product decisions and improve production reliability through automation and resilient design.
Design and enhance observability, monitoring, and alerting across a complex microservices environment.
Build and evolve an internal developer platform to enable product teams to deploy and manage infrastructure safely and independently.
Automate repetitive operational work through hands-on coding, primarily using Python or similar languages, and contribute to infrastructure-as-code and cloud platform evolution.
Solve complex reliability challenges across Kubernetes, cloud infrastructure, messaging systems, and databases.
Technical Skills Required
Python Kubernetes Infrastructure as Code
Benefits & Perks
100% remote work model with offices in Barcelona available for workshops and occasional meetings
Private health insurance, with travel, dental, and psychological coverage
Possibility of relocation and visa support for professionals joining from outside Spain
Nice to Have
Knowledge of Kafka, RabbitMQ, PostgreSQL, SQL Server, or similar distributed systems

Job Description


Build the reliability culture, not just the incident response. 🚀


Join a profitable, growing product company transforming how vacation-rental hosts manage their businesses. With ~160 people and a mature engineering organisation, the company is focused on product stability, collaboration and sustainable growth. As a Senior Site Reliability Engineer, you’ll help evolve Platform from reactive operations to a proactive, data-driven SRE culture. You’ll code, automate and enable teams to build and run reliable services autonomously.


Working with ~75 microservices, 800 concurrent containers and 3B monthly requests, you’ll strengthen observability, define reliability targets and help build an internal developer platform that makes infrastructure and deployments easier. Join a small, senior team and make a measurable impact on a product used by customers worldwide.


💼 Responsibilities

  • Drive the transition from reactive DevOps to a preventive, scalable SRE model, embedding SRE practices and increasing engineering teams’ operational autonomy.
  • Define and evolve SLIs, SLOs and error budgets to support better engineering and product decisions, while improving production reliability through automation, resilient design and sustainable operations.
  • Design and enhance observability, monitoring and alerting across a complex microservices environment.
  • Build and evolve an internal developer platform that enables product teams to deploy and manage infrastructure safely and independently.
  • Automate repetitive operational work through hands-on coding, primarily with Python or similar languages, and contribute to infrastructure-as-code and the continuous evolution of the cloud platform.
  • Solve complex reliability challenges across Kubernetes, cloud infrastructure, messaging systems and databases.


🔎 Profile Requirements

  • 5+ years of experience in Site Reliability Engineering, with a clear preventive and reliability-first mindset.
  • Hands-on experience building automation and writing production-quality code in Python or a similar language.
  • Strong practical knowledge of Kubernetes and cloud environments; experience with Google Cloud Platform is highly valuable.
  • Solid experience with infrastructure as code, ideally using Pulumi or comparable tooling.
  • Proven ability to design and improve observability, monitoring and actionable alerting.
  • Practical knowledge of SLIs, SLOs and error budgets—and the judgment to use them to drive meaningful improvements.
  • A systems-thinking approach: you look beyond immediate fixes and design solutions that last.
  • Experience collaborating with software engineering teams and helping them adopt reliable, autonomous ways of working.
  • Knowledge of Kafka, RabbitMQ, PostgreSQL, SQL Server or similar distributed systems is a plus.
  • Fluent English at C1 level, as you will work in an international environment.


💰 Benefits

  • 100% remote work model, with offices in Barcelona available for workshops and occasional meetings.
  • 25 vacation days per year.
  • Private health insurance, with travel, dental, and psychological coverage.
  • Flexible compensation through meal and transport vouchers.
  • Full equipment and setup to work from home.
  • Flexible working hours, with a start time between 8:00 and 10:00 and no strict time tracking.
  • Reduced working hours in August: 35 hours per week.
  • Referral program; for recommending candidates.
  • Possibility of relocation and visa support for professionals joining from outside Spain.


If you want to build reliability into the way an engineering organisation operates—not merely keep the lights on—this is the challenge for you. Apply and let’s talk.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Amazon Web Services (AWS)

Spain

Senior Technical Release Manager

Devops
•
12h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Scopely

Spain
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Barcelona Supercomputing Cente...

Spain

Subscribe our newsletter

New Things Will Always Update Regularly