J

Site Reliability Engineering (SRE) Leader

Remote
This Job is No Longer Active This position is no longer accepting applications

Job Description


Job Title: Site Reliability Engineering (SRE) Lead / Manager

Location: Remote

Experience Level: 10+ years total, 7+ years in SRE/DevOps


About the Role:

We are seeking an experienced and visionary Site Reliability Engineering (SRE) Leader to own the reliability, scalability, and performance of our critical multi-cloud infrastructure. In this leadership role, you will guide a talented team of SREs and collaborate across engineering to champion a culture of automation, robust observability, and continuous improvement. You will be directly responsible for our infrastructure strategy across AWS and Azure, ensuring our platforms are secure, efficient, and highly available.


Key Responsibilities:

  • Lead, mentor, and grow a distributed team of Site Reliability Engineers.
  • Architect, implement, and manage multi-cloud infrastructure (AWS & Azure) using Terraform as the primary Infrastructure as Code (IaC) tool.
  • Own the end-to-end health of our systems by designing and implementing comprehensive monitoring, alerting, and observability solutions with Datadog.
  • Develop, maintain, and optimize secure and efficient CI/CD pipelines using GitHub Actions to accelerate developer velocity.
  • Define, track, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to drive data-informed decisions on reliability.
  • Lead incident response, post-mortems, and root cause analysis, fostering a blameless culture of continuous learning.
  • Manage cloud spending and implement cost-optimization strategies across multiple providers.


Required Qualifications & Skills:

  • 7+ years of hands-on experience in a Site Reliability Engineering (SRE), DevOps, or Cloud Infrastructure role.
  • Proven track record of leading, mentoring, and scaling high-performing engineering teams.
  • Expert-level knowledge and hands-on experience with Terraform for managing multi-cloud environments.
  • Strong proficiency in designing and managing CI/CD pipelines with GitHub Actions.
  • Deep, practical experience with both AWS and Azure cloud platforms.
  • Extensive experience implementing and leveraging observability tools, with a strong preference for Datadog.
  • Strong scripting skills for automation (e.g., Python, Bash).
  • Excellent communication and collaboration skills, with the ability to bridge the gap between technical and non-technical stakeholders.


Preferred Qualifications (Nice-to-Have):

  • Relevant certifications (e.g., AWS Certified DevOps Engineer, HashiCorp Terraform Associate).
  • Hands-on experience with container orchestration platforms, particularly Kubernetes (EKS).
  • Familiarity with ArgoCD, Prometheus, or service mesh technologies.
  • Knowledge of chaos engineering principles and security best practices in the cloud.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Numerator

India
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Thomson Reuters

India

Google Cloud Developer

Devops
2d ago
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Associate

aceolution

India

Subscribe our newsletter

New Things Will Always Update Regularly