Senior Site Reliability Engineer (SRE) - Cloud & Platform Infrastructure
Lead the design and maintenance of scalable, observable, and resilient cloud infrastructure at superbot, shaping reliability practices and on-call culture for a growing engineering team. Own Kubernetes, AWS, and observability stacks while partnering with product teams to drive infrastructure-as-code and CI/CD improvements. Requires 4+ years of hands-on SRE/DevOps experience with production-scale ownership.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
About The Role
You’ll sit at the intersection of software engineering and infrastructure, partnering with product and platform teams to keep our systems observable, scalable, and resilient. From shaping our on-call culture to driving infrastructure-as-code adoption, you’ll have real influence over how BOT builds and operates software at scale — for 500+ people and growing.
Key Responsibilities
- Own reliability across critical services — define and track SLIs/SLOs/SLAs, lead blameless post-mortems, and drive down MTTR through systematic incident management.
- Design, build, and maintain cloud infrastructure on AWS (primary) using Terraform, ensuring environments are reproducible, version-controlled, and auditable.
- Scale and optimize our Kubernetes-based container platform — capacity planning, resource tuning, autoscaling, and cluster lifecycle management.
- Strengthen observability end-to-end: instrument services with Prometheus and Grafana, build actionable alerting, and reduce alert noise through continuous tuning.
- Accelerate CI/CD pipelines (GitHub Actions / GitLab CI) to support frequent, safe deployments — shift reliability left by embedding checks into the delivery workflow.
- Partner with engineering teams as an internal reliability advisor — run game days, advocate for SRE best practices, and help developers build services that operate well from the start.
Searching for Devops roles that provide visa sponsorship? Connect with international employers through Devops Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
- 4+ years in an SRE, DevOps, or platform engineering role with production ownership at scale.
- Hands-on Kubernetes experience — deployment, scaling, networking, and troubleshooting in a production environment.
- Infrastructure-as-code fluency with Terraform (or Pulumi) across a major cloud provider (AWS preferred).
- Observability stack experience — Prometheus, Grafana, and/or equivalent tools with a track record of building meaningful dashboards and alerts.
- Proficiency in at least one scripting or programming language (Python, Go, or Bash) for automation and tooling.
- AWS certification (Solutions Architect, DevOps Engineer, or SysOps Administrator).
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Interested in opportunities specifically in United Arab Emirates? Discover our dedicated Visa Sponsorship Jobs in United Arab Emirates page featuring roles from top employers in this location.
- Competitive Compensation: Enjoy a salary package tailored to your skills and experience, along with performance-based bonuses.
- Comprehensive Benefits: We support your well-being with accommodation, meal allowances, and assistance with work visa processing.
- Work-Life Balance: Unwind with generous holiday and New Year bonuses.
- Top-Tier Equipment: Stay productive with the latest tools, including a MacBook and iPhone.
- Thriving Culture: Immerse yourself in a dynamic, inclusive work environment that fosters growth.
Similar Jobs
Explore other opportunities that match your interests