J

Site Reliability Engineer

Jobgether • United State
Remote
Apply
AI Summary

Site Reliability Engineer to enhance platform reliability, scalability, and operational performance across cloud-based SaaS environments. Lead incident investigations, implement proactive reliability solutions, and collaborate with engineering and cloud operations teams. Requires 5+ years SRE experience, strong observability and cloud platform expertise, and excellent communication skills.

Key Highlights
Improve platform resilience and reduce incident impact
Lead post-incident analysis and root cause investigations
Develop and implement preventative reliability strategies
Configure and maintain observability tools for monitoring and alerting
Collaborate with engineering, cloud operations, and business teams
Key Responsibilities
Enhance platform resilience and improve incident management processes
Lead post-incident investigations and perform detailed root cause analyses
Create clear and actionable Root Cause Analysis (RCA) documentation
Develop and implement preventative strategies to improve system reliability
Monitor and improve key reliability metrics including time to resolution
Configure and maintain observability tools for accurate monitoring and alerting
Build client-focused dashboards and alerts to proactively identify performance challenges
Collaborate with Engineering, Cloud Operations, and SRE teams to implement improvements
Provide guidance on effective use of observability tools during investigations
Create feedback loops with engineering teams to improve development practices
Contribute to automation initiatives that streamline incident response workflows
Support SaaS platform reliability by identifying risks and reducing incident impact
Troubleshoot application and infrastructure issues across cloud environments
Technical Skills Required
Site Reliability Engineering Observability Cloud Operations Incident Management Automation Performance Optimization AWS JavaScript SQL C#/.NET Windows environments SQL Server
Benefits & Perks
Competitive base salary range of $100,000 - $120,000 per year
Eligibility for performance-based bonus opportunities
Medical, dental, and vision coverage
Fully paid vision insurance
Short-term and long-term disability insurance
Basic life insurance
Flexible paid time off options
Paid company holidays
Hybrid work environment with many positions eligible for fully remote work
Generous family leave options
401(k) retirement plan with company match up to 4%
Flexible spending accounts
Health savings accounts
Commuter benefits
Dependent care savings accounts
Employee Assistance Program
Education assistance programs
Wellness reimbursement programs
Additional voluntary benefits including pet insurance, critical illness coverage, and life insurance options

Job Description


This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in United States.

The Site Reliability Engineer will play a critical role in improving platform reliability, scalability, and operational performance across cloud-based SaaS environments.

This position focuses on reducing incident impact, accelerating resolution times, and building proactive solutions that prevent future disruptions.

You will investigate complex technical issues, lead post-incident analysis, and collaborate with engineering teams to strengthen system stability.

The role combines reliability engineering, observability, automation, and performance optimization to support mission-critical platforms.

You will work closely with Cloud Operations, Engineering, and business teams to improve operational maturity and customer outcomes.

This is an opportunity for an experienced SRE professional to influence reliability practices within a growing technology environment.

Accountabilities

The Site Reliability Engineer will be responsible for enhancing platform resilience, improving incident management processes, and driving proactive reliability initiatives. This role requires strong technical expertise, analytical thinking, and collaboration across multiple teams.

  • Lead post-incident investigations and perform detailed root cause analyses to identify failures and prevent recurring issues.
  • Create clear and actionable Root Cause Analysis (RCA) documentation for internal teams and customer delivery.
  • Develop and implement preventative strategies to improve system reliability and reduce operational disruptions.
  • Monitor and improve key reliability metrics, including time to resolution and incident response effectiveness.
  • Configure and maintain observability tools to ensure accurate monitoring, alerting, and performance visibility.
  • Build client-focused dashboards and alerts to proactively identify application and platform performance challenges.
  • Collaborate with Engineering, Cloud Operations, and SRE teams to implement improvements that enhance scalability and stability.
  • Provide guidance and knowledge sharing on effective use of observability tools during investigations.
  • Create feedback loops with engineering teams to improve development practices, operational patterns, and platform performance.
  • Contribute to automation initiatives that streamline incident response and operational workflows.
  • Support SaaS platform reliability by identifying risks, improving processes, and reducing incident impact.
  • Troubleshoot application and infrastructure issues across cloud environments and enterprise systems.

Requirements

The ideal candidate brings strong experience in site reliability engineering, cloud operations, incident management, and modern software development practices. They should be comfortable working in complex environments and solving ambiguous technical challenges.

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of professional experience in a Site Reliability Engineering role.
  • Strong understanding of SRE principles, reliability practices, and incident management processes.
  • Experience with observability platforms such as New Relic, Datadog, Sumo Logic, or similar tools.
  • Proficiency reading and writing code using technologies such as JavaScript, .NET, and SQL.
  • Experience troubleshooting C#/.NET web applications and identifying performance issues.
  • Familiarity with cloud platforms such as AWS or Azure, with AWS experience strongly preferred.
  • Experience operating within public cloud environments and SaaS platforms.
  • Strong understanding of cloud architecture patterns and operational best practices.
  • Knowledge of CI/CD pipelines and Infrastructure as Code (IaC) environments is preferred.
  • Experience troubleshooting Windows environments and SQL Server is preferred.
  • Strong analytical, problem-solving, and data-driven decision-making skills.
  • Ability to work effectively in environments with evolving processes and varying operational maturity.
  • Excellent written and verbal communication skills with the ability to collaborate across technical and business teams.

Benefits

  • Competitive base salary range of $100,000 - $120,000 per year, depending on experience, skills, location, and business needs.
  • Eligibility for performance-based bonus opportunities.
  • Medical, dental, and vision coverage for employees and eligible dependents.
  • Fully paid vision insurance, short-term and long-term disability insurance, and basic life insurance.
  • Flexible paid time off options and paid company holidays where available.
  • Hybrid work environment with many positions eligible for fully remote work.
  • Generous family leave options, including adoption and foster care support.
  • 401(k) retirement plan with company match up to 4%.
  • Flexible spending accounts, health savings accounts, commuter benefits, and dependent care savings accounts.
  • Employee Assistance Program providing confidential personal and professional support.
  • Education assistance programs for certifications and professional development.
  • Wellness reimbursement programs supporting healthy habits and employee wellbeing.
  • Additional voluntary benefits, including pet insurance, critical illness coverage, and life insurance options.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


Similar Jobs

Explore other opportunities that match your interests

Senior Linux Graphics Engineer

Programming
•
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Jobgether

United State

Engineering Team Lead

Programming
•
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Director

aweber

United State

Manager, Global Technical Recruiting

Programming
•
8h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Outreach

United State

Subscribe our newsletter

New Things Will Always Update Regularly