J

Senior Reliability Monitoring Engineer

Jobgether • United State
Remote
Apply
AI Summary

Design, build, and maintain observability solutions. Improve system reliability and performance. Collaborate with teams and stakeholders.

Key Highlights
Design and maintain observability platforms
Improve monitoring maturity and reduce operational noise
Collaborate with engineering and business teams
Key Responsibilities
Design, implement, and maintain observability solutions
Build and operate scalable telemetry collection pipelines
Configure and optimize platforms such as Prometheus, Grafana
Develop monitoring strategies to improve system visibility
Support SRE practices through SLOs, error budgets, and reliability improvements
Technical Skills Required
Observability platforms Cloud-native infrastructure Prometheus, Grafana, OpenTelemetry
Benefits & Perks
Competitive annual salary range of $100,000-$150,000
Fully remote work opportunity within the United States
Opportunity to work on modern reliability and monitoring technology initiatives
Nice to Have
Experience with Thanos, Mimir, Cortex, Loki, or Tempo
Familiarity with eBPF-based monitoring tools
Exposure to regulated environments requiring audit-ready logging and monitoring

Job Description


This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Reliability Monitoring Engineer based in the United States.

The Reliability Monitoring Engineer will design, build, and maintain observability solutions that help engineering teams understand, operate, and improve complex systems.

This role focuses on creating reliable monitoring strategies across metrics, logs, traces, dashboards, and alerting workflows.

The ideal candidate will transform large volumes of operational data into actionable insights that improve system reliability and performance.

You will work across modern cloud-native environments, supporting scalable platforms while optimizing visibility and operational efficiency.

This position requires strong technical expertise, a proactive approach to problem-solving, and the ability to collaborate with engineering teams and stakeholders.

It is an opportunity to shape observability practices, improve incident response capabilities, and strengthen the reliability of critical technology platforms.

Accountabilities

The Reliability Monitoring Engineer will own the development and operation of observability platforms, ensuring systems provide accurate, actionable, and efficient operational insights. This role will focus on improving monitoring maturity, reducing operational noise, and enabling teams to make data-driven reliability decisions.

  • Design, implement, and maintain observability solutions covering metrics, logging, tracing, dashboards, and alerting workflows.
  • Build and operate scalable telemetry collection pipelines and monitoring infrastructure.
  • Configure and optimize platforms such as Prometheus, Grafana, and commercial observability solutions.
  • Develop monitoring strategies that improve system visibility, reliability, and operational efficiency.
  • Implement distributed tracing, structured logging, and OpenTelemetry-based solutions.
  • Manage high-volume, high-cardinality metrics and log environments while ensuring performance and scalability.
  • Create and maintain dashboards, alerts, and reporting solutions that provide actionable insights for engineering teams.
  • Support Site Reliability Engineering (SRE) practices through service level objectives (SLOs), error budgets, and reliability improvements.
  • Integrate observability tools with CI/CD pipelines, incident management systems, and operational workflows.
  • Analyze monitoring data to identify trends, improve performance, and optimize infrastructure costs.
  • Collaborate with engineering and business teams to translate technical signals into meaningful operational outcomes.

Requirements

The ideal candidate will have strong experience in reliability engineering, observability platforms, and cloud-native infrastructure. They should be comfortable designing scalable monitoring solutions, troubleshooting complex environments, and communicating technical insights across teams.

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of experience in SRE, platform engineering, reliability engineering, or observability-focused roles.
  • Hands-on experience with Prometheus, Grafana, and at least one major observability platform such as Datadog, New Relic, or Splunk.
  • Strong understanding of OpenTelemetry, distributed tracing, telemetry pipelines, and structured logging practices.
  • Proficiency in at least one programming language such as Go, Python, or Java.
  • Experience managing high-throughput metrics and log processing systems.
  • Solid understanding of SRE principles, including SLOs, SLIs, and error budgets.
  • Experience integrating observability solutions with CI/CD platforms and incident response processes.
  • Strong knowledge of Linux systems, networking concepts, and container technologies.
  • Excellent troubleshooting, analytical, communication, and collaboration skills.
  • Experience with observability technologies such as Thanos, Mimir, Cortex, Loki, or Tempo is a plus.
  • Familiarity with eBPF-based monitoring tools, open-source observability projects, and cost optimization strategies is preferred.
  • Exposure to regulated environments requiring audit-ready logging and monitoring is beneficial.

Benefits

  • Competitive annual salary range of $100,000-$150,000, depending on experience and qualifications.
  • Fully remote work opportunity within the United States.
  • Full-time employment with opportunities for professional growth and career advancement.
  • Opportunity to work on modern reliability, monitoring, and cloud technology initiatives.
  • Collaborative environment focused on innovation and technical excellence.
  • Exposure to large-scale systems, distributed platforms, and advanced observability practices.
  • Benefits package and employee support programs.
  • Opportunity to contribute to impactful technology solutions across enterprise environments.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


Similar Jobs

Explore other opportunities that match your interests

Algorithmic Trading Developer

Programming
•
56m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State

Junior Backend Developer

Programming
•
1h ago
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Entry level

hiredbuddy

United State

Full Stack Senior Python Engineer

Programming
•
1h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Perficient

United State

Subscribe our newsletter

New Things Will Always Update Regularly