T

Senior Observability Engineer

talenthop • United State
Remote
Apply Now
AI Summary

Design and manage comprehensive observability platforms to enhance system performance and operational efficiency. Architect scalable monitoring solutions using Prometheus, Grafana, and OpenTelemetry to transform telemetry into actionable insights. Requires 10+ years of experience in SRE or platform engineering with strong proficiency in Go, Python, or Java.

Key Highlights
Design and operate observability platforms covering metrics, logs, traces, events, and synthetic monitoring
Define and enforce SLOs, SLIs, and error budgets to support system reliability
Lead the adoption of OpenTelemetry and develop standards for service instrumentation
Key Responsibilities
Design and operate observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
Architect deployments of Prometheus, Thanos, Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog for scalability and high availability.
Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging.
Define and enforce SLOs, SLIs, and error budgets; build dashboards and alerts to support them.
Create alerting strategies that reduce noise and integrate with on-call tools such as PagerDuty and Opsgenie.
Manage large-scale time-series and log storage balancing retention, performance, and cost.
Design distributed tracing pipelines to assist in diagnosing latency and reliability issues.
Develop self-service tools and libraries to promote adoption of observability standards.
Drive cost management and label cardinality discipline in observability infrastructure.
Lead improvements in incident response readiness through dashboards, alert hygiene, and post-incident analysis tooling.
Collaborate with SRE and platform teams to integrate observability with deployment pipelines and delivery workflows.
Evaluate and recommend observability tools and vendors based on cost, capability, and maturity.
Mentor teams on observability best practices, debugging, and SLO-driven operations.
Maintain documentation, onboarding materials, and runbooks for the observability platform.
Technical Skills Required
Prometheus Grafana OpenTelemetry
Benefits & Perks
Fully Remote
Salary range: $100,000–$160,000 annually

Job Description


This is a Fully Remote Job


1. About Our Client:

This organization operates in the technology consulting and software development industry, providing cloud, AI, data, and enterprise solutions across the United States. It addresses challenges related to scalable and reliable technology infrastructure by delivering expert services that support businesses in these areas.


2. About the Opportunity:

The Observability Engineer role is focused on designing and managing comprehensive observability platforms that enhance engineering teams' confidence in system performance. This position is critical for ensuring the usability, quality, and operational efficiency of monitoring solutions that transform telemetry data into actionable insights for engineering and business stakeholders.


3. Responsibilities:

  • Design and operate observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
  • Architect deployments of Prometheus, Thanos, Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog for scalability and high availability.
  • Develop standards for service instrumentation, including OpenTelemetry adoption, metric naming, label cardinality, and structured logging.
  • Define and enforce SLOs, SLIs, and error budgets; build dashboards and alerts to support them.
  • Create alerting strategies that reduce noise and integrate with on-call tools such as PagerDuty and Opsgenie.
  • Manage large-scale time-series and log storage balancing retention, performance, and cost.
  • Design distributed tracing pipelines to assist in diagnosing latency and reliability issues.
  • Develop self-service tools and libraries to promote adoption of observability standards.
  • Drive cost management and label cardinality discipline in observability infrastructure.
  • Lead improvements in incident response readiness through dashboards, alert hygiene, and post-incident analysis tooling.
  • Collaborate with SRE and platform teams to integrate observability with deployment pipelines and delivery workflows.
  • Evaluate and recommend observability tools and vendors based on cost, capability, and maturity.
  • Mentor teams on observability best practices, debugging, and SLO-driven operations.
  • Maintain documentation, onboarding materials, and runbooks for the observability platform.


4. Requirements:

  • Bachelor’s degree in Computer Science or related field.
  • 10+ years experience in SRE, platform engineering, or observability roles.
  • Extensive hands-on experience with Prometheus, Grafana, and at least one commercial observability platform (Datadog, New Relic, or Splunk).
  • Strong knowledge of OpenTelemetry, distributed tracing, and structured logging.
  • Proficiency in at least one programming language such as Go, Python, or Java.
  • Experience managing high-cardinality, high-throughput metrics and log pipelines.
  • Solid understanding of SLOs, error budgets, and SRE principles.
  • Experience integrating observability with CI/CD and incident management tools.
  • Good knowledge of Linux internals, networking, and container platforms.
  • Excellent communication and collaboration skills.


5. Pay Range and Compensation Package:

  • Salary range: $100,000–$160,000 annually.


Equal Opportunity Statement:

Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.


Note:

TalentHop is a recruitment partner of this role. Please note that all employment decisions, including candidate assessment, interviews, hiring, compensation, and employment terms, are made exclusively by the hiring employer.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

vulncheck

United State
Visa Sponsorship Relocation Remote
Job Type Part-time
Experience Level Associate

torentify

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

torentify

United State

Subscribe our newsletter

New Things Will Always Update Regularly