P

Site Reliability Engineer

picogrid • United State
Relocation
Apply
AI Summary

Picogrid is seeking a Site Reliability Engineer to own production reliability across cloud and edge, ensuring systems can be relied upon by warfighters in tough battlefield conditions. The ideal candidate will have experience with Kubernetes, observability, and incident response. The role requires 3+ years of experience as an SRE or related role, with deep Kubernetes operations experience and a strong understanding of security principles.

Key Highlights
Own production reliability across cloud and edge
Define and drive reliability SLIs and SLOs
Participate in on-call and incident response
Key Responsibilities
Own, define and drive reliability SLIs and SLOs for cloud deployments
Own, define and drive reliability SLIs and SLOs for edge devices
Participate in on-call and incident response
Technical Skills Required
Kubernetes AWS Terraform
Benefits & Perks
Base salary range: $170,000 - $195,000 per year
Significant stock options
Full health coverage
Nice to Have
GovCloud, FIPS, or other regulated or air-gapped environment experience
Constrained edge hardware experience
Standing up SLO and error-budget tooling from scratch

Job Description


Who We Are

Picogrid is a leading venture-backed defense technology company founded to bridge the decades-long gap between modern technology and the critical demands of national security. Today, we're building the essential infrastructure to unify sensors, autonomy, and operators with our technology deployed in active operations around the world. Our mission is to deliver an operational advantage to secure the United States and its allies.

About The Role

As Picogrid's first Site Reliability Engineer you will own production reliability across cloud and edge, from observability and incident response through node lifecycle, stateful workloads, and a fleet of hardware edge devices in the field. You will help build and define the systems, processes and best practices that ensure Picogrid's systems can be relied upon by our warfighters in even the toughest battlefield conditions. You will work with engineers to build a strong on-call culture where issues are root caused swiftly, and ensure our alerting and monitoring have exceptional coverage and signal-to-noise ratio.

Security is a shared responsibility across all our DevSecOps roles, and as part of a scrappy startup team you will be expected to help stand up new infrastructure and other related DevSecOps tasks as needed.

Responsibilities

  • Own, define and drive our reliability SLIs and SLOs for cloud deployments
  • Own, define and drive our reliability SLIs and SLOs for our edge devices deployed in remote and sometimes contested areas
  • Own the observability stack: Grafana, Prometheus, Loki, and OpenTelemetry, with dashboards versioned in git and alerting rules checked in alongside the code they watch
  • Participate in on-call and incident response: log-first troubleshooting, blameless postmortems, and follow-up hardening
  • Encode reliability into infrastructure as code

Required Qualifications

  • 3+ years of experience as an SRE or related roles
  • Deep Kubernetes operations experience: node lifecycle, workload scheduling, StatefulSets, graceful drains, and live cluster debugging
  • Experience designing comprehensive observability dashboards and high signal-to-noise ratio alerting rules
  • You are a competent and experienced incident responder practicing methodical evidence-first triage, blameless postmortems, and turning incidents into durable guardrails
  • Production Terraform or OpenTofu experience
  • Fluent in AWS including IAM, networking, multi-account environments, and account and workload hardening
  • Experience managing high availability database deployments
  • IoT or edge fleet operation experience
  • Comfortable operating in scrappy, fast-paced environments, and turning ambiguous requirements into concrete solutions
  • You optimize for providing value early in projects and short iteration cycles

Preferred Qualifications

  • GovCloud, FIPS, or other regulated or air-gapped environment experience
  • Constrained edge hardware such as NVIDIA Jetson platforms (AGX Thor, Orin Nano), including shared CPU and GPU memory and thermal constraints
  • Overlay or mesh networking operations: Nebula, WireGuard, Tailscale, or similar
  • Standing up SLO and error-budget tooling (sloth, Pyrra, or equivalent) from scratch
  • Active security clearance

Compensation & Benefits

  • Base salary range: $170,000 - $195,000 per year. Base salary is just one part of your total compensation package at Picogrid.
  • Significant stock options with a high potential upside as an early-stage company
  • 401(k) with employer matching
  • Full health coverage (medical, dental, and vision insurance)
  • Relocation assistance provided (if applicable)
  • Unlimited PTO (two-week minimum) and 11 paid holidays per year
  • Paid parental leave for both parents
  • Lunch provided when working in-office and a fully stocked kitchenette
  • Free EV charging at the HQ
  • Unique office in El Segundo, CA stocked with quality coffee, snacks, and craft beer

Export Control Requirements

To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C.

  • 1157, or (iv) Asylee under 8 U.S.C.
  • 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.

Equal Employment Opportunity (EEO) Policy

Picogrid is committed to providing a professional work environment free from discrimination, harassment, and retaliation. We are an equal opportunity employer and make all employment decisions based on merit, qualifications, and business needs.

Equal Employment Opportunity (EEO) Policy

Picogrid is committed to providing a professional work environment free from discrimination, harassment, and retaliation. We are an equal opportunity employer and make all employment decisions based on merit, qualifications, and business needs.

Compensation Range: $170K - $195K


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

TalentAlly

United State

Junior AI Engineer - AI Solutions Development

Devops
•
8h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Techstra Solutions

United State

Cloud Services Software Engineer III

Devops
•
1d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

allen institute

United State

Subscribe our newsletter

New Things Will Always Update Regularly