M

Site Reliability Engineer (SRE) for Blockchain Infrastructure

mpch • New York City Metropolitan Area
Remote
Apply
AI Summary

Join MPCH as a Site Reliability Engineer (SRE) to ensure the stability, scalability, and security of our blockchain infrastructure. You will be responsible for operating, maintaining, and upgrading Canton Network validator nodes and participant nodes in production environments. This role requires strong technical skills, including proficiency with Terraform, Kubernetes, and Linux systems administration.

Key Highlights
Operate, maintain, and upgrade Canton Network validator nodes and participant nodes in production environments
Manage cloud infrastructure on Google Cloud Platform (GCP) and Amazon Web Services (AWS)
Participate in a 24/7 on-call rotation with defined escalation paths
Key Responsibilities
Operate, maintain, and upgrade Canton Network validator nodes and participant nodes in production environments
Manage Canton ledger health, synchronization status, and domain connectivity
Participate in network governance processes including identity and namespace management
Technical Skills Required
Terraform Kubernetes Linux systems administration
Benefits & Perks
Salary range: $120-140k, based upon experience
Equity options
Fully remote team
Nice to Have
Direct production experience running Canton Network validator or domain infrastructure
Familiarity with Daml smart contract concepts and Canton ledger internals
Experience with Helm, Kustomize, or ArgoCD for Kubernetes workload management

Job Description


About the Role

MPCH is looking for a hands-on Site Reliability Engineer (SRE) to join our Engineering (InfraOps) team, responsible for the stability, scalability, and security of our blockchain infrastructure. You will be a key operator in a distributed blockchain network environment, maintaining validator nodes and network-critical components across cloud platforms — with a strong bias toward automation, observability, and resilience.


This role sits at the intersection of blockchain infrastructure, cloud engineering, and SRE practice. You will be part of an on-call rotation providing 24/7 coverage, and you will be expected to respond to, investigate, and resolve incidents with urgency and rigor.


Key Responsibilities

Canton Network Operations

  • Operate, maintain, and upgrade Canton Network validator nodes and participant nodes in production environments
  • Monitor Canton ledger health, synchronization status, and domain connectivity; respond to degradation events
  • Manage Canton domain and sequencer configurations, participant onboarding, and topology changes
  • Execute planned maintenance windows including node upgrades, certificate rotations, and key management procedures
  • Participate in network governance processes including identity and namespace management

Cloud Infrastructure (GCP / AWS)

  • Provision and manage infrastructure on Google Cloud Platform (primary) and AWS (secondary) using both consoles and CLI tooling
  • Administer compute (GKE, GCE, EKS, EC2), networking (VPCs, firewall rules, load balancers, peering), storage, and IAM across environments
  • Manage Kubernetes clusters hosting Canton workloads — deployments, statefulsets, resource quotas, PodDisruptionBudgets, and cluster upgrades
  • Own cloud cost hygiene: right-sizing, reserved capacity planning, and resource tagging discipline

Infrastructure as Code & Automation

  • Author, review, and maintain Terraform modules for all infrastructure provisioning — no manual snowflakes
  • Enforce GitOps workflows; all infrastructure changes go through version control and CI/CD pipelines
  • Leverage AI-assisted tooling (e.g., GitHub Copilot, Claude, or similar) to accelerate IaC authoring, runbook generation, and incident analysis
  • Continuously improve automation to reduce toil and MTTR

SRE / On-Call Duties

  • Participate in a 24/7 on-call rotation with defined escalation paths
  • Respond to pages, triage alerts, and lead incident response for Canton and infrastructure-layer issues
  • Write thorough post-mortems with blameless root cause analysis and tracked action items
  • Build and refine alerting rules, dashboards, and runbooks in partnership with the broader InfraOps team
  • Maintain and test disaster recovery and business continuity procedures for validator infrastructure

Security & Compliance

  • Apply hardening standards to all node infrastructure; manage secrets via Vault or equivalent secrets management tooling
  • Own TLS/mTLS certificate lifecycle for Canton components
  • Support audit and compliance activities including access reviews, change logs, and evidence collection


Required Qualifications

  • 3+ years of experience in infrastructure, cloud operations, or SRE roles in production environments
  • Solid, hands-on experience with Google Cloud Platform (Compute Engine, GKE, Cloud Networking, IAM, Cloud Monitoring); familiarity with AWS is a plus
  • Proficiency with Terraform — you write modules, not just apply them; comfortable with state management, workspaces, and remote backends
  • Strong Kubernetes operations experience: cluster management, workload troubleshooting, networking (CNI, Ingress, NetworkPolicy), and storage
  • Experience operating blockchain or distributed ledger nodes — Canton / Daml, Hyperledger, Ethereum, or similar; understanding of consensus, finality, and node synchronization concepts
  • Familiarity with Canton Network architecture — validators, sequencers, domains, participants, and topology management
  • Proficiency with Linux systems administration, shell scripting (bash/zsh), and standard SRE tooling
  • Experience with observability stacks: metrics (Prometheus/Cloud Monitoring), logging (ELK, Cloud Logging, Loki), tracing, and alerting
  • Demonstrated ability to work an on-call rotation, respond to incidents under pressure, and write high-quality post-mortems
  • Strong written communication skills — clear incident updates, runbooks, and architecture documentation


Foundational IT & General Computing Skills

While this is a cloud and blockchain infrastructure role, candidates must also demonstrate solid general computing competency to operate effectively in a distributed, remote-first team environment:

  • Operating Systems: Comfortable working across macOS, Windows, and Linux desktop/laptop environments; able to navigate system settings, manage files, configure network adapters, and perform basic OS-level troubleshooting (e.g., DNS resolution issues, VPN connectivity, permission problems)
  • Hardware Troubleshooting: Ability to diagnose and resolve common computer hardware and peripheral issues — including display, keyboard/mouse, storage, and connectivity problems — and coordinate with IT support or vendors when escalation is needed
  • Email & Calendar: Proficient with Microsoft Outlook (or equivalent) for email management, calendar scheduling, meeting coordination, and distribution list communication; familiar with shared calendar and resource booking workflows
  • Word Processing: Able to produce clear, well-formatted documents using Microsoft Word or Google Docs — runbooks, incident reports, onboarding guides, and internal proposals
  • Spreadsheets: Comfortable with Microsoft Excel or Google Sheets for basic data tasks: building and maintaining tables, applying formulas (SUM, VLOOKUP, IF statements), filtering/sorting, and producing simple charts for capacity tracking or incident metrics
  • Remote Collaboration Tools: Proficient with video conferencing platforms (Zoom, Google Meet, or Teams), chat/messaging tools (Slack or Teams), and shared documentation platforms (Confluence, Notion, or Google Workspace)
  • File & Document Management: Able to organize, version, and share files effectively using cloud storage (Google Drive, OneDrive, or SharePoint); understands basic folder structure hygiene and access permission management
  • Basic Networking Concepts: Understands home/office network setup sufficiently to self-troubleshoot VPN, Wi-Fi, and connectivity issues that may affect on-call response time

Nice to Have

  • Direct production experience running Canton Network validator or domain infrastructure
  • Familiarity with Daml smart contract concepts and Canton ledger internals
  • Experience with Helm, Kustomize, or ArgoCD for Kubernetes workload management
  • Exposure to HashiCorp Vault or equivalent secrets/PKI management
  • Working knowledge of Python or Go for operational tooling and automation
  • Experience in regulated or compliance-heavy environments (SOC 2, ISO 27001)
  • Familiarity with AI-assisted development workflows and prompt engineering for operational use cases


Benefits

  • Salary range: $120-140k, based upon experience
  • Equity options
  • Fully remote team

Similar Jobs

Explore other opportunities that match your interests

Senior Manager, Operations

Networking
•
2h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Berkeley Payments

France
Visa Sponsorship Relocation Remote
Job Type Part-time
Experience Level Mid-Senior level

premium-gruppe

United State

Major Account Manager

Networking
•
5h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly