Join MPCH as a Site Reliability Engineer (SRE) to ensure the stability, scalability, and security of our blockchain infrastructure. You will be responsible for operating, maintaining, and upgrading Canton Network validator nodes and participant nodes in production environments. This role requires strong technical skills, including proficiency with Terraform, Kubernetes, and Linux systems administration.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
About the Role
MPCH is looking for a hands-on Site Reliability Engineer (SRE) to join our Engineering (InfraOps) team, responsible for the stability, scalability, and security of our blockchain infrastructure. You will be a key operator in a distributed blockchain network environment, maintaining validator nodes and network-critical components across cloud platforms — with a strong bias toward automation, observability, and resilience.
This role sits at the intersection of blockchain infrastructure, cloud engineering, and SRE practice. You will be part of an on-call rotation providing 24/7 coverage, and you will be expected to respond to, investigate, and resolve incidents with urgency and rigor.
Key Responsibilities
Canton Network Operations
- Operate, maintain, and upgrade Canton Network validator nodes and participant nodes in production environments
- Monitor Canton ledger health, synchronization status, and domain connectivity; respond to degradation events
- Manage Canton domain and sequencer configurations, participant onboarding, and topology changes
- Execute planned maintenance windows including node upgrades, certificate rotations, and key management procedures
- Participate in network governance processes including identity and namespace management
Cloud Infrastructure (GCP / AWS)
- Provision and manage infrastructure on Google Cloud Platform (primary) and AWS (secondary) using both consoles and CLI tooling
- Administer compute (GKE, GCE, EKS, EC2), networking (VPCs, firewall rules, load balancers, peering), storage, and IAM across environments
- Manage Kubernetes clusters hosting Canton workloads — deployments, statefulsets, resource quotas, PodDisruptionBudgets, and cluster upgrades
- Own cloud cost hygiene: right-sizing, reserved capacity planning, and resource tagging discipline
Infrastructure as Code & Automation
- Author, review, and maintain Terraform modules for all infrastructure provisioning — no manual snowflakes
- Enforce GitOps workflows; all infrastructure changes go through version control and CI/CD pipelines
- Leverage AI-assisted tooling (e.g., GitHub Copilot, Claude, or similar) to accelerate IaC authoring, runbook generation, and incident analysis
- Continuously improve automation to reduce toil and MTTR
Interested in remote work opportunities in IT & Network Engineering? Discover IT & Network Engineering Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
SRE / On-Call Duties
- Participate in a 24/7 on-call rotation with defined escalation paths
- Respond to pages, triage alerts, and lead incident response for Canton and infrastructure-layer issues
- Write thorough post-mortems with blameless root cause analysis and tracked action items
- Build and refine alerting rules, dashboards, and runbooks in partnership with the broader InfraOps team
- Maintain and test disaster recovery and business continuity procedures for validator infrastructure
Security & Compliance
- Apply hardening standards to all node infrastructure; manage secrets via Vault or equivalent secrets management tooling
- Own TLS/mTLS certificate lifecycle for Canton components
- Support audit and compliance activities including access reviews, change logs, and evidence collection
Required Qualifications
- 3+ years of experience in infrastructure, cloud operations, or SRE roles in production environments
- Solid, hands-on experience with Google Cloud Platform (Compute Engine, GKE, Cloud Networking, IAM, Cloud Monitoring); familiarity with AWS is a plus
- Proficiency with Terraform — you write modules, not just apply them; comfortable with state management, workspaces, and remote backends
- Strong Kubernetes operations experience: cluster management, workload troubleshooting, networking (CNI, Ingress, NetworkPolicy), and storage
- Experience operating blockchain or distributed ledger nodes — Canton / Daml, Hyperledger, Ethereum, or similar; understanding of consensus, finality, and node synchronization concepts
- Familiarity with Canton Network architecture — validators, sequencers, domains, participants, and topology management
- Proficiency with Linux systems administration, shell scripting (bash/zsh), and standard SRE tooling
- Experience with observability stacks: metrics (Prometheus/Cloud Monitoring), logging (ELK, Cloud Logging, Loki), tracing, and alerting
- Demonstrated ability to work an on-call rotation, respond to incidents under pressure, and write high-quality post-mortems
- Strong written communication skills — clear incident updates, runbooks, and architecture documentation
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Foundational IT & General Computing Skills
While this is a cloud and blockchain infrastructure role, candidates must also demonstrate solid general computing competency to operate effectively in a distributed, remote-first team environment:
- Operating Systems: Comfortable working across macOS, Windows, and Linux desktop/laptop environments; able to navigate system settings, manage files, configure network adapters, and perform basic OS-level troubleshooting (e.g., DNS resolution issues, VPN connectivity, permission problems)
- Hardware Troubleshooting: Ability to diagnose and resolve common computer hardware and peripheral issues — including display, keyboard/mouse, storage, and connectivity problems — and coordinate with IT support or vendors when escalation is needed
- Email & Calendar: Proficient with Microsoft Outlook (or equivalent) for email management, calendar scheduling, meeting coordination, and distribution list communication; familiar with shared calendar and resource booking workflows
- Word Processing: Able to produce clear, well-formatted documents using Microsoft Word or Google Docs — runbooks, incident reports, onboarding guides, and internal proposals
- Spreadsheets: Comfortable with Microsoft Excel or Google Sheets for basic data tasks: building and maintaining tables, applying formulas (SUM, VLOOKUP, IF statements), filtering/sorting, and producing simple charts for capacity tracking or incident metrics
- Remote Collaboration Tools: Proficient with video conferencing platforms (Zoom, Google Meet, or Teams), chat/messaging tools (Slack or Teams), and shared documentation platforms (Confluence, Notion, or Google Workspace)
- File & Document Management: Able to organize, version, and share files effectively using cloud storage (Google Drive, OneDrive, or SharePoint); understands basic folder structure hygiene and access permission management
- Basic Networking Concepts: Understands home/office network setup sufficiently to self-troubleshoot VPN, Wi-Fi, and connectivity issues that may affect on-call response time
Nice to Have
- Direct production experience running Canton Network validator or domain infrastructure
- Familiarity with Daml smart contract concepts and Canton ledger internals
- Experience with Helm, Kustomize, or ArgoCD for Kubernetes workload management
- Exposure to HashiCorp Vault or equivalent secrets/PKI management
- Working knowledge of Python or Go for operational tooling and automation
- Experience in regulated or compliance-heavy environments (SOC 2, ISO 27001)
- Familiarity with AI-assisted development workflows and prompt engineering for operational use cases
Benefits
- Salary range: $120-140k, based upon experience
- Equity options
- Fully remote team
Similar Jobs
Explore other opportunities that match your interests