O

Staff Platform Engineer (AI-Driven Infrastructure & Agentic Engineering)

orion digital Canada
Remote
Apply
AI Summary

Lead the design and implementation of a production-grade platform engineering layer that enables AI-driven, agentic workflows across multi-engine fintech operations. Own Kubernetes, GitOps, and multi-cloud infrastructure while reducing operational toil and ensuring compliance. Drive infrastructure modernization, security, and autonomous remediation at scale for a regulated, publicly listed fintech.

Key Highlights
Build a production-grade platform enabling agentic engineering across regulated fintech domains (lending, payments, wealth).
Own Kubernetes, Terraform, GitOps (ArgoCD), and multi-cloud (AWS/OCI) infrastructure with hands-on deployment ownership.
Reduce operational toil by automating remediation, improving observability, and enforcing policy-as-code for autonomous agents.
Key Responsibilities
Design and implement execution environments for agentic workflows, ensuring sandboxing, reproducibility, and permission boundaries.
Stand up and maintain the MCP (Multi-Cloud Platform) and tool-gateway layer to govern agent interactions with infrastructure.
Transform CI/CD into a deterministic validation loop with autonomous retry mechanisms and human-in-the-loop oversight.
Modernize multi-cloud infrastructure (AWS/OCI) with IaC, GitOps, and dependency mapping while ensuring compliance and security.
Converge divergent platform configurations across business lines while respecting regulatory and functional differences.
Push autonomous remediation for operational alerts, targeting self-resolving incidents and agent-driven incident response.
Track and improve key metrics: deploy frequency, lead time, change failure rate, MTTR, and operational autonomy.
Collaborate with operations, compliance, and support teams to absorb repetitive workflows via platform automation.
Technical Skills Required
Kubernetes Terraform GitOps
Benefits & Perks
Fully Remote Work
Comprehensive Health Benefits (medical, dental, vision, Health Spending Account)
Meaningful equity participation (stock options)
Nice to Have
Experience building internal developer platforms, MCP servers, or agent tooling adopted by engineers.
Non-human identity, secrets architecture, or policy-as-code in regulated environments.
Experience with large-scale re-platforming, especially in zero-downtime scenarios.
Opinions on evaluating infrastructure agents (rarely discussed topic).
Sustained real-world use of coding/ops agents (e.g., Claude Code, Cursor, Codex) in production systems.

Job Description


Location: Remote (Canada)

Compensation: $135,000 - $170,000 CAD + bonus + meaningful equity

About Orion Digital

Orion Digital Corp (NASDAQ/TSX: ORIO) is a publicly listed Canadian fintech operating a multi-engine financial platform spanning lending (Mogo), payments (Carta), and wealth (Intelligent Investing).

We have moved past the traditional fintech model. Capital allocation, decisioning, and risk management sit at the core of how the business creates value, and AI runs as an enabling layer across the platform rather than as a standalone product narrative. That layer is real and in use, and the work now is making it faster, safer, and harder to break.

Why this role exists

This role builds that layer.

Platform engineering here is not a support function. Multiple business lines, regulated entities, and a public company's reporting obligations all run on infrastructure this team owns. Every deployment path, permission boundary, and audit trail either makes faster decisioning possible or quietly caps how fast the business can move.

Which means the platform is where the strategy succeeds or stalls. If deploying takes a day, if an agent cannot safely touch production, if nobody can tell what changed and why, the enabling layer stays theoretical. We are hiring a Staff Platform Engineer to build it properly.

The shift we are hiring for

Most platform teams build for human developers. Ours builds for humans and agents, across multiple business lines and regulated entities. That changes the job:

  • Golden paths become machine-executable contracts, not wiki pages. An agent has to be able to discover them, invoke them, and verify the result
  • Every probabilistic step needs a deterministic gate behind it. Tests, policy checks, and schema validation stop being hygiene and become the control system that lets autonomy scale safely
  • Agents are identities. They need scoped credentials, bounded blast radius, and audit trails that hold up across lending, payments, and securities regulation
  • Observability has to answer a new question. Not just "is the service healthy" but "did the agent do the right thing, and how do we know."

If that is already how you think about the problem, keep reading.

What you will do

Build the substrate for agentic engineering

  • Design and run the execution environments agents work in: sandboxed, reproducible, permissioned, disposable
  • Stand up and own our MCP and tool-gateway layer so agents reach infrastructure through governed interfaces instead of ad hoc credentials
  • Turn CI/CD into a validation loop: fast deterministic gates an agent can retry against until it passes, with humans on the loop rather than in it

Own the platform

  • Multi-cloud Kubernetes, infrastructure as code, and GitOps (ArgoCD) delivery. Hands on keyboard, not diagrams
  • Drive infrastructure modernization across our cloud environments, including the unglamorous parts: dependency mapping, and documentation that survives your absence
  • Serve every product line without building a separate platform for each one. Where they genuinely differ, respect it. Where they have drifted for no reason, converge them
  • Treat security and compliance as design constraints from the first commit rather than a review at the end

Take toil off the company

  • Push autonomous remediation into the paths that page people today. The target is alerts that resolve themselves and incidents where the first responder is an agent arriving with a hypothesis already tested
  • Own the numbers: deploy frequency, lead time, change failure rate, MTTR, and the share of operational work running without a human in the middle. Move them quarter over quarter
  • Work outside engineering too. Operations, compliance, and support are full of repetitive work a well-built platform can absorb

What you need

Non-negotiable

  • Deep, current production experience with Kubernetes, Terraform, and a modern GitOps deployment stack. You have owned what you built, including running it in production rather than handing it off once it shipped
  • Real cloud depth. We run across AWS and OCI, so experience in both is an advantage, but what matters is genuine production depth in at least one and the appetite to get fluent in the other quickly
  • You run coding and ops agents as part of your daily work, and you can show us. Sustained real use on real systems (Claude Code, Cursor, Codex, or equivalent, plus whatever you have wired together yourself), not a trial and an opinion. Be ready to walk through something you shipped this month where an agent did most of the typing and you did the judging
  • Strong opinions about where autonomy belongs and where it does not, informed by having been burned at least once
  • Writing clear enough to be read by a junior engineer and a model

Advantage

  • You have built internal developer platforms, MCP servers, or agent tooling that other engineers actually adopted
  • Non-human identity, secrets architecture, or policy-as-code in a regulated environment
  • Re-platforming at scale, especially where the old system could not go down
  • Opinions about how to evaluate infrastructure agents. Almost nobody has these yet

We are not screening on years of experience as a proxy for capability. We will look at what you have built and how fast you build.

How we work

  • Small teams, high ownership, short distance between decision and deployment
  • We would rather collaborate than dictate, and we know when to do it well versus when to do it quickly
  • Output is the measure

Benefits of working with us:

  • Fully Remote Work - Unless otherwise specified in the job posting, Orion offers flexible remote work, supported by the tools and resources needed to succeed.
  • Comprehensive Health Benefits - Access to medical, dental, and vision coverage plus a Health Spending Account
  • Stock Options - Meaningful equity participation aligned with long-term value creation and company growth
  • Flex Vacation - Benefit from paid time off, including vacation and personal days
  • Wellbeing Programs - Access counselling services, mental health support, and additional wellness resources

Powered by JazzHR

gyQpzvXuwD

Similar Jobs

Explore other opportunities that match your interests

Senior Cloud Security Engineer (CNAPP) - Multi-Cloud Security & DevSecOps

Devops
2d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

SCIEX

Canada
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Jobgether

Canada

Senior System Quality Assurance Analyst

Devops
1w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

telus health

Canada

Subscribe our newsletter

New Things Will Always Update Regularly