C

Platform Engineer

clera United State
Remote Visa Sponsorship
Apply
AI Summary

Own reliability, scale, and performance of core AI/ML platform infrastructure. Build and maintain AWS systems using Terraform, Kubernetes/EKS, Docker, and CI/CD pipelines. Ensure high availability, cost efficiency, and developer productivity for production services.

Key Highlights
Production ownership of uptime, latency, cost, and incident response
AWS infrastructure with Terraform, Kubernetes/EKS, Docker, and CI/CD
Observability, alerting, and on-call workflows for rapid failure resolution
Key Responsibilities
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management
Design and improve backend and platform systems for scale including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths
Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows for rapid failure detection and resolution
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows to improve developer productivity and reduce production risk
Write clean, maintainable code to automate systems, improve backend services, and create internal tooling
Technical Skills Required
AWS Kubernetes Terraform CI/CD
Benefits & Perks
Equity participation
Visa sponsorship
Remote work
Nice to Have
Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services
Background operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms
Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability

Job Description


About The Role

We're a fast-moving AI/ML platform startup building infrastructure for reinforcement learning environments, post-training data pipelines, and large-scale agent evaluation. Our engineering team of ~15 includes exceptional technical talent — competition medalists, serial AI startup founders, and published researchers.

As a Platform Engineer, you'll own the reliability, scale, performance, and developer experience of our core infrastructure and systems. This is a backend-architecture-heavy role with real production ownership — your work directly shapes how fast, reliable, and cost-effective our platform is to build on and run.

What You'll Do

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Design and improve backend and platform systems for scale — including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
  • Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
  • Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
  • Write clean, maintainable code to automate systems, improve backend services, and create internal tooling.

What We're Looking For

Required

  • 2–4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost.
  • Deep hands-on experience with AWS and containerized systems; strong familiarity with Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
  • Proven track record building or operating CI/CD, release automation, observability, alerting, and incident response systems.
  • Strong backend engineering judgment — ability to reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes.
  • Ability to write clean, maintainable code and apply software engineering judgment across infrastructure, backend systems, and developer workflows.
  • High ownership mindset; comfortable being accountable for production systems end-to-end.

Nice to Have

  • Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
  • Background operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms.
  • Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability.

Location

This role supports a few location arrangements:

  • San Francisco, CA (on-site) — preferred for US-based candidates.
  • Singapore (on-site) — for Southeast Asia-based candidates.
  • Fully remote (independent contractor) — open to candidates elsewhere, particularly in Europe.

Visa sponsorship is available.

Compensation & Benefits

  • Salary: $150,000 – $250,000 USD annually (full-time, US-based).
  • Equity participation in an early-stage, well-funded AI startup.
  • Work alongside a world-class technical team on infrastructure that operates at real scale.
  • High degree of autonomy and direct impact on product and platform direction.

Similar Jobs

Explore other opportunities that match your interests

Azure Cloud Engineer

Devops
8h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State

AWS Cloud Engineer

Devops
10h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

moon tiger

United State

Subscribe our newsletter

New Things Will Always Update Regularly