L

Senior Cloud Infrastructure Engineer (AI Platform)

limrun (yc p26) United State
Relocation
Apply
AI Summary

Build and scale a cloud-based platform for AI agents, designing and maintaining Go services, Kubernetes controllers, and tenant isolation systems. Own projects from conception to production, including incident response and operational reliability improvements. Collaborate with founders to shape product direction and infrastructure for AI-driven development workflows.

Key Highlights
Own end-to-end development of cloud services, APIs, and Kubernetes controllers in Go
Drive scalability, performance, and reliability for AI agent infrastructure
Investigate and resolve production incidents with cross-system debugging expertise
Key Responsibilities
Design, build, and launch cloud services, APIs, and Kubernetes controllers in Go
Optimize startup times, scheduling, and compute resource utilization for scaling demands
Implement and maintain tenant isolation, access controls, and networking for secure workloads
Automate host provisioning, upgrades, and failure recovery processes
Collaborate with customers to refine APIs and tooling based on real-world workflows
Technical Skills Required
Go Kubernetes Cloud Infrastructure
Benefits & Perks
Remote-friendly environment with flexible working hours
San Francisco office in South Park with relocation support
Regular team offsites and company-covered work travel across Europe and the US
Nice to Have
Experience with sandboxing, virtual machines, or Linux/macOS host management
Hands-on work with DNS, routing, proxies, or private networking
Experience building APIs, SDKs, CLIs, CI systems, or remote build tools
Experience using or building AI agents for platform maintenance and production operations
Familiarity with mobile development workflows using simulators and emulators

Job Description


Help us expand what AI agents can build. Limrun runs development tools and applications as cloud services that agents can use from wherever they run. We're hiring an engineer to build and scale that platform.

The role

You'll help decide what to build and own projects from design through ongoing maintenance. You'll also join the on-call rotation, investigate production incidents, and fix their causes.

What you'll do

• Build and launch new tools, features, and cloud services, including the Go services, APIs, and Kubernetes controllers behind them.
• Reduce startup times, improve scheduling, and make better use of compute as demand grows.
• Build and maintain the tenant isolation, access controls, and networking that protect customer workloads and data.
• Automate host provisioning, upgrades, and recovery from failures.
• Work with customers to understand their workflows and turn problems they encounter into better APIs, tooling, and product features.

What we're looking for

• You've built, operated, and scaled production cloud infrastructure, with responsibility for deployment, observability, and reliability.
• You've written and maintained Go services in production and can reason about concurrent code.
• You understand Kubernetes architecture and controller patterns well enough to build on them and debug their behavior.
• You can investigate failures across application code, networking, and operating systems.
• You've shipped production software or infrastructure using AI-assisted development.

We care about your work and the responsibility you've taken for it more than titles or years of experience.

Nice to have, not required

• Experience with sandboxing, virtual machines, or managing Linux and macOS hosts.
• Hands-on work with DNS, routing, proxies, tunnels, or private networking.
• Experience building APIs, SDKs, CLIs, CI systems, or remote build tools.
• Experience using or building AI agents for platform maintenance and production operations.
• Familiarity with mobile development workflows using simulators and emulators.

Why Limrun

As an early engineer, you'll work directly with the founders to decide what Limrun builds next. Each new service can open up more development workflows to agents. You'll help establish how the engineering team works, with room to take on broader responsibilities as the company grows.

We're a small team using AI to tackle big infrastructure problems. Limrun gives you room to explore how far AI can take your engineering work, test ideas, and see the results in systems that developers and agents depend on.

What we offer

• A remote-friendly environment with flexible working hours.
• A San Francisco office in South Park, with relocation support.
• Regular team offsites.
• Company-covered work travel across Europe and the US.

Apply

Send us your resume or profile and a short note about an engineering project you're proud of. This could be:

• A system or service you took from design to production.
• A scaling, performance, or reliability problem you solved.
• An operational improvement that reduced failures, manual work, or recovery time.

Tell us what you personally owned, the decisions you made, and what improved. If AI contributed, tell us how you used it and checked the results. Links to code, projects, or technical writing are welcome where you can share them.


Similar Jobs

Explore other opportunities that match your interests

Senior Machine Learning Infrastructure Engineer (AI Validation Platform)

Devops
12h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

General Motors

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Storm4

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

preference model

United State

Subscribe our newsletter

New Things Will Always Update Regularly