G

Founding Engineer, Agent Harness for Production AI Systems

goaly ai • United State
Visa Sponsorship
Apply Now
AI Summary

Build the agent harness layer between frontier models and real users, including loops, tools, context, sandboxing, and evals. Ship and operate production agent workflows that run at scale, with full observability, replayability, and reliable execution. Own end-to-end problem solving and evaluation, requiring strong production software experience and core agent/tooling and debugging skills.

Key Highlights
Founding role building the agent harness layer for real customer usage
Own the full lifecycle from prototype to production and improve environments/evals/models from production traces
Design sandboxed, reproducible execution and an eval stack that detects reward hacking and silent failures
Key Responsibilities
Own the agent harness components including the agent loop, tool interface, system prompts, permissions, stop conditions, retries, budget caps, and output checks.
Build sandboxed execution environments (containers/microVMs) that are isolated, reproducible, and fast enough for thousands of parallel agent rollouts for serving and RL/evals.
Design context engineering strategies such as windowing, compaction, file-backed memory/state, retrieval, and sub-agent hand-offs for long-running tasks.
Make agent runs observable and replayable by tracing model calls, tool calls, and state changes for reproduction of failures.
Build the evaluation stack using offline suites, tests on production traces, LLM-as-judge, and checks for reward hacking and silent failures.
Ship the system with customers by embedding with real users, taking workflows from prototype to stable production, and turning patterns into reusable building blocks.
Close the loop with research by converting production traces and failures into tasks, environments, and reward signals for post-training.
Own model call routing and manage cost, latency, and quality trade-offs across where different models run.
Technical Skills Required
Python Docker Linux
Benefits & Perks
Hybrid in Palo Alto: 4+ days/week in office
Visa sponsorship: H-1B and OPT/CPT support available, with immigration counsel
Complimentary lunch, dinner, snacks, and drinks
Nice to Have
Experience building agent platforms, workflow systems, model-serving products, ML infrastructure, or other developer-facing AI systems.
Experience with multi-tenant control planes, identity and access management, metering or billing, enterprise security, or high-throughput APIs.
Familiarity with LLM inference, fine-tuning, evaluation, tool use, sandboxes, or the operational lifecycle of production agents.
Experience owning open-source SDKs, developer tools, API documentation, examples, or community-facing integrations.
Full-stack ability and an eye for interaction design for product problems beyond the backend.

Job Description


About Us

We’re building toward a world where every company can become its own AI lab.

Goaly is a stealth AI startup founded by ex-Meta Superintelligence Labs engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last.

Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.

About The Role

You will be one of the first engineers building the layer between frontier models and real users: the agent harness. That means the loop, the tools, the context, the sandbox the agent runs in, and the evals that tell us whether it actually works. You will ship this into production for real customers, then turn what we learn from those runs into better environments, better evals, and better models.

This is a founding role. There is no playbook yet; you help write it. You will own problems end to end, from the first prototype with a customer to a system that runs thousands of agent rollouts a day without anyone watching it.

What You'll Do

  • Own the harness: the agent loop, tool interface, system prompts, permissions, stop conditions, retries, budget caps, and output checks that turn a model into a dependable agent.
  • Build sandboxed execution environments (containers / microVMs) that are isolated, reproducible, and fast enough to run thousands of rollouts in parallel, used both for serving and for RL and evals.
  • Design context engineering: what goes into the window, when to compact, file-backed memory and state, retrieval, and hand-offs between sub-agents on long-running tasks.
  • Make every agent run observable and replayable: trace every model call, tool call, and state change, so a failure at 3am can be reproduced at 9am.
  • Build the eval stack: offline suites, tests on production traces, LLM-as-judge, and checks that catch reward hacking and silent failures before they poison training data.
  • Ship with customers: embed with real users, take a workflow from prototype to stable production, and turn repeated patterns into reusable building blocks (tools, MCP servers, agent skills, playbooks).
  • Close the loop with research: convert production traces and failures into tasks, environments, and reward signals for post-training.
  • Make the model calls: pick which model runs where, and own the cost, latency, and quality trade-offs.

You may be a good fit if you have

  • 5+ years building production software, with real experience running systems at scale (many users, high request volume, on-call scars).
  • Built and shipped an agentic system with tool use that real users depend on, and debugged it when it broke.
  • First-hand experience with the failure modes: context overflow, agents looping, tool misuse, prompt injection, runaway cost.
  • Strong Python plus at least one of TypeScript, Go, or Rust; comfortable with Docker, Linux, and cloud infrastructure.
  • A habit of building your own evals instead of trusting public benchmarks; you measure before you claim.
  • A strong pull toward the simplest system that works, and a low tolerance for complexity that does not earn its keep.
  • Comfort operating without a playbook, in ambiguity, and being wrong quickly in front of customers.
  • The ability to explain trade-offs clearly to engineers, researchers, and customers.

Strong pluses

  • Experience building agent platforms, workflow systems, model-serving products, ML infrastructure, or other developer-facing AI systems.
  • Experience with multi-tenant control planes, identity and access management, metering or billing, enterprise security, or high-throughput APIs.
  • Familiarity with LLM inference, fine-tuning, evaluation, tool use, sandboxes, or the operational lifecycle of production agents.
  • Experience owning open-source SDKs, developer tools, API documentation, examples, or community-facing integrations.
  • Full-stack ability and an eye for interaction design when a product problem requires work beyond the backend.

How we work

  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.
  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.
  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.
  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.
  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.
  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location, visa sponsorship & benefits

  • Hybrid in Palo Alto: 4+ days/week in office.
  • Visa sponsorship: H-1B and OPT/CPT support available, with immigration counsel.
  • Meals & perks: Complimentary lunch, dinner, snacks, and drinks.

A note on qualifications. We value exceptional ability over perfect keyword matches. If the work excites you and you can show strong technical ability, learning speed, or ownership, we encourage you to apply.

Equal opportunity

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

BNSF Railway

United State

Senior Staff Full-stack Engineer

Programming
•
1m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Capital One

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

goaly ai

United State

Subscribe our newsletter

New Things Will Always Update Regularly