Q

Senior AI Engineer, GTM Automation

qwntl labs San Francisco Bay Area
Visa Sponsorship
Apply
AI Summary

Build and operate production LLM agents for end-to-end GTM workflows including sourcing, research, and CRM management. Design evaluation frameworks, human-in-the-loop gates, and instrumentation to ensure agent reliability and alignment. Requires strong Python/TypeScript skills, experience with LLM systems and MCP, and the ability to ship without detailed specs.

Key Highlights
Compensation of $300k plus equity
3/3 hybrid work model in San Francisco
Visa sponsorship available
Key Responsibilities
Build and run agents that own whole GTM workflows end to end: sourcing, person-level research, drafting, enrichment, CRM updates, reporting
Design the human gates: approval steps, hand-offs, escalation, and the short list of things an agent never does on its own
Write the evals including transcript review, scoring rubrics, and regression sets
Instrument every model and tool call and tie it to pipeline metrics
Ship MCP servers, agent skills, and small internal web apps against CRM, enrichment, analytics, and social data
Set the conventions for a shared codebase used by co-founders and leads
Sit in on sales calls and read the outbound to ensure systems match actual work
Track model releases and agent frameworks to determine future build priorities
Technical Skills Required
Python TypeScript SQL
Benefits & Perks
$300k + equity
Visa sponsorship
Nice to Have
2+ years as a software, ML, or forward-deployed engineer
Serious time in Claude Code or the Claude Agent SDK
Worked on or next to a revenue team with CRM, enrichment, or outbound tooling
Sold to or built for developers
An applied ML or experimentation background
A habit of writing clearly for non-technical stakeholders

Job Description


About QWNTL Labs

QWNTL Labs is a venture-backed research lab in London/San Francisco working on the frontier of agent performance and alignment

Our view is that the models are already smart enough. What keeps agents stuck in demos is trust over time. An agent can ace a task on day one and quietly stop following its rules by week three, and almost nobody is measuring that. The difference between AI that helps people work and fully autonomous workers and companies is whether an agent can be trusted across the horizon a real job runs on. Closing that gap is what we exist to do.


About the role

You will work directly with the co-founder who owns growth and with the sales and marketing leads, building the agents, evals, and tooling that run sourcing, research, outbound, content, and pipeline for a lab that sells to engineers.


-> Application Process: <-

First Gate - Submit application to https://qwntl.com/early-access

Second Gate - If accepted for access to harness, you will be given a tool which is not publicly available and is used internally by all QWNTL team members to build and manage fully autonomous workers. The harness enables 30-day horizon agents that achieve 98% reliability

Third Gate - Submit what you've built when given tools that surpass what is commercially available to the public. Submission can be done through the team discord which will be shared when first gate is passed.


We hire on the work, not the resume.

If you demonstrate your ability to outperform, you will be given the opportunity to continue doing so.


What you will do

  • Build and run agents that own whole GTM workflows end to end: sourcing, person-level research, drafting, enrichment, CRM updates, reporting
  • Design the human gates: approval steps, hand-offs, escalation, and the short list of things an agent never does on its own (send, post, promise)
  • Write the evals. Transcript review, scoring rubrics, regression sets. Know when an agent is ready for production and know when it has drifted
  • Instrument every model and tool call and tie it to pipeline: replies, meetings, design-partner conversations
  • Ship MCP servers, agent skills, and small internal web apps against our CRM, enrichment, analytics, and social data
  • Set the conventions for a shared codebase that a co-founder, a sales lead, and a marketing lead will all commit to
  • Sit in on sales calls and read the outbound so the systems match how the work is actually done
  • Track model releases and agent frameworks, and know which ones change what we should build next


You might be a fit if you

  • Write production Python or TypeScript and have shipped software other people depend on
  • Have built LLM systems that run in production: context engineering, agents, tool use, MCP, evals
  • Have used evals and transcript analysis to make an LLM system better, and can show the before and after
  • Are fluent in SQL and comfortable in messy data
  • Ship without a spec. You will get a goal and a co-founder's time, not a ticket queue
  • Would rather build one system that works for years than ten demos


Strong candidates may also have

  • 2+ years as a software, ML, or forward-deployed engineer
  • Serious time in Claude Code or the Claude Agent SDK, or the equivalent elsewhere
  • Worked on or next to a revenue team: CRM, enrichment, outbound tooling, attribution
  • Sold to or built for developers, and know why engineers ignore most outreach
  • An applied ML or experimentation background
  • A habit of writing clearly for people who do not read code


Logistics

  • Location: San Francisco, 3/3 hybrid
  • Compensation: $300k + equity
  • Visa sponsorship: yes
  • Start: as soon as we find the right person



Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

RemoteStar

San Francisco Bay Area
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

inventure

San Francisco Bay Area
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

a4assist

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly