C

Senior AI/Software Engineer (Voice AI & Agent Reliability) - San Francisco

cekura United State
Visa Sponsorship Relocation
Apply
AI Summary

Build core voice AI simulation engines, LLM-powered evaluation systems, and self-improving agent pipelines at a high-growth startup. Own real-time voice infrastructure, adversarial testing, and end-to-end observability for enterprise-grade voice agents. Requires strong Python, distributed systems, and LLM evaluation expertise with a generalist mindset.

Key Highlights
Build simulation engines for voice, chat, and phone agent testing with realistic personas, noise, and edge cases
Develop LLM-powered evaluators and autonomous self-improvement loops for agent reliability
Own real-time voice infrastructure (SIP, WebRTC, Twilio, STT/TTS) with focus on latency, audio quality, and observability
Key Responsibilities
Design and ship simulation engines for testing voice agents across voice, chat, and phone with realistic scenarios
Build LLM-powered evaluators, metrics, and closed-loop self-improvement systems for agent reliability
Develop adversarial red-teaming for jailbreaks, PII leaks, and off-script behavior detection
Implement real-time voice infrastructure (SIP, WebRTC, STT/TTS) with low-latency and high audio quality
Own end-to-end observability pipelines for production monitoring and live drift detection
Shape architectural decisions and technical standards for a small, senior engineering team
Technical Skills Required
Python Distributed Systems LLM Evaluation & Agentic Systems
Benefits & Perks
Top-of-market cash compensation
Meaningful equity with liquidity opportunities
Zero commute tax (daily lunches/dinners, office support)
Nice to Have
Experience with Twilio, SIP, WebRTC, Vapi, Retell, LiveKit, or Pipecat
Hands-on LLM evaluation frameworks or AI observability tools
Early-stage startup or AI/voice AI industry experience
Open-source contributions or public technical writing

Job Description


Read This First

We work with unusual intensity. In-person in San Francisco, six days a week, long days, most weekends. This is not a phase we'll grow out of. It's how we've chosen to build, because we're in a market where speed decides who wins.

We're telling you this in the first paragraph, not the last, because we only want people who read that and feel pulled in, not talked into it. If you want a 9-to-5 (genuinely, no judgment), this isn't your role, and we'd rather you know now.

Here's What You Get In Exchange

  • Top-of-market cash. We don't pay average salaries and ask for extraordinary hours. The comp reflects the commitment.
  • Meaningful equity you can believe in. Significant grants, employee-friendly terms, and our intent to create liquidity opportunities as we raise.
  • Zero commute tax. We support you living close to the office, with dinner at the office every night. Your hours go into building, not commuting.
  • Founders in the trenches. We work the same schedule we ask of you. This is a shared war, not extraction.
  • Compression of a decade into two years. You'll ship more, own more, and grow faster here than anywhere paying you to coast.

About Cekura

Cekura (YC F24) is building the voice AI engineer. Teams use Cekura to test agents before going live, monitor real production calls, and self-improve continuously. Cekura doesn't just flag issues and suggest fixes: it reproduces failures in simulation, fixes them, tests the fix thoroughly, and raises PRs. The platform spans pre-production simulation, LLM-powered evaluation, adversarial red-teaming, production monitoring with live drift detection, and cross-provider benchmarking (Vapi, Retell, Pipecat, LiveKit, ElevenLabs, and more).

We're trusted where reliability is non-negotiable: customers include Five9, HighLevel, Twin Health, PwC, Deloitte, Jobber, and Jotform, with HIPAA, SOC 2, and GDPR as defaults, not checkboxes. We're growing fast and backed by top investors.

About The Role

You'll build the core of Cekura: the simulation engines, evaluation systems, self-improvement loops, and observability pipelines our customers rely on to ship voice agents with confidence.

We deliberately don't split this into "software engineer" vs. "AI engineer." The interesting problems live at the boundary: real-time voice infrastructure meets LLM-as-judge evaluation, distributed systems meet RL-style self-improvement loops, telephony meets audio and speech analysis (ASR quality, barge-in, latency, prosody), and classic NLP meets frontier agentic behavior. You'll work across that whole surface.

What You'll Do

Build the testing and simulation engine. Design and ship systems that simulate thousands of realistic conversations against customer agents across voice, chat, and phone, with control over personas, interruptions, background noise, and edge cases.

Push the frontier of agent evaluation. Build LLM-powered evaluators, metrics, and the closed self-improvement loop at the heart of the product: detect failure, reproduce in simulation, generate the fix, test it thoroughly, and raise the PR, all autonomously. This includes adversarial red-teaming for jailbreaks, PII leaks, and off-script behavior, plus production monitoring with live drift detection.

Do applied audio and speech research. Go beyond the transcript. Separate background noise from actual speech, and map the paralinguistic layer of conversation (emotion, tone, silences, hesitations, speaking rate, overlaps) into structured signals, so agents are evaluated on how something was said, not just what.

Own real-time voice infrastructure. SIP, WebRTC, WebSockets, STT/TTS pipelines, and providers like Twilio, Vapi, Retell, LiveKit, and Pipecat. Latency, barge-in, and audio quality are first-class problems here.

Ship end-to-end. Take features from design to production. You own the full stack of what you build: backend, infra, evals, and the product surface customers touch.

Shape how we build. We're a small, senior team. Your architectural decisions, code standards, and technical taste will compound as the team grows.

About You

  • A strong generalist engineer excited to work across systems and AI, not someone looking to stay in one lane.
  • You've built and operated production systems and care about reliability, latency, and correctness.
  • Fluent in Python and comfortable picking up whatever the problem needs.
  • Strong instincts about LLMs: where they fail, how to evaluate them, how to build reliable systems on unreliable models.
  • You like ambiguity, move fast, and raise the bar for the people around you.
  • You've read the first section of this JD twice and you're still here.
  • Bonus: you're an ex-founder or aspire to be one.

Minimum Qualifications

  • Strong coding ability in Python, TypeScript, or Go.
  • Experience with at least one of: distributed systems, real-time infrastructure, or LLM-based products.
  • Strong written and verbal communication.

Nice to Have

  • Hands-on experience with LLM evals, agent frameworks, or AI observability.
  • Voice AI experience: Twilio, SIP, WebRTC, Vapi, Retell, LiveKit, Pipecat, or STT/TTS pipelines.
  • Early-stage startup experience, or early engineer at a dev-tool, infra, or AI company.
  • Open-source contributions or public technical work.

This Is Not for You If

  • You want a narrowly scoped role or a fixed tech stack.
  • You prefer working only on models, or only on infrastructure, never both.
  • You need rigid processes or heavy structure.
  • You're optimizing for work-life balance right now. (Later in your career, maybe. Here, now, no.)
  • You don't want to work in-person in San Francisco.

Why Cekura

  • The hardest problems in AI agent reliability: simulation, voice evals, real-time voice at scale.
  • A 12-person engineering team post-seed: early enough to matter, funded enough to move fast.
  • Work directly with founders and a highly technical team.
  • Meaningful equity, top-of-market compensation, fast growth.
  • US visa sponsorship available.
  • Medical, dental, vision, daily lunches and dinners, relocation support to live near the office.

Similar Jobs

Explore other opportunities that match your interests

Fuel Systems Test Engineering Manager

Programming
6h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Caterpillar Inc.

United State

Applied AI Field Researcher

Programming
10h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

anthropic

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

torentify

United State

Subscribe our newsletter

New Things Will Always Update Regularly