Own generation quality, cost, and prompt architecture for an AI workforce transition platform. Build and maintain evaluation suites, manage model strategy on AWS Bedrock, and optimize RAG integration. Requires deep expertise in LLM prompting, systematic evaluation, and Python or PHP coding.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
About Cortextual
Cortextual is a workforce transition platform built for companies restructuring around AI. It gives founders, COOs, and Heads of People a live view of org structure, workforce costs, and scenario planning - so they can make workforce decisions with real data before committing to them.
Most companies are navigating some version of the same challenge: teams are changing, AI is being adopted, roles are shifting, and the cost of getting it wrong is higher than ever. But most leaders are making these decisions without seeing the full picture - what each role actually costs, where dependencies exist, and what could happen if someone leaves.
Cortextual brings org structure, workforce costs, and scenario modeling into one live view - with visibility into dependencies and knowledge risk when you need it. Not a static diagram someone updated last quarter. A living view of how your organization works and what a change could cost before you make it.
The Role
We are looking for a mid to senior AI Engineer to own generation quality and cost — the prompts, model choices, and evaluations that turn retrieved context into correct, grounded, well-formatted answers — along with the AI chat product logic around them. You will join Cortextual's existing engineering team of five and report directly to the Head of Development, working closely alongside the Retrieval (RAG) Engineer on the retrieval-to-generation boundary.
What You Will Own
- The prompt architecture: intent classification and query reformulation, answer synthesis over retrieved context, and the conversational and account data paths.
- Model strategy on AWS Bedrock (currently Nova Pro; evaluating newer models including Claude): prompt caching (static versus dynamic context split), token cost, and latency.
- The generation quality evaluation suite (built jointly with the Retrieval (RAG) Engineer): hallucination and groundedness scoring, answer format adherence, intent routing accuracy, and A/B testing of prompt variants.
- AI chat product behavior: streaming responses, follow-up handling, and short-circuiting wasteful "no answer found" model calls.
Interested in remote work opportunities in Development & Programming? Discover Development & Programming Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
Key Responsibilities
- Reduce hallucination and improve citation and source attribution reliability.
- Resolve relative dates and ordering intent in the classification prompt; improve query reformulation for follow-up questions.
- Own model migrations and cost and latency tradeoffs; keep prompt cache points correct.
- Partner with the Retrieval (RAG) Engineer on the retrieval and prompt boundary (top K, context packing).
- Collaborate closely with the Head of Development and the wider Cortextual engineering team of five as the platform scales toward full autonomy.
What You Bring
Must Have
- Demonstrated depth in LLM prompting and systematic evaluation — building measurable eval harnesses, not just tweaking prompts.
- Solid coding (Python and/or PHP) to integrate prompts, run evaluations, and read the codebase.
- Strong grasp of RAG, tokenization, and context window and caching economics.
Nice to Have
- Bedrock, Nova, or Claude experience.
- Streaming (SSE or WebSocket) generation; LLM-as-judge evaluation design.
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Mindset
- Builder - energised by ambiguity and comfortable owning a system end to end.
- Rigorous - trusts evaluation data over intuition and builds the harness to prove it.
- Collaborative - works closely with a small, senior engineering team without ego.
- Curious about AI - genuinely excited by retrieval, search, and the knowledge intelligence space.
What We Offer
- Competitive salary and performance-based incentives.
- Fully remote-first environment with flexible working.
- A founding-stage role on a five-person engineering team, with direct influence over Cortextual's retrieval architecture.
- Fast-moving, international team with a strong culture of ownership.
- Learning & development budget.
Join our team at Cortextual, a fast-growing global start-up. We're looking for driven, passionate individuals to join us as we continue to expand our reach and impact!
Similar Jobs
Explore other opportunities that match your interests