Seeking a Principal Evals Engineer for a major AI program in Abu Dhabi to define and build AI evaluation standards, frameworks, and infrastructure. This hands-on role involves ensuring AI systems are production-ready by designing shared evaluation architectures and automating quality measurement. Requires deep experience in modern AI evaluation, LLM quality assessment, and building robust, scalable evaluation systems.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Principal Evals Engineer
Location: Abu Dhabi, UAE (relocation support available)
About the opportunity
We are hiring a Principal Evals Engineer to join a major new AI programme in Abu Dhabi. The organization is building a large scale AI platform designed to support next generation AI systems across multiple sectors.The work spans AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence, with a strong focus on building reliable AI systems that can operate at scale in real world environments.
This is a highly strategic and hands on role for someone who has deep experience evaluating modern AI systems and wants to help define how AI quality is measured across an entire engineering organization. You will own the evaluation standards, frameworks and infrastructure used to determine whether AI systems are ready to move into production.
The role
You will design and build the shared evaluation architecture used across multiple AI products and engineering teams. This includes creating robust evaluation harnesses, managing golden datasets, automating grading, detecting behavioural regressions and ensuring that engineering teams can make release decisions based on evidence rather than subjective judgement.
This is a senior individual contributor role. You will remain highly hands on, building frameworks, running experiments and setting the technical standard for AI evaluation across the organization.
Responsibilities:
• Define the overall AI evaluation strategy and reference architecture across the engineering organization
• Build shared evaluation frameworks, harnesses and reusable libraries
• Develop and maintain golden datasets and dataset versioning strategies
• Design automated evaluation and grading systems
• Implement regression detection for non deterministic AI behaviour
• Evaluate LLM output quality, hallucination, grounding, faithfulness and citation accuracy
• Build evaluation frameworks for RAG and retrieval systems
• Evaluate agentic systems, including tool usage, multi step reasoning and failure recovery
• Design and calibrate LLM as a judge systems against human evaluation
• Establish production monitoring, online evaluation and drift detection
• Develop adversarial testing methodologies covering prompt injection, jailbreaks and data leakage
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
• Integrate AI evaluation into CI/CD pipelines and release processes
• Help engineering teams understand and apply rigorous evaluation standards independently
• Contribute to building the broader evaluation engineering function as the organisation scales
Key requirements:
• Staff or Principal level engineering experience
• Strong hands on Python engineering skills
• Deep experience with LLM or Generative AI evaluation
• Experience evaluating RAG, retrieval or agentic systems
• Strong understanding of hallucination, grounding and output quality measurement
• Experience building regression detection for non deterministic AI systems
• Practical experience with LLM as a judge techniques
• Strong understanding of rubric design and calibration against human labels
• Statistical literacy including sampling, confidence intervals, significance and inter rater agreement
• Strong CI/CD and software quality engineering experience
• Experience building reusable frameworks or shared engineering libraries
• Hands on experience working with production AI systems
Highly desirable
Experience with one or more of the following would be highly valuable:
• Ragas
• DeepEval
• Promptfoo
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
• Braintrust
• LangSmith
• Langfuse
• LangChain or LangGraph
• Vector databases such as pgvector or Qdrant
• AI red teaming or adversarial testing
• Voice or conversational AI evaluation
• Human evaluation programme design
• Data pipeline or performance testing
• Arabic language AI evaluation, including golden datasets, dialect coverage and right to left validation
• Experience working within regulated, government or highly secure environments
Technology environment
The primary language is Python, with additional use of TypeScript and Java.
The wider environment includes custom evaluation harnesses, LLM as a judge patterns, golden datasets, Langfuse, Ragas, DeepEval, Promptfoo, LangChain, LangGraph, Microsoft Agent Framework, pgvector, Qdrant, Playwright, Cypress, Great Expectations, Grafana, Docker, Kubernetes and Azure. The team is pragmatic about tooling and is more interested in strong engineering judgement than experience with one specific framework.
Why join
This is an opportunity to help shape the evaluation discipline within one of the most ambitious AI programmes currently being built in the UAE.
You will work alongside experienced AI, platform and product engineers in small senior teams, with significant technical ownership and direct influence over how production AI systems are approved and released.
The organization also provides a competitive tax free UAE package and comprehensive relocation support for candidates and their families.
If you are a senior AI engineer who enjoys building the systems that determine whether AI truly works in production, this role offers the opportunity to define that standard from the ground up.
Similar Jobs
Explore other opportunities that match your interests