D

Principal AI Evaluation Engineer

Relocation
Apply
AI Summary

Seeking a Principal Evals Engineer for a major AI program in Abu Dhabi to define and build AI evaluation standards, frameworks, and infrastructure. This hands-on role involves ensuring AI systems are production-ready by designing shared evaluation architectures and automating quality measurement. Requires deep experience in modern AI evaluation, LLM quality assessment, and building robust, scalable evaluation systems.

Key Highlights
Define and own AI evaluation strategy, standards, frameworks, and infrastructure.
Design and build shared evaluation architecture for multiple AI products and engineering teams.
Ensure AI systems are production-ready through rigorous, evidence-based evaluation.
Key Responsibilities
Define the overall AI evaluation strategy and reference architecture across the engineering organization
Build shared evaluation frameworks, harnesses and reusable libraries
Develop and maintain golden datasets and dataset versioning strategies
Design automated evaluation and grading systems
Implement regression detection for non deterministic AI behaviour
Evaluate LLM output quality, hallucination, grounding, faithfulness and citation accuracy
Build evaluation frameworks for RAG and retrieval systems
Evaluate agentic systems, including tool usage, multi step reasoning and failure recovery
Design and calibrate LLM as a judge systems against human evaluation
Establish production monitoring, online evaluation and drift detection
Develop adversarial testing methodologies covering prompt injection, jailbreaks and data leakage
Integrate AI evaluation into CI/CD pipelines and release processes
Help engineering teams understand and apply rigorous evaluation standards independently
Contribute to building the broader evaluation engineering function as the organisation scales
Technical Skills Required
Python LLM Evaluation Generative AI Evaluation
Benefits & Perks
Relocation support available
Competitive tax-free UAE package
Nice to Have
Ragas
DeepEval
Promptfoo
Braintrust
LangSmith
Langfuse
LangChain or LangGraph
Vector databases such as pgvector or Qdrant
AI red teaming or adversarial testing
Voice or conversational AI evaluation
Human evaluation programme design
Data pipeline or performance testing
Arabic language AI evaluation, including golden datasets, dialect coverage and right to left validation
Experience working within regulated, government or highly secure environments

Job Description


Principal Evals Engineer

Location: Abu Dhabi, UAE (relocation support available)


About the opportunity

We are hiring a Principal Evals Engineer to join a major new AI programme in Abu Dhabi. The organization is building a large scale AI platform designed to support next generation AI systems across multiple sectors.The work spans AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence, with a strong focus on building reliable AI systems that can operate at scale in real world environments.


This is a highly strategic and hands on role for someone who has deep experience evaluating modern AI systems and wants to help define how AI quality is measured across an entire engineering organization. You will own the evaluation standards, frameworks and infrastructure used to determine whether AI systems are ready to move into production.


The role

You will design and build the shared evaluation architecture used across multiple AI products and engineering teams. This includes creating robust evaluation harnesses, managing golden datasets, automating grading, detecting behavioural regressions and ensuring that engineering teams can make release decisions based on evidence rather than subjective judgement.

This is a senior individual contributor role. You will remain highly hands on, building frameworks, running experiments and setting the technical standard for AI evaluation across the organization.


Responsibilities:

• Define the overall AI evaluation strategy and reference architecture across the engineering organization

• Build shared evaluation frameworks, harnesses and reusable libraries

• Develop and maintain golden datasets and dataset versioning strategies

• Design automated evaluation and grading systems

• Implement regression detection for non deterministic AI behaviour

• Evaluate LLM output quality, hallucination, grounding, faithfulness and citation accuracy

• Build evaluation frameworks for RAG and retrieval systems

• Evaluate agentic systems, including tool usage, multi step reasoning and failure recovery

• Design and calibrate LLM as a judge systems against human evaluation

• Establish production monitoring, online evaluation and drift detection

• Develop adversarial testing methodologies covering prompt injection, jailbreaks and data leakage

• Integrate AI evaluation into CI/CD pipelines and release processes

• Help engineering teams understand and apply rigorous evaluation standards independently

• Contribute to building the broader evaluation engineering function as the organisation scales


Key requirements:

• Staff or Principal level engineering experience

• Strong hands on Python engineering skills

• Deep experience with LLM or Generative AI evaluation

• Experience evaluating RAG, retrieval or agentic systems

• Strong understanding of hallucination, grounding and output quality measurement

• Experience building regression detection for non deterministic AI systems

• Practical experience with LLM as a judge techniques

• Strong understanding of rubric design and calibration against human labels

• Statistical literacy including sampling, confidence intervals, significance and inter rater agreement

• Strong CI/CD and software quality engineering experience

• Experience building reusable frameworks or shared engineering libraries

• Hands on experience working with production AI systems


Highly desirable

Experience with one or more of the following would be highly valuable:

• Ragas

• DeepEval

• Promptfoo

• Braintrust

• LangSmith

• Langfuse

• LangChain or LangGraph

• Vector databases such as pgvector or Qdrant

• AI red teaming or adversarial testing

• Voice or conversational AI evaluation

• Human evaluation programme design

• Data pipeline or performance testing

• Arabic language AI evaluation, including golden datasets, dialect coverage and right to left validation

• Experience working within regulated, government or highly secure environments


Technology environment

The primary language is Python, with additional use of TypeScript and Java.

The wider environment includes custom evaluation harnesses, LLM as a judge patterns, golden datasets, Langfuse, Ragas, DeepEval, Promptfoo, LangChain, LangGraph, Microsoft Agent Framework, pgvector, Qdrant, Playwright, Cypress, Great Expectations, Grafana, Docker, Kubernetes and Azure. The team is pragmatic about tooling and is more interested in strong engineering judgement than experience with one specific framework.


Why join

This is an opportunity to help shape the evaluation discipline within one of the most ambitious AI programmes currently being built in the UAE.


You will work alongside experienced AI, platform and product engineers in small senior teams, with significant technical ownership and direct influence over how production AI systems are approved and released.


The organization also provides a competitive tax free UAE package and comprehensive relocation support for candidates and their families.


If you are a senior AI engineer who enjoys building the systems that determine whether AI truly works in production, this role offers the opportunity to define that standard from the ground up.



Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

hyre

Emea
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

proghres

Emea
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Director

rgg capital

Emea

Subscribe our newsletter

New Things Will Always Update Regularly