Join Duxre as an AI Evaluation and Quality Assurance Engineer to ensure the highest standards of performance, accuracy, and reliability for Dash's RAG-based AI agent. Design and implement automated QA pipelines, develop metrics, and collaborate with engineers and data scientists. Contribute to continuous improvement of QA processes and retrieval versioning.
Key Highlights
Technical Skills Required
Benefits & Perks
Job Description
About Duxre / Dash:
Duxre is a funded startup that is redefining how commercial real estate professionals interact with data through Dash, our intelligent AI brokerage assistant. Dash streamlines workflows, analyzes property data, and enhances decision-making through AI-driven insights. We’re looking for a passionate AI Evaluation & QA Engineer to work with our senior AI team to ensure that Dash’s RAG-based AI agent consistently meets the highest standards of performance, accuracy, and reliability.
Role Overview:
We are seeking an experienced AI Evaluation & QA Engineer (3-5 years) to join our seasoned product and engineering team. This role is responsible for testing, validating, and improving Dash’s retrieval-augmented generation (RAG) engine and AI agent systems. You will design automated and manual evaluation pipelines, develop testing frameworks for retrieval quality and agent reasoning, and collaborate closely with engineers and data scientists to drive continuous improvement in AI quality and performance.
Key Responsibilities:
- Design and implement automated QA pipelines for RAG engine components and AI agent behavior testing.
- Develop metrics and benchmarks for evaluating retrieval accuracy, generation relevance, and agent reliability, including measures such as faithfulness, context precision, retrieval recall, factual consistency, answer completeness, and user intent alignment.
- Utilize evaluation tools such as Ragas, LangSmith, Weights & Biases, or similar frameworks to measure and improve RAG pipeline performance.
- Build custom Python scripts and utilities to test, monitor, and analyze retrieval and response quality at scale.
- Work directly with AI agents and backend engineers to identify edge cases, hallucination issues, and inconsistencies in retrieval and generation flows.
- Create synthetic and real-world test datasets for evaluating document retrieval, context injection, and agent reasoning quality.
- Develop dashboards and reports for performance tracking and regression monitoring.
- Contribute to continuous improvement of QA processes, retrieval versioning, and test automation.
- Embrace the startup environment by demonstrating adaptability, initiative, and a strong work ethic, contributing beyond defined responsibilities when needed to drive product success
Qualifications:
- 3-5 years of professional experience in software QA, AI/ML evaluation, or intelligent system testing.
- Proficiency in Python and scripting for automation and data analysis.
- Hands-on experience with AI evaluation tools (e.g., Ragas, LangChain evaluation tools, Weights & Biases, PromptLayer, etc.).
- Solid understanding of retrieval-augmented generation (RAG) systems, vector databases, and embedding evaluation.
- Experience with modern MLOps tools, CI/CD systems, and version control (GitHub, Jenkins, etc.).
- Strong analytical and problem-solving skills, with a data-driven approach to quality assurance.
- Excellent communication skills and ability to collaborate cross-functionally with AI engineers, data scientists, and product managers.
- Experience in user feedback loop integration and reinforcement learning from human feedback (RLHF).
- US Timezone: You must work in the US timezones.
- English Speaking: You must speak English very well.
- References: You must be able to provide 2-3 references.
Nice to Have:
- Familiarity with retrieval pipelines, document ranking, and semantic search systems.
- Exposure to real estate or SaaS platforms is a plus.
What We Offer:
- The opportunity to shape the future of AI in commercial real estate.
- A fully remote, collaborative, and innovation-driven culture.
- Competitive salary and early-stage equity participation.
- Career growth in AI product quality and applied RAG system evaluation.
- Access to AI evaluation frameworks, research tools, and a supportive environment to experiment and learn.
Additional Job Application Terms:
This job is part of LinkedIn’s Full-Service Hiring beta program. Eligibility is limited to candidates located in and performing services in the United States, excluding those based in Alaska, Hawaii, Nevada, South Carolina, or West Virginia.
We’re committed to making our hiring process as smooth and timely as possible, and we understand that waiting to hear back can add to the anticipation. If you’re a potential fit, our team will reach out within two weeks to progress you to the next stage. If you don’t hear from us in that time, we encourage you to explore other opportunities with our team in the future, and we wish you the very best in your job search.
Similar Jobs
Explore other opportunities that match your interests