V

Senior Applied ML Research Engineer

verona • United State
Relocation
Apply
AI Summary

Develop rigorous evaluation methods and infrastructure to measure AI system reliability, safety, and business impact. Build datasets, simulations, and automated judges that capture real-world workflow quality. Partner with production teams to translate research insights into deployed systems that continuously improve.

Key Highlights
Bridge research and production engineering for continuous AI system improvement
Build evaluation infrastructure for task success, answer quality, robustness, safety, and business outcomes
Analyze production failures to form testable hypotheses and implement effective interventions
Key Responsibilities
Develop evaluations for task success, answer quality, robustness, safety, cost, and business outcomes
Build high-signal datasets, simulations, adversarial tests, and human-review protocols
Establish evaluation integrity by validating automated judges and monitoring drift
Investigate production traces and unexpected failures to explain root causes
Partner with forward deployed engineers to implement research ideas in customer systems
Turn field results into reusable methods, internal standards, and product capabilities
Technical Skills Required
Python Statistics Machine Learning
Benefits & Perks
Competitive compensation and equity
Generous health benefits
Unlimited PTO
Paid parental leave
Daily lunches and dinners
Transportation and relocation support
Retirement plans
Nice to Have
Experience with LLM or agent evaluation, post-training, red teaming, synthetic data, reward modeling, or human-feedback systems
Experience designing evaluations for open-ended workflows with subjective quality
Systems-level understanding of models, retrieval, tools, prompts, data, and application code in production
Experience working with messy domain data or subject-matter experts to turn tacit judgment into reliable measurement

Job Description


About The Role

Applied research at Verona begins with a demanding real-world question: how do we know an AI system is useful, reliable, and improving on the workflow that matters? You will develop the evaluation methods, datasets, experiments, and learning systems that make those answers rigorous.

This role sits between research and production engineering. You will study failures from live enterprise deployments, form precise hypotheses about model and system behavior, build the infrastructure needed to test them, and turn the results into systems that improve continuously in the field.

What you’ll do

  • Develop evaluations for task success, answer quality, robustness, safety, cost, and the business outcome a system is intended to improve.
  • Build high-signal datasets, simulations, adversarial tests, and human-review protocols that capture what actually matters in a customer workflow.
  • Establish evaluation integrity by validating automated judges, monitoring drift, and making sure reported improvements represent real progress.
  • Investigate production traces and unexpected failures deeply enough to explain why they happened and which intervention is most likely to work.
  • Partner with forward deployed engineers to move research ideas into customer systems and measure their impact under real operating conditions.
  • Turn field results into reusable methods, internal standards, research infrastructure, and product capabilities across Verona.

What we’re looking for

  • A record of strong applied machine-learning research or engineering work, demonstrated through shipped systems, publications, open-source contributions, or equivalent projects.
  • Excellent experimental judgment: you can turn surprising system behavior into testable hypotheses and distinguish a meaningful result from a misleading metric.
  • Fluency in Python and the engineering ability to build durable research and evaluation infrastructure, not only one-off notebooks.
  • Strong foundations in statistics, machine learning, and the practical limitations of automated evaluation.
  • Clear technical communication and the ability to collaborate with researchers, production engineers, customer teams, and domain experts.
  • A bias toward research whose value can be observed in deployed systems and real user outcomes.

You might excel here if

  • Experience with LLM or agent evaluation, post-training, red teaming, synthetic data, reward modeling, or human-feedback systems.
  • Experience designing evaluations for open-ended workflows where quality is subjective, multidimensional, or difficult to observe directly.
  • Systems-level understanding of how models, retrieval, tools, prompts, data, and application code interact in production.
  • Experience working with messy domain data or subject-matter experts to turn tacit judgment into reliable measurement.

Benefits

  • Competitive compensation and equity
  • Generous health benefits
  • Unlimited PTO
  • Paid parental leave
  • Daily lunches and dinners
  • Transportation and relocation support
  • Retirement plans

Similar Jobs

Explore other opportunities that match your interests

Systems Engineer II - EO/IR Signal Processing

Programming
•
2h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Raytheon

United State

Flight Software Engineer

Programming
•
2h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

round-peg solutions (rps)

United State

Software Integration Engineer

Programming
•
2h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Jobot

United State

Subscribe our newsletter

New Things Will Always Update Regularly