E

Member of Technical Staff - AI Evaluation & Benchmark Research

employia • United State
Visa Sponsorship
Apply
AI Summary

Lead research in LLM evaluations, benchmark design, and data quality for a stealth AI infrastructure startup. Design novel evaluation methodologies, build datasets for frontier AI models, and publish impactful research. Requires 6–10 years of AI/ML research experience, a PhD in a quantitative field, and published work at premier conferences.

Key Highlights
Early dedicated research hire with significant influence over company's research direction and evaluation strategy
Lead research that directly influences how frontier AI models are evaluated and improved
Own benchmark design, evaluation methodology, and dataset strategy from the ground up
Key Responsibilities
Design novel evaluation methodologies for LLM evaluations
Build datasets that improve frontier AI models
Conduct publishable research at premier conferences
Collaborate with domain experts to define the next generation of AI evaluation infrastructure
Own benchmark design, evaluation methodology, and dataset strategy
Technical Skills Required
Python PyTorch Machine Learning
Benefits & Perks
$250K–$350K base plus equity
Visa sponsorship available
Hybrid work model (4–5 days in office)
Nice to Have
PhD in Machine Learning, Computer Science, Computer Vision, Statistics, or another quantitative discipline from a top-tier university
Published research at NeurIPS, ICML, ACL, or equivalent as a lead or first author
Previous experience at a frontier AI lab or AI infrastructure company

Job Description


About the Company

We're partnering with a stealth AI infrastructure startup building the data foundation for the next generation of frontier AI models. Rather than training models directly, the company develops the benchmarks, evaluation frameworks, and high-quality datasets that enable leading AI labs to build more capable, reliable, and production-ready systems. Working across domains such as computer use, finance, multimodal generation, and enterprise productivity, the team is redefining how AI models are evaluated and improved through world-class research and data infrastructure.


Company Snapshot

  • $30M funded stealth AI infrastructure startup founded by leaders from a frontier AI lab and an AI infrastructure decacorn.
  • Team includes researchers and engineers from Meta, Scale AI, Anthropic, and other leading AI organizations.
  • Already partnering with top frontier AI labs across computer use, finance, multimodal AI, and enterprise productivity.
  • One of the earliest dedicated research hires with significant influence over the company's research direction and evaluation strategy.
  • Mission-driven team building the data infrastructure that enables frontier AI models to perform valuable real-world work.
  • Competitive compensation of $250K–$350K base plus equity, with visa sponsorship available.


Why You Should Join

  • Lead research that directly influences how frontier AI models are evaluated and improved.
  • Own benchmark design, evaluation methodology, and dataset strategy from the ground up.
  • Publish impactful research while collaborating with some of the world's leading AI organizations.
  • Join an exceptional team of researchers with deep experience across frontier AI labs and infrastructure companies.
  • Shape the future of AI data quality and evaluation at one of the industry's most promising stealth startups.


About the Role

You'll join as a Member of Technical Staff leading research across LLM evaluations, benchmark design, and data quality. You'll design novel evaluation methodologies, build datasets that improve frontier models, conduct publishable research, and work closely with domain experts to define the next generation of AI evaluation infrastructure.


Qualifications

  • 6–10 years of combined academic and industry AI/ML research experience.
  • PhD in Machine Learning, Computer Science, Computer Vision, Statistics, or another quantitative discipline from a top-tier university (strongly preferred). Exceptional Master's candidates with outstanding publications will also be considered.
  • Published research at premier conferences such as NeurIPS, ICML, ACL, or equivalent, ideally as a lead or first author.
  • Deep expertise in LLM evaluation, benchmark design, dataset creation, and AI data quality.
  • Strong engineering skills with Python and PyTorch, capable of designing experiments and collaborating on evaluation infrastructure.
  • Previous experience at a frontier AI lab or AI infrastructure company is highly preferred.


Pay range and compensation package

  • Compensation: $250K–$350K Base + Competitive Equity
  • Location: New York City or San Francisco
  • Work Model: Hybrid (4–5 days in office)
  • Visa Sponsorship: Available, including visa transfer

Similar Jobs

Explore other opportunities that match your interests

Stamping Vision Engineer

Programming
•
23m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

General Motors

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

BeaconFire Inc.

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

rs global services

United State

Subscribe our newsletter

New Things Will Always Update Regularly