G

Research Engineer, Reinforcement Learning

Goliath Partners • San Francisco Bay Area
Relocation
Apply Now
AI Summary

Own the end-to-end reinforcement learning loop for a healthcare AI startup building frontier reasoning models. Design RL environments, generate training data, and execute post-training experiments to improve model reasoning. Requires 1+ years of technical experience in RL or LM post-training and a strong academic background in a technical field.

Key Highlights
Full-cycle ownership from environment design to trained model deployment
Focus on teaching frontier AI models to reason in high-stakes healthcare domains
High equity potential (up to 1.0%) in a $100M valued Y Combinator backed startup
Key Responsibilities
Own end-to-end model training, from data and environment design through post-training, evaluation, and iteration
Build RL environments from the ground up, including task design, action spaces, tool interfaces, verifiers, and evaluation harnesses
Run post-training experiments on language models and agents using SFT, RLVR, RLHF/RLAIF, and reward modeling
Train multi-step agents to reason over long, complex, unstructured real-world data
Build the evaluation and reward layer to define correct reasoning for frontier AI models
Analyze model failures to improve data, rewards, and feedback signals
Help set the research direction and technical foundations of the company
Technical Skills Required
Reinforcement Learning Language Model Post-training Machine Learning Systems
Benefits & Perks
Up to $350,000 base salary
Up to 1.0% equity
Relocation and joining bonuses
Comprehensive benefits
Unlimited AI tool budgets
Nice to Have
Master's degree or PhD in Machine Learning, Computer Science, or a related field
Publications at top venues such as NeurIPS, ICML, ICLR, or ACL
Experience at a frontier AI lab, elite research group, or top-tier startup
Experience with RL for language models, agentic systems, or large-scale distributed training

Job Description


Our client is a hyper-growth AI startup building the post-training infrastructure for the next generation of frontier models in healthcare. Backed by Y Combinator and executives from Google DeepMind and OpenAI, and now valued at $100M, they are building the foundational data layer and simulation environments where advanced AI models learn to reason in one of the most complex, high-stakes domains in the world. This is a small, elite team where every engineer shapes the product, and they are seeking a powerhouse Research Engineer to own the full reinforcement learning loop from end to end.


This isn't a role where you tune one piece of a pipeline someone else designed. You'll own the entire cycle: designing the environments, generating the data, running the training, building the evaluations, and deciding what comes next based on what the models get wrong. If you've ever wanted to see what happens when one person can move from idea to trained model to real-world

impact without waiting on five other teams, this is it.


Role & Impact

  • Own end-to-end model training, from data and environment design through post-training, evaluation, and iteration
  • Build RL environments from the ground up, including task design, action spaces, tool interfaces, verifiers, and evaluation harnesses
  • Run cutting-edge post-training experiments on language models and agents using techniques like SFT, RLVR, RLHF/RLAIF, and reward modeling
  • Train multi-step agents to reason over long, complex, unstructured real-world data, where getting the answer right matters
  • Build the evaluation and reward layer that defines what "correct reasoning" means for frontier AI models
  • Dig into model failures and close the loop, continuously improving the data, rewards, and feedback signals that shape what these systems learn
  • Help set the research direction and technical foundations of the company as an early, high-ownership member of a small team


Essential Skills

  • 1+ years of highly technical experience in reinforcement learning, agent environments, LM post-training, or related ML systems
  • Hands-on experience running training end to end, not just one stage of the pipeline
  • Exceptional judgment about data quality, including assessing signal fidelity, coverage, and domain relevance for meaningful training tasks
  • A strong ability to build scalable pipelines, debug complex training runs in large ML codebases, and move rapidly from research concepts to working prototypes
  • A degree in Computer Science, Mathematics, Physics, or a related technical field from a top-tier university, or an equally exceptional academic track record
  • Evidence of exceptional ability, such as top honors, a strong research record, or standout achievements in competitive programming, olympiads, or similar
  • A builder's mindset: you want to own the problem, not just a ticket


Nice to have

  • A Master's degree or PhD in Machine Learning, Computer Science, or a related field from a leading research institution
  • Publications at top venues such as NeurIPS, ICML, ICLR, or ACL
  • Experience at a frontier AI lab, elite research group, or top-tier startup
  • Experience with RL for language models, agentic systems, or large-scale distributed training


Why this role

  • Full-cycle ownership: you run the entire loop from environment to trained model, with no handoffs and no silos
  • A small team with outsized impact, where your work directly shapes the core product and the company's trajectory
  • Exceptional talent density: work alongside researchers and engineers from top universities and leading AI labs
  • Backed by Y Combinator and leaders from Google DeepMind and OpenAI, so you'll be building alongside people who have shaped frontier AI
  • Meaningful equity of up to 1.0% in a $100M-valued company
  • Work on one of the hardest open problems in AI: teaching models to reason reliably when the stakes are real


Location: San Francisco, CA (in person, with relocation and joining bonuses available)


Compensation: Up to $350,000 base + up to 1.0% equity + comprehensive benefits, including unlimited AI tool budgets


If this aligns with your background, reply and Goliath Partners will be in touch!


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Goliath Partners

San Francisco Bay Area

Research Engineer

Graphic Design
•
4d ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Goliath Partners

San Francisco Bay Area

Brand Art Director

Graphic Design
•
2h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

CD PROJEKT RED

Poland

Subscribe our newsletter

New Things Will Always Update Regularly