G

Senior Research Engineer in Reinforcement Learning for Healthcare AI

Goliath Partners • San Francisco Bay Area
Relocation
Apply Now
AI Summary

Build and optimize reinforcement learning (RL) environments and reward systems for cutting-edge healthcare AI models focused on drug discovery and cancer treatment. Design complex RL tasks, evaluate multi-step agents, and improve data/feedback signals to enhance clinical reasoning. Requires 1+ years of expertise in RL, agent environments, and post-training ML systems.

Key Highlights
Own the full reinforcement learning loop, including task design, evaluation, and reward modeling for healthcare AI
Work with advanced post-training techniques like RLHF/RLAIF, SFT, and RLVR to refine language models and agents
Analyze and improve model failures to enhance data quality, reward signals, and clinical relevance
Key Responsibilities
Design and implement comprehensive reinforcement learning environments, including task design, action spaces, and evaluation harnesses
Conduct sophisticated post-training experiments on language models and agents using techniques like SFT, RLVR, RLHF/RLAIF, and reward modeling
Train and evaluate multi-step agents to navigate complex patient histories and unstructured medical data for clinical reasoning accuracy
Analyze model failures to iteratively improve data quality, reward signals, and feedback mechanisms
Technical Skills Required
Reinforcement Learning Machine Learning Post-Training Language Model Optimization
Benefits & Perks
Up to $350,000 base salary
Up to 1.0% equity
Unlimited AI tool budgets

Job Description


Our client is a hyper-growth, heavily backed startup that is revolutionizing the frontier of healthcare AI to radically improve drug discovery and solve cancer. Having recently secured a $100M raise, they are building the foundational data layer and post-training simulation environments where the next generation of medical AI models learn. They are seeking a powerhouse Research Engineer to own the full reinforcement learning loop and build the crucial evaluation and reward layer for frontier AI models using real-world clinical and genomic context.


Role & Impact

  • Build comprehensive RL environments, encompassing task design, action spaces, tool interfaces, verifiers, and evaluation harnesses.
  • Run sophisticated post-training experiments on language models and agents utilizing techniques like SFT, RLVR, RLHF/RLAIF, and reward modeling.
  • Train and evaluate multi-step agents to navigate complex patient histories and unstructured medical data to capture correctness in clinical reasoning.
  • Analyze model failures to continuously improve the data, rewards, and feedback signals that dictate what these systems learn.


Essential Skills

  • 1+ years of highly technical experience focusing on reinforcement learning, agent environments, LM post-training, or related ML systems.
  • Exceptional judgment around data quality, capable of assessing signal fidelity, coverage, and clinical relevance to support meaningful training tasks.
  • Strong ability to build scalable pipelines, debug complex training runs within large ML codebases, and move rapidly from research concepts to working prototypes.


Location: San Francisco, CA (In-person, with relocation and joining bonuses available)


Compensation: Up to $350,000 base + up to 1.0% equity + comprehensive benefits including unlimited AI tool budgets


If this aligns with your background, reply and Goliath Partners will be in touch!


Similar Jobs

Explore other opportunities that match your interests

Research Engineer

Graphic Design
•
2d ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Goliath Partners

San Francisco Bay Area

Senior Product Designer

Graphic Design
•
13h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Bending Spoons

Poland

Product Designer (AI-Integrated)

Graphic Design
•
14h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Bending Spoons

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly