Own the end-to-end reinforcement learning loop for a healthcare AI startup building frontier reasoning models. Design RL environments, generate training data, and execute post-training experiments to improve model reasoning. Requires 1+ years of technical experience in RL or LM post-training and a strong academic background in a technical field.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Our client is a hyper-growth AI startup building the post-training infrastructure for the next generation of frontier models in healthcare. Backed by Y Combinator and executives from Google DeepMind and OpenAI, and now valued at $100M, they are building the foundational data layer and simulation environments where advanced AI models learn to reason in one of the most complex, high-stakes domains in the world. This is a small, elite team where every engineer shapes the product, and they are seeking a powerhouse Research Engineer to own the full reinforcement learning loop from end to end.
This isn't a role where you tune one piece of a pipeline someone else designed. You'll own the entire cycle: designing the environments, generating the data, running the training, building the evaluations, and deciding what comes next based on what the models get wrong. If you've ever wanted to see what happens when one person can move from idea to trained model to real-world
impact without waiting on five other teams, this is it.
Role & Impact
- Own end-to-end model training, from data and environment design through post-training, evaluation, and iteration
- Build RL environments from the ground up, including task design, action spaces, tool interfaces, verifiers, and evaluation harnesses
- Run cutting-edge post-training experiments on language models and agents using techniques like SFT, RLVR, RLHF/RLAIF, and reward modeling
- Train multi-step agents to reason over long, complex, unstructured real-world data, where getting the answer right matters
- Build the evaluation and reward layer that defines what "correct reasoning" means for frontier AI models
- Dig into model failures and close the loop, continuously improving the data, rewards, and feedback signals that shape what these systems learn
- Help set the research direction and technical foundations of the company as an early, high-ownership member of a small team
Looking to advance your Graphic Design & Art career with relocation support? Explore Graphic Design & Art Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
Essential Skills
- 1+ years of highly technical experience in reinforcement learning, agent environments, LM post-training, or related ML systems
- Hands-on experience running training end to end, not just one stage of the pipeline
- Exceptional judgment about data quality, including assessing signal fidelity, coverage, and domain relevance for meaningful training tasks
- A strong ability to build scalable pipelines, debug complex training runs in large ML codebases, and move rapidly from research concepts to working prototypes
- A degree in Computer Science, Mathematics, Physics, or a related technical field from a top-tier university, or an equally exceptional academic track record
- Evidence of exceptional ability, such as top honors, a strong research record, or standout achievements in competitive programming, olympiads, or similar
- A builder's mindset: you want to own the problem, not just a ticket
Nice to have
- A Master's degree or PhD in Machine Learning, Computer Science, or a related field from a leading research institution
- Publications at top venues such as NeurIPS, ICML, ICLR, or ACL
- Experience at a frontier AI lab, elite research group, or top-tier startup
- Experience with RL for language models, agentic systems, or large-scale distributed training
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
Why this role
- Full-cycle ownership: you run the entire loop from environment to trained model, with no handoffs and no silos
- A small team with outsized impact, where your work directly shapes the core product and the company's trajectory
- Exceptional talent density: work alongside researchers and engineers from top universities and leading AI labs
- Backed by Y Combinator and leaders from Google DeepMind and OpenAI, so you'll be building alongside people who have shaped frontier AI
- Meaningful equity of up to 1.0% in a $100M-valued company
- Work on one of the hardest open problems in AI: teaching models to reason reliably when the stakes are real
Location: San Francisco, CA (in person, with relocation and joining bonuses available)
Compensation: Up to $350,000 base + up to 1.0% equity + comprehensive benefits, including unlimited AI tool budgets
If this aligns with your background, reply and Goliath Partners will be in touch!
Similar Jobs
Explore other opportunities that match your interests