Design and build datasets and evaluation systems for frontier AI models at a Series B AI company. Develop reward signals, evaluation rubrics, and frameworks to measure model performance across various workflows. Requires strong Python skills and 1-4 years of experience in production ML/LLM workflows.
Key Highlights
Key Responsibilities
Searching for Development & Programming roles that provide visa sponsorship? Connect with international employers through Development & Programming Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
Technical Skills Required
Benefits & Perks
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Nice to Have
Job Description
🚀 Software Engineer — RL Environments
We're partnering with a fast-growing (Series B) AI company that builds the datasets and evaluation systems frontier AI labs use to train their models. You'd be working hand-in-hand with top research teams designing the tasks models practise on, and the scoring that decides whether they're really getting smarter.
What you'd own:
🔹 Design data that exposes where models fail across finance, code & enterprise workflows
🔹 Build reward signals and evaluation rubrics for RLHF / RLVR training pipelines
🔹 Develop frameworks to measure dataset quality and its real impact on model performance
🔹 Turn ambiguous research goals into concrete, shippable systems
You must have:
This is a short excerpt. All rights to the full description belong to its original publisher.
Similar Jobs
Explore other opportunities that match your interests