B

Software Engineer – Reinforcement Learning Environments

Visa Sponsorship
Apply
AI Summary

Design datasets and evaluation frameworks for AI model training and evaluation. Build reward signals and metrics for RLHF/RLVR pipelines. Collaborate with AI researchers to improve model performance and alignment.

Key Highlights
Build training data and evaluation infrastructure for frontier AI labs
Design datasets exposing AI model failure modes across domains
Develop reward signals and metrics for RLHF and RLVR training
Key Responsibilities
Design datasets that expose meaningful AI model failure modes across domains such as software engineering, finance, and enterprise workflows
Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines
Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact
Technical Skills Required
Python Reinforcement Learning Data Pipeline Development
Benefits & Perks
Competitive equity
Profit sharing
Visa sponsorship
Nice to Have
Experience with RLHF, RLVR, AI benchmarking, or AI evaluation frameworks
Experience working with AI safety organizations or reinforcement learning environments
Hands-on experience building real-world and synthetic data pipelines
Startup or early-stage company experience

Job Description


Software Engineer – RL Environments (3 Open Positions)

Location: San Francisco, CA (Onsite)

Employment Type: Full-Time

Experience: 1–4 Years

Compensation: $180,000–$220,000 Base + Competitive Equity + Profit Share (Total Compensation up to ~$500K)

Visa Sponsorship: H-1B, O-1, and OPT supported

About the Role

Our client, an innovative AI startup, is hiring Software Engineers – RL Environments to help build the training data and evaluation infrastructure used by leading frontier AI labs.

In this role, you will design datasets, evaluation frameworks, and reward signals that directly influence how next-generation AI models are trained and evaluated. You will collaborate closely with AI researchers, experiment with data collection strategies, identify model failure modes, and develop metrics that improve model performance and alignment.

Key Responsibilities

  • Design datasets that expose meaningful AI model failure modes across domains such as software engineering, finance, and enterprise workflows.
  • Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines.
  • Design and execute lightweight experiments to improve model capabilities.
  • Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact.
  • Create and maintain real-world and synthetic data pipelines.
  • Partner with AI research teams to translate training objectives into scalable data and evaluation solutions.

Required Qualifications

  • 1–4 years of professional Software Engineering experience.
  • Strong programming and problem-solving skills.
  • Passion for AI, reinforcement learning, and data-driven systems.
  • Ability to design experiments and derive actionable insights from complex datasets.
  • Comfortable working across multiple domains, including software engineering, finance, and enterprise applications.
  • Proven ability to build, ship, and iterate quickly in a fast-paced environment.

Preferred Qualifications

  • Experience with RLHF, RLVR, AI benchmarking, or AI evaluation frameworks.
  • Experience working with AI safety organizations or reinforcement learning environments.
  • Hands-on experience building real-world and synthetic data pipelines.
  • Startup or early-stage company experience is a plus.

Why Join?

  • Work on cutting-edge AI infrastructure used by leading frontier AI labs.
  • Highly competitive compensation package.
  • Equity participation and significant profit-sharing opportunities.
  • Visa sponsorship available (H-1B, O-1, OPT).
  • Opportunity to work alongside world-class AI researchers and engineers in a fast-growing startup environment.


Similar Jobs

Explore other opportunities that match your interests

.NET Full Stack Developer

Programming
47m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State

Embedded Software Engineer

Programming
2h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

varda space industries

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Actalent

United State

Subscribe our newsletter

New Things Will Always Update Regularly