Build next-generation AI evaluation benchmarks and perform red teaming for cutting-edge LLMs. Design, curate, and implement complex evaluation frameworks using Python and Jupyter/Colab environments. Requires MSc/PhD in STEM, proven research engineering experience, and strong analytical problem-solving skills.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
Work Location: US, remote
Engagement Model: Freelancer / Independent Contractor
Work Schedule: Flexible hours, based on the applicant's availability
Language skills required: English
Qualification Requirements: Successful completion of a role-specific coding/research assessment and background check.
Remote, United States · Fully Remote Research Engineer (RE) Curator DataForce by TransPerfect seeks a STEM Research Engineer to build next-generation AI evaluation benchmarks, perform red teaming, and advance LLM experimentation. Apply for this job
DataForce by TransPerfect is seeking exceptional Research Engineers and AI Specialists for the role of Research Engineer (RE) Curator to drive the creation, execution, and refinement of next-generation AI evaluation benchmarks. In this high-priority role, you will leverage your deep analytical mindset, Python proficiency, and experimental research background to curate, evaluate, red team, and benchmark cutting-edge Large Language Models (LLMs) and generative AI systems, ensuring top-tier accuracy, safety, and reasoning capabilities.
Main Responsibilities
Interested in remote work opportunities in Development & Programming? Discover Development & Programming Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
- Design, curate, and implement complex AI evaluation benchmarks and dataset frameworks to test advanced LLM capabilities and limitations.
- Conduct experimental research, red teaming, and systematic testing to identify model failure modes, edge cases, and reasoning gaps.
- Utilize Python, Jupyter/Colab environments, and modern IDEs to write scalable code, process datasets, and automate evaluation pipelines.
- Collaborate with AI research and engineering teams to translate complex subject-matter concepts into actionable model benchmarks.
- Document research findings, track codebase and dataset changes via Git, and deliver precise sourcing and evaluation reports.
Role Requirements
- Education: MSc or PhD in a STEM discipline (Computer Science, AI/ML, Statistics, Mathematics, Physics, Computational Sciences, or related field).
- Experience: Proven background as a Research Engineer, Applied Scientist, ML Engineer, Research Scientist, or AI Researcher.
- Programming & Tools: Strong hands-on Python programming skills, with daily fluency in Git, modern IDEs, and Jupyter/Colab environments.
- AI/ML Technical Expertise: Core experience in Machine Learning, LLMs, data analysis, and structured experimental research.
- Specialized Background (Preferred): Direct experience in AI evaluation, benchmark curation/development, red teaming, or model testing frameworks.
- Mindset: Strong research mindset paired with exceptional analytical and problem-solving capabilities.
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
DataForce by TransPerfect is part of the TransPerfect family of companies, the world’s largest provider of language and technology solutions for global business, with offices in more than 100 cities worldwide. We offer high-quality data for Human-Machine Interaction to some of the most prestigious technology companies in the world. Our department focuses on gathering, enriching, and processing data for Machine Learning in different AI domains. To learn more about DataForce please visit us at https://www.transperfect.com/dataforce
For more information on the TransPerfect Family of Companies, please visit our website at www.transperfect.com
Similar Jobs
Explore other opportunities that match your interests