Data Engineer specializing in NLP and large-scale data processing to support cutting-edge foundation model research. Responsible for collecting, curating, and preprocessing high-quality datasets, developing web crawling solutions, and implementing scalable data pipelines. Requires extensive experience in Python, data engineering, and automation, with a Bachelor's degree in a technical field.
Key Highlights
Specialization in Natural Language Processing (NLP) and large-scale data processing for foundation models.
Development of efficient web crawling solutions, APIs, and automated workflows for data collection.
Implementation of scalable data pipelines for efficient processing, storage, and distribution to research teams.
Searching for Data Science roles that provide visa sponsorship? Connect with international employers through Data Science Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
Key Responsibilities
Rapidly collect, curate, and preprocess datasets based on detailed specifications provided by NLP researchers.
Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.
Technical Skills Required
Python
Data Engineering
Web Crawling
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Benefits & Perks
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off, sick leave and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability
Nice to Have
Master’s degree or PhD degree or equivalent experience in Computer Science, Data Engineering, or related technical fields
Proven track record of supporting NLP or AI research teams with rapid and reliable data delivery
Experience working with large language models, including evaluation, efficient inference, and prompt engineering
Experience with refining outputs from large-scale AI models, such as LLM-generated data
Contributions to open-source projects, coding competitions, or high visibility in coding communities (e.g., GitHub, Stack Overflow)
Familiarity with the latest advancements in NLP data processing and large language model technologies
Job Description
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.