Optimize performance for machine learning software stacks focusing on training and inference for foundation models. Design performance benchmarks, automate workload analysis, and triage system bottlenecks to enhance GPU utilization. Requires a PhD with 1+ years experience or a Master's with 2+ years experience in CS, EE, or CSEE.
Key Highlights
Focus on optimizing deep learning workloads on state-of-the-art hardware and software platforms
Develop kernels and systems for new model architectures and algorithms
Establish the institution as a global hub for high-performance computing in deep learning
Key Responsibilities
Understand, analyze, profile, and optimize deep learning workloads on state-of-the-art hardware and software platforms.
Design and implement performance benchmarks and testing methodologies to evaluate application performance.
Build tools to automate workload analysis, workload optimization, and other critical workflows.
Triage system issues and identify bottlenecks and inefficiencies to enhance GPU utilization.
Support the team in developing appropriate kernels and systems for new model architectures and algorithms.
Participate in or lead design reviews with peers and stakeholders to evaluate available technologies.
Review code developed by other developers to ensure best practices in style, accuracy, testability, and efficiency.
Contribute to existing documentation or educational content based on product updates and user feedback.
Represent the institution at industry conferences and events to showcase HPC and deep learning capabilities.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
Technical Skills Required
Parallel Computing
System Level Coding
Distributed Machine Learning
Benefits & Perks
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off, sick leave and holidays
Paid Parental Leave
Employee Assistance Program
Life insurance and disability
Visa sponsorship
Job Description
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.The Role
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.