S

AI Systems Engineer

sundayy • United Kingdom
Relocation
Apply
AI Summary

Design, optimize, and deploy high-performance AI inference systems for real-time multimodal models. Implement model acceleration techniques and containerize research models for production at scale. Requires deep expertise in GPU optimization, distributed systems, and full-cycle ownership from research to deployment.

Key Highlights
High-performance AI inference systems at scale
Model acceleration techniques (quantization, distillation, caching)
Full-cycle ownership from research to production
Key Responsibilities
Design, develop, and optimize high-performance AI inference systems for real-time applications
Implement and improve model acceleration techniques to enhance inference speed and efficiency
Build and maintain distributed systems capable of scaling to thousands of concurrent queries
Containerize research models and ensure their reliable deployment in production environments
Collaborate with research teams to translate innovative models into scalable production solutions
Profile and optimize code to maximize GPU performance and resource utilization
Troubleshoot and resolve system bottlenecks, latency issues, and reliability challenges
Stay updated with the latest advancements in AI hardware and software optimization techniques
Technical Skills Required
C++ CUDA Python Kubernetes Ray
Benefits & Perks
Competitive base salary within the range of £140,000 – £200,000
Equity options providing ownership in a fast-growing AI company
Comprehensive benefits package including health, dental, and vision insurance
Flexible working arrangements and supportive work environment
Opportunities for professional growth and continuous learning in cutting-edge AI research
Potential relocation support for candidates interested in moving to the San Francisco Bay Area

Job Description


About The Company

Inworld is a leading product-oriented research laboratory composed of top-tier AI researchers and engineers dedicated to advancing the frontiers of artificial intelligence. Our core focus is on developing best-in-class real-time multimodal models and the only real-time orchestration platform optimized for handling thousands of queries per second. With substantial backing, having raised over $125 million from prominent investors such as Lightspeed, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta, and Stanford, we have established ourselves as a pioneer in the AI industry. Our innovative technology has powered experiences for renowned companies including NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella, and Bible Chat. Recognized globally, we have been named one of CB Insights’ 100 Most Promising AI Companies and ranked among LinkedIn’s Top 10 Startups in the USA, underscoring our leadership and potential in the AI ecosystem.

About The Role

We are seeking a highly skilled and motivated AI Systems Engineer to join our dynamic team. In this role, you will be instrumental in designing, optimizing, and deploying high-performance AI inference systems at scale. Your work will directly impact the development of our multimodal models and real-time orchestration platform, ensuring they operate with minimal latency, maximum throughput, and exceptional reliability. You will collaborate closely with research teams to take cutting-edge models from concept to production, containerize and optimize them, and ensure seamless deployment across distributed systems. The ideal candidate thrives in an environment of ambiguity, demonstrates rapid learning, and possesses a passion for building scalable, high-performance AI infrastructure. Your expertise will help push the boundaries of what’s possible in real-time AI applications, contributing to the future of intelligent systems that are both powerful and efficient.

Qualifications

  • PhD in Computer Science, Physics, Mathematics, or an equivalent practical experience in building backend or ML systems
  • Deep understanding of modern serving frameworks and techniques such as vLLM or TRT-LLM
  • Hands-on experience with model acceleration methods including quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in programming languages such as C++, CUDA, Rust, or highly optimized Python
  • Experience with profiling code and optimizing performance on NVIDIA GPUs
  • Knowledge of distributed systems and scaling solutions, including Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference
  • Experience with handling thousands of concurrent connections reliably
  • Contributions to open-source inference engines or technical deep-dives in relevant areas
  • Full-cycle ownership capability from research model to containerization, optimization, and production deployment

Responsibilities

  • Design, develop, and optimize high-performance AI inference systems for real-time applications
  • Implement and improve model acceleration techniques to enhance inference speed and efficiency
  • Build and maintain distributed systems capable of scaling to thousands of concurrent queries
  • Containerize research models and ensure their reliable deployment in production environments
  • Collaborate with research teams to translate innovative models into scalable production solutions
  • Profile and optimize code to maximize GPU performance and resource utilization
  • Contribute to open-source projects and technical documentation to advance the field
  • Troubleshoot and resolve system bottlenecks, latency issues, and reliability challenges
  • Stay updated with the latest advancements in AI hardware and software optimization techniques

Benefits

  • Competitive base salary within the range of £140,000 – £200,000, commensurate with experience and location
  • Equity options providing ownership in a fast-growing AI company
  • Comprehensive benefits package including health, dental, and vision insurance
  • Flexible working arrangements and supportive work environment
  • Opportunities for professional growth and continuous learning in cutting-edge AI research
  • Potential relocation support for candidates interested in moving to the San Francisco Bay Area in the future

Equal Opportunity

Inworld is committed to creating an inclusive environment for all employees. We are an equal opportunity employer and do not discriminate based on race, ethnicity, gender, sexual orientation, age, disability, or any other protected characteristic. We believe diversity enhances innovation and are dedicated to fostering a workplace where everyone can thrive and contribute to our mission of advancing AI technology.

Similar Jobs

Explore other opportunities that match your interests

Senior Machine Learning Scientist

Machine Learning
•
2w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

monzo

United Kingdom

Lead Machine Learning Scientist

Machine Learning
•
3w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

monzo

United Kingdom

Multimodal ML Engineer

Machine Learning
•
3w ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

white circle

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly