Own the end-to-end architecture of Wayve's AI model development platform, ensuring reliability, scalability, and observability for autonomous driving systems. Lead cross-domain technical initiatives spanning distributed compute, ML Ops, data pipelines, and experiment scheduling to accelerate model iteration and deployment. Requires 10+ years of experience in large-scale distributed systems and ML infrastructure, with at least 3 years in a staff or principal-level engineering role.
Key Highlights
Design and evolve the architecture for a large-scale AI model lifecycle platform supporting autonomous driving.
Lead cross-domain technical strategy unifying web applications, distributed training, and data pipelines.
Build systems for optimizing model testing and scheduling using linear programming and heuristic optimization.
Key Responsibilities
Design and evolve the platform's architecture for reliability, observability, and scalability, setting performance, latency, and availability targets.
Unify the platform across front-end UIs, distributed training, Spark data pipelines, and optimization-based experiment scheduling.
Take on the hardest problems across subteams, leading architectural reviews and proposing pragmatic solutions.
Build systems that optimize how models are tested in simulation and on-road using linear programming and heuristic optimization.
Architect pipelines that ingest, transform, and enrich petabytes of fleet sensor data, driving efficient compute use across GPU, CPU, cloud, and edge.
Work with Product, Research, and Operations to align architecture with user needs and co-own the platform's long-term roadmap.
Technical Skills Required
Distributed Systems
Machine Learning Infrastructure
System Architecture
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
Benefits & Perks
Meaningful equity
Relocation support
Visa sponsorship
Hybrid working
Learning and development budgets
Comprehensive benefits including health insurance, dental, and enhanced parental leave
Nice to Have
Experience applying algorithmic or mathematical optimization (e.g. linear programming, graph algorithms) to operational or scheduling problems
Familiarity with end-to-end model lifecycle tooling, from data ingestion and training CI to model artifact tracking and evaluation workflows
Prior exposure to autonomous systems, robotics or other safety-critical domains
Experience with modern web frameworks (e.g. React, Flask, FastAPI) and how they integrate with backend systems
Understanding of data privacy, compliance and secure handling practices for large-scale sensor data
Job Description
Before the detail, here's the challenge you'd help us solve.We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that.Here’s what this particular role covers.The Model Development Platform team builds the infrastructure and tooling behind Wayve's AI model lifecycle, from data ingestion and training to experiment scheduling and on-road testing. Our work spans AI research, large-scale distributed systems and robotic operations, and lets researchers and engineers iterate fast and deploy autonomous driving models safely.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.