Magic is seeking a Senior Distributed Systems Engineer to build data and coordination systems for ultra-long context AI inference and training. You will focus on high-performance storage, deep learning framework internals, and fault detection. This role requires deep expertise in distributed systems and public cloud platforms.
Key Highlights
Build data and coordination systems for ultra-long context AI inference and training.
Develop high-performance storage, caching, and fault detection/recovery systems.
Troubleshoot complex issues across GPUs, network, storage, OS, and cloud environments.
Technical Skills Required
Distributed Systems Design
Public Cloud Platforms
High-Availability Systems
High-Throughput Data Systems
Distributed DBMS Internals
Batch Processing Systems
Stream Processing Systems
Distributed File Systems
Deep Learning Frameworks (internals)
Benefits & Perks
Annual salary range: $225K - $550K
Significant equity compensation
401(k) plan with 6% salary matching
Generous health, dental, and vision insurance
Unlimited paid time off
Visa sponsorship
Relocation stipend to SF
Job Description
Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more reliably than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal.About the role:As a distributed systems engineer, you will build the data and coordination systems that enable ultra-long context inference and training on Magic’s GPU clusters.What you might work on:
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.