S

Reliability Engineer - AI Systems

sundayy • United Kingdom
Visa Sponsorship
Apply
AI Summary

As a Reliability Engineer at Anthropic, you will ensure the dependability of large language model serving systems, focusing on high availability, low latency, and safety. Your responsibilities include defining SLOs, designing monitoring systems, leading incident response, and enhancing infrastructure resilience across cloud providers. You need a strong background in distributed systems, ML infrastructure, and chaos engineering with excellent cross-team collaboration skills.

Key Highlights
Design and implement SLOs for LLM serving systems
Lead incident response for critical AI services
Develop high-availability infrastructure across multiple regions
Ensure reliability and safety of safeguard model serving
Experience with ML hardware accelerators and networking optimizations
Key Responsibilities
Develop and define Service Level Objectives (SLOs) for large language model serving systems
Design, implement, and maintain monitoring and observability systems across the token processing pipeline
Lead incident response efforts for critical AI services
Support the reliability and safety of safeguard model serving
Collaborate with cross-functional teams to identify system vulnerabilities and implement resilience enhancements
Technical Skills Required
Distributed systems ML infrastructure Chaos engineering
Benefits & Perks
Competitive annual salary ranging from £325,000 to £390,000 GBP
Comprehensive health and wellness benefits
Flexible hybrid work policy with minimum 25% in-office presence
Visa sponsorship available

Job Description


About Anthropic

Anthropic is dedicated to developing reliable, interpretable, and steerable artificial intelligence systems. Our mission is to ensure that AI technologies are safe and beneficial for users and society at large. We are a rapidly expanding team comprising researchers, engineers, policy experts, and business leaders who collaborate closely to build AI systems that prioritize safety, transparency, and societal good. Our commitment to innovation and responsible AI development positions us at the forefront of the industry, working tirelessly to create AI that aligns with human values and needs.

About The Role

As a Reliability Engineer at Anthropic, you will play a crucial role in maintaining and enhancing the dependability of our AI systems, particularly focusing on Claude, our flagship large language model. The AI Reliability Engineering (AIRE) team collaborates across various departments to improve the robustness and resilience of our critical serving pathways—from SDKs and network layers to API infrastructure and hardware accelerators. Your work will involve designing and implementing systems that ensure high availability and low latency, managing incident responses, and supporting infrastructure that underpins our AI safety commitments. This position offers a unique opportunity to influence the reliability of cutting-edge AI systems at a company committed to safety and societal benefit, providing a dynamic and cross-disciplinary environment that values holistic system thinking.

Qualifications

  • Bachelor’s degree or equivalent in a relevant field such as Computer Science, Engineering, or related disciplines
  • Strong background in distributed systems, infrastructure, or reliability engineering
  • Experience in operating large-scale model serving or training infrastructure, preferably with over 1000 GPUs
  • Knowledge of ML hardware accelerators including GPUs, TPUs, or Trainium
  • Understanding of ML-specific networking optimizations like RDMA and InfiniBand
  • Familiarity with AI-specific observability tools and frameworks
  • Experience with chaos engineering and resilience testing methodologies
  • Contributions to open-source infrastructure or ML tooling are advantageous
  • Excellent communication and collaboration skills with the ability to build strong cross-team relationships
  • Demonstrated ownership and user-centric approach to system reliability

Responsibilities

  • Develop and define Service Level Objectives (SLOs) for large language model serving systems, balancing availability, latency, and development velocity
  • Design, implement, and maintain monitoring and observability systems across the token processing pipeline
  • Assist in designing and deploying high-availability serving infrastructure across multiple regions and cloud providers
  • Lead incident response efforts for critical AI services, ensuring rapid resolution, conducting thorough incident reviews, and implementing systematic improvements
  • Support the reliability and safety of safeguard model serving, ensuring alignment with safety commitments and operational excellence
  • Collaborate with cross-functional teams to identify system vulnerabilities and implement resilience enhancements
  • Contribute to the development of best practices for system reliability, scalability, and safety

Benefits

  • Competitive annual salary ranging from £325,000 to £390,000 GBP
  • Comprehensive health and wellness benefits
  • Opportunities for professional growth and development in a pioneering AI environment
  • Flexible hybrid work policy with a minimum of 25% in-office presence
  • Visa sponsorship available for eligible candidates
  • Collaborative and innovative work culture focused on societal impact and safety

Equal Opportunity

Anthropic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, ethnicity, gender, sexual orientation, age, disability, or any other protected characteristic. We believe that diverse perspectives and backgrounds foster innovation and excellence, and we actively encourage candidates from all backgrounds to apply.

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

earlham institute

United Kingdom

Senior Validation Engineer

Programming
•
16h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

IC Resources

United Kingdom

Senior Full-Stack Software Engineer (Golang/React, Cloud-Native Architectures)

Programming
•
19h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

American Express

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly