Design and operate scalable ML infrastructure and platform capabilities to support the full machine learning lifecycle. Lead complex technical initiatives from architecture to production to improve developer productivity and operational excellence. Requires 5+ years of software engineering experience with strong expertise in Python, Kubernetes, and distributed systems.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr Platform Engineer, ML Infrastructure based in United States.
This senior engineering role focuses on building the foundational platforms and infrastructure that enable machine learning teams to develop and operate systems more efficiently. You’ll design scalable tooling and services supporting the full ML lifecycle across cloud and on-premises environments. The role combines platform engineering, distributed systems, Kubernetes, and ML infrastructure in a highly technical environment. You’ll independently lead complex initiatives from architecture through production and ongoing operations. Working closely with ML and infrastructure engineers, you’ll translate technical needs into reliable, easy-to-use platform capabilities. Your work will improve developer productivity, operational excellence, scalability, and the delivery of production ML systems.
Accountabilities
- Design, build, and operate scalable ML infrastructure and platform capabilities supporting experimentation, training, deployment, and production operations.
- Develop developer tooling, services, automation, and infrastructure that help ML and engineering teams build and operate production systems more efficiently.
- Lead complex technical initiatives independently, from problem definition and architecture through implementation, rollout, and operational ownership.
- Make architectural decisions that balance immediate delivery needs with long-term scalability, reliability, maintainability, and developer experience.
- Partner with ML engineers, infrastructure teams, and other stakeholders to understand needs and deliver effective platform solutions.
- Identify and solve challenging infrastructure problems involving performance, reliability, scalability, and operational efficiency.
- Drive adoption and continuous improvement by incorporating feedback from engineering teams using the platform.
- Maintain high standards for software quality, production readiness, observability, and operational excellence.
- Deliver platform capabilities that create measurable engineering and business impact across multiple teams and use cases.
Searching for Development & Programming roles that provide visa sponsorship? Connect with international employers through Development & Programming Jobs with Visa Sponsorship opportunities actively seeking talented professionals.
- 5+ years of professional software engineering experience, particularly in platform engineering, infrastructure, or distributed systems.
- Strong Python engineering skills, including experience developing production services, SDKs, automation, or platform tooling.
- Proven experience designing, building, and operating production platforms used by multiple engineering teams.
- Solid understanding of ML platform architecture and the end-to-end machine learning lifecycle, including experimentation, distributed training, model deployment, and production operations.
- Experience building and operating applications on Kubernetes and cloud platforms, with AWS experience preferred.
- Strong understanding of production reliability, observability, scalability, and operational best practices.
- Strong technical judgment and the ability to independently drive complex initiatives from discovery through production while collaborating across technical teams.
- Experience with developer platforms, internal tooling, or services that improve engineering productivity and reduce operational complexity is preferred.
- Familiarity with workflow orchestration or distributed computing technologies such as Airflow, Kubeflow, Ray, Spark, or similar systems is a plus.
- Experience designing or optimizing distributed, GPU-intensive compute platforms for ML training, inference, or large-scale image processing is preferred.
- Experience supporting production ML platforms in computer vision, robotics, or related technical domains is advantageous.
- Demonstrated technical leadership through architecture, mentoring, or influencing technical direction across teams.
Explore our comprehensive directory of visa sponsorship jobs from employers worldwide who are ready to sponsor talented international professionals.
- Base salary range of $160,000–$287,000 per year, depending on experience, qualifications, education, location, and skills.
- Eligibility for an annual performance bonus.
- Competitive benefits package.
- Full-time, remote position within the United States.
- Visa sponsorship may be available for this position.
- Opportunities for career development, mentorship, and learning and development.
- Inclusive and collaborative work environment focused on meaningful, technically challenging work.
- Opportunity to work on advanced machine learning, robotics, and intelligent machinery technologies with cross-disciplinary teams.
Interested in opportunities specifically in United State? Discover our dedicated Visa Sponsorship Jobs in United State page featuring roles from top employers in this location.
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Similar Jobs
Explore other opportunities that match your interests