We are seeking a Senior HPC Platform Engineer to design, build, and operate large-scale computing infrastructure for AI and simulation workloads. This role involves managing compute, storage, and networking systems, optimizing performance, and mentoring junior engineers. Candidates should have 5+ years of HPC experience and expertise in parallel computing, distributed storage, and GPU-based systems.
Key Highlights
Design and implement HPC infrastructure across servers, storage, networking, and data center systems.
Optimize system performance, reliability, and capacity planning.
Collaborate with engineering, operations, and research teams to support new systems.
Mentor junior engineers and document architectures and procedures.
Looking to advance your Devops career with relocation support? Explore Devops Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
Key Responsibilities
Design and implement HPC infrastructure across servers, storage, networking, and related data center systems.
Find performance bottlenecks and improve system throughput, reliability, and resource use.
Partner with engineering, operations, and research teams to install, configure, test, and support new systems.
Use monitoring data to troubleshoot complex issues and plan for future capacity needs.
Assess new technologies and vendor solutions, and recommend improvements to the platform.
Document architectures, configurations, and operating procedures while helping junior engineers develop their skills.
Technical Skills Required
HPC
Parallel Computing
GPU
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
Benefits & Perks
Medical, dental, and vision insurance
401(k)
25 days of PTO
HSA contribution
Gym membership
Lunch on office days
Nice to Have
Experience with ZFS
Experience with NiFi
Experience with computational fluid dynamics workloads
Job Description
HPC Platform Engineer High Performance Computing / AI Infrastructure Dallas, TX Direct hire $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus Hybrid; three days in the Dallas office and two days remote. The team manager determines the in-office schedule. This position is eligible for medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, a gym membership, and lunch on office days.Our client develops advanced computing and cloud infrastructure for demanding AI, research, and simulation workloads. The organization is investing in the systems and engineering teams needed to expand its computing capacity.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.