Lead performance benchmarking, validation, and optimization of large-scale GPU clusters for AI and HPC workloads. Develop automated testing, profiling tools, and CI/CD pipelines to ensure cluster readiness and efficiency. Collaborate with engineering teams to drive data-driven infrastructure decisions and resolve performance bottlenecks.
Key Highlights
Direct hire role with $180K–$260K base salary + $50K–$100K bonus potential
Hybrid work model (3 days in Dallas office, 2 days remote) with flexible scheduling
Focus on GPU cluster performance, AI workload profiling, and distributed HPC systems
Looking to advance your Devops career with relocation support? Explore Devops Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
Key Responsibilities
Design and implement automated tests to validate GPU node readiness, health, and utilization across large-scale HPC environments
Benchmark compute, storage, and network performance using established and workload-specific tests, including InfiniBand/RoCE environments
Profile AI and research workloads to identify bottlenecks and collaborate with engineering teams to optimize performance
Develop validation tools and CI/CD pipelines using Python, Go, and Kubernetes for continuous testing workflows
Establish monitoring and reporting systems to track cluster health and performance trends using observability tools
Document benchmark findings and guide architecture/capacity decisions based on data-driven insights
Lead technical investigations and collaborate with infrastructure and research teams to scale the platform
Technical Skills Required
Python
Kubernetes
Benchmarking and Profiling Tools (e.g., NVIDIA Nsight, MLPerf)
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
Benefits & Perks
100% paid medical, dental, and vision insurance
401(k) matching
25 days of paid time off (PTO)
Nice to Have
Experience with observability tools (Prometheus, Grafana, OpenTelemetry, ELK stack)
Familiarity with DCGM, ClusterKit, or other GPU performance tools
Degree in Computer Science or related field (though experience is prioritized)
Job Description
HPC Performance and Validation Engineer High Performance Computing / AI Infrastructure Dallas, TX Direct hire $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus Hybrid; three days in the Dallas office and two days remote. The manager determines the in-office days. This position is eligible for 100% paid medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, lunch on office days, and a gym membership.Our client is expanding the computing infrastructure used for large-scale AI, research, and simulation workloads.
Want the full job description?
Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn
This is a short excerpt. All rights to the full description belong to its original publisher.