A

Senior HPC Performance and Validation Engineer (AI/Research Infrastructure)

Addison Group • United State
Relocation
Apply Now

Lead performance benchmarking, validation, and optimization of large-scale GPU clusters for AI and HPC workloads. Develop automated testing, profiling tools, and CI/CD pipelines to ensure cluster readiness and efficiency. Collaborate with engineering teams to drive data-driven infrastructure decisions and resolve performance bottlenecks.

Key Highlights
Direct hire role with $180K–$260K base salary + $50K–$100K bonus potential
Hybrid work model (3 days in Dallas office, 2 days remote) with flexible scheduling
Focus on GPU cluster performance, AI workload profiling, and distributed HPC systems
Key Responsibilities
Design and implement automated tests to validate GPU node readiness, health, and utilization across large-scale HPC environments
Benchmark compute, storage, and network performance using established and workload-specific tests, including InfiniBand/RoCE environments
Profile AI and research workloads to identify bottlenecks and collaborate with engineering teams to optimize performance
Develop validation tools and CI/CD pipelines using Python, Go, and Kubernetes for continuous testing workflows
Establish monitoring and reporting systems to track cluster health and performance trends using observability tools
Document benchmark findings and guide architecture/capacity decisions based on data-driven insights
Lead technical investigations and collaborate with infrastructure and research teams to scale the platform
Technical Skills Required
Python Kubernetes Benchmarking and Profiling Tools (e.g., NVIDIA Nsight, MLPerf)
Benefits & Perks
100% paid medical, dental, and vision insurance
401(k) matching
25 days of paid time off (PTO)
Nice to Have
Experience with observability tools (Prometheus, Grafana, OpenTelemetry, ELK stack)
Familiarity with DCGM, ClusterKit, or other GPU performance tools
Degree in Computer Science or related field (though experience is prioritized)

Job Description

HPC Performance and Validation Engineer High Performance Computing / AI Infrastructure Dallas, TX Direct hire $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus Hybrid; three days in the Dallas office and two days remote. The manager determines the in-office days. This position is eligible for 100% paid medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, lunch on office days, and a gym membership.Our client is expanding the computing infrastructure used for large-scale AI, research, and simulation workloads.
Want the full job description? Read the complete details on LinkedIn, the original posting.
Continue on LinkedIn

This is a short excerpt. All rights to the full description belong to its original publisher.

See Jaabz jobs first on Google 1 tap · free · in AI Overviews Jaabz is on your Google Manage Preferred Sources

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

aic

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

enhance it

United State

Cloud Engineer

Devops
•
2d ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Toyota North America

United State

Subscribe our newsletter

New Things Will Always Update Regularly