O

HPC Systems Specialist

Optomi United State
Relocation
Apply
AI Summary

Support large-scale AI/HPC infrastructure environments, monitor and troubleshoot HPC and AI cluster environments, and ensure system health and uptime.

Key Highlights
Monitor and troubleshoot HPC and AI cluster environments
Support and troubleshoot enterprise storage systems
Investigate performance, connectivity, and network-related storage issues
Key Responsibilities
Monitor and troubleshoot HPC and AI cluster environments
Support and troubleshoot enterprise storage systems
Investigate performance, connectivity, and network-related storage issues
Technical Skills Required
HPC Enterprise Storage Kubernetes
Benefits & Perks
On-call rotation
Relocation assistance available
Direct Hire

Job Description


Optomi, in partnership with our client, are seeking an IOC Systems Specialist to support a large-scale AI/HPC infrastructure environment focused on high-performance compute and data-intensive workloads.


This role sits in a 24×7 operations center and is responsible for monitoring, troubleshooting, and ensuring reliability of distributed HPC systems and enterprise storage platforms.


  • Onsite Fort Worth, TX - relocation assistance available!
  • Direct Hire
  • On-call rotation


What you’ll do:

  • Monitor and troubleshoot HPC and AI cluster environments in a Tier 2 IOC/NOC setting
  • Support and troubleshoot enterprise storage systems (WEKA, VAST, Dell Isilon/PowerScale, or similar SAN/NAS)
  • Investigate performance, connectivity, and network-related storage issues (including VLAN and configuration validation)
  • Work with Kubernetes and Slurm-based compute environments
  • Use monitoring tools (Grafana) and ticketing systems (Jira) for incident management
  • Perform root cause analysis and collaborate with engineering teams for resolution
  • Ensure system health, uptime, and performance across distributed infrastructure


What we’re looking for:

  • Experience supporting enterprise storage or data center storage environments
  • Strong troubleshooting skills across storage, network, and compute systems
  • Familiarity with HPC or high-throughput infrastructure environments
  • Understanding of networking concepts (VLANs, connectivity, throughput)
  • Experience in operational support environments (IOC/NOC or Tier 2 support)
  • Exposure to Kubernetes, Slurm, or similar orchestration/workload tools is a plus
  • Join a cutting-edge AI infrastructure company building sustainable, large-scale GPU compute environments powering next-generation workloads.



Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

careerselite.com

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

thinking machines lab

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Jobot

United State

Subscribe our newsletter

New Things Will Always Update Regularly