R

Senior Linux Administrator (AI & HPC Infrastructure)

Raas Infotek United State
Visa Sponsorship
Apply
AI Summary

Senior Linux Administrator responsible for maintaining and optimizing large-scale AI and HPC environments, including GPU clusters, high-performance storage, and data center networks. The role requires deep troubleshooting expertise across OS, hardware, and software layers, automation proficiency, and cross-team collaboration to ensure system reliability. Requires 7+ years of hands-on experience in enterprise Linux environments and advanced infrastructure management.

Key Highlights
Senior Linux administration for AI training, inference, and HPC workloads
Deep troubleshooting across OS, kernel, firmware, GPU drivers, and storage layers
Automation of administration tasks using Bash and Python
Key Responsibilities
Administer large-scale Linux environments supporting AI and HPC workloads
Troubleshoot complex OS, kernel, boot, package, firmware, driver, filesystem, and GPU driver issues across bare-metal server fleets
Build and maintain golden images, provisioning pipelines, and post-deployment validation procedures
Partner with cross-functional teams to diagnose and resolve cross-domain failures affecting cluster readiness
Investigate performance anomalies involving CPU, memory, NUMA, I/O, and kernel tuning
Automate administration and remediation tasks using Bash and Python
Technical Skills Required
Linux Administration (RHEL, Ubuntu, Rocky) GPU Server & Cluster Management Shell Scripting & Python
Nice to Have
Experience with Slurm, Kubernetes, or AI cluster schedulers
Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis
Support for validation labs or pre-production cluster certification

Job Description


Hi,

I hope you are doing well.

We have an urgent below position .If you are interested, please share your updated resume with the rate expectation.


Role: Senior Linux Administration

Location: Santa Clara, CA(Onsite)

Visa Status: GC/USC

Job description:


Skills: AI and HPC Infrastructure


The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.


WHAT THIS CANDIDATE WILL BE DOING

· Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.

· Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.

· Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.

· Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.

· Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.

· Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.

· Automate repeatable administration and remediation tasks with Bash and Python.

· Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.


WHAT WE NEED TO SEE

· 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.

· Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.

· Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.

· Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.

· Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.

· Strong shell scripting and Python-based automation capability.

· Working knowledge of storage and network dependencies affecting Linux host health.

· Ability to operate independently in ambiguous, high-severity production situations.


PREFERRED EXPERIENCE

· Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.

· Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.

· Experience supporting validation labs or pre-production cluster certification.


--

Thanks & Regards,

Anil Kumar

Raas Infotek Corporation.

262 Chapman Road, Suite 105A,

Newark, DE -19702

Direct No: 302-286-9932 Ext: 133

Email: [email protected]


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Tech Observer

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Hire IT People, Inc

United State

Principal Software Engineer - Cloud Management Platform

Devops
9h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United State

Subscribe our newsletter

New Things Will Always Update Regularly