Job Description
We are seeking a highly skilled HPC / GPU Infrastructure Engineer to design, deploy, optimize, and manage NVIDIA GPU environments on Dell PowerEdge servers within large‑scale Linux high‑performance computing systems. The role will support advanced workloads including GenAI and HPC applications.
Key Responsibilities
- Deploy, configure, and validate GPU clusters using NVIDIA Base Command Manager.
- Perform multi‑node testing, benchmarking, and performance tuning.
- Execute cluster‑level code upgrades aligned with approved versions and compatibility standards.
- Manage and maintain server infrastructure, including iDRAC configuration, monitoring, troubleshooting, firmware updates (BIOS, NICs, storage, and system components).
- Utilize Redfish APIs for automation, monitoring, and system management.
- Configure and manage NVIDIA BlueField DPUs.
- Support GPU observability and performance profiling.
- Ensure...
Ready to Apply?
Take the next step in your AI career. Submit your application to TEKsystems today.
Submit Application