Job Description
Job Purpose
Build and operate large?scale GPU compute pods to deliver predictable, high?throughput, low?latency training and inference services across 10K GPU cluster.
Roles & Responsibilities
Stand up multi?pod GPU clusters (rack/power/cooling layouts; TOR/leaf connectivity; IB/Ethernet host configs).
Implement GPU partitioning (MIG/vGPU profiles) and quota policies for multi?tenant environments.
Integrate cluster schedulers (Slurm/Kubernetes) with GPU device plugins, node feature discovery, and accounting/quotas.
Own day?2 operations across firmware/driver/DCGM/NVML lifecycles; execute change windows with zero/low downtime.
Capacity planning (GPU/CPU/Memory/NIC) and bin?packing strategies for heterogeneous GPU SKUs.
Tune NCCL/UCX, GPU clocks/persistence, GPU Direct Storage, NUMA/locality, and CUDA runtime parameters.
...
Ready to Apply?
Take the next step in your AI career. Submit your application to Larsen & Toubro today.
Submit Application