Job Description

Job Purpose

Build and operate large?scale GPU compute pods to deliver predictable, high?throughput, low?latency training and inference services across 10K GPU cluster.

Roles & Responsibilities

  • Implementation
  • Stand up multi?pod GPU clusters (rack/power/cooling layouts; TOR/leaf connectivity; IB/Ethernet host configs).

    Implement GPU partitioning (MIG/vGPU profiles) and quota policies for multi?tenant environments.

    Integrate cluster schedulers (Slurm/Kubernetes) with GPU device plugins, node feature discovery, and accounting/quotas.

  • Operations
  • Own day?2 operations across firmware/driver/DCGM/NVML lifecycles; execute change windows with zero/low downtime.

    Capacity planning (GPU/CPU/Memory/NIC) and bin?packing strategies for heterogeneous GPU SKUs.

  • Performance & Optimization
  • Tune NCCL/UCX, GPU clocks/persistence, GPU Direct Storage, NUMA/locality, and CUDA runtime parameters.

    ...

    Ready to Apply?

    Take the next step in your AI career. Submit your application to Larsen & Toubro today.

    Submit Application