Job Description

Job Purpose

Provide high?throughput, consistent storage tiers (Scratch/HPS + Object) for large?scale training data ingest, checkpoints, and inference artifacts.

Roles & Responsibilities

  • Implementation
  • Design/expand Lustre/BeeGFS HPS; NVMe?oF and Object (S3) tiers; align with AI dataflow and GDS.

    Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.

  • Operations
  • Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention.

    Proactive detection of hot spots and metadata contention; schema for small?file handling.

  • Performance & Optimization
  • Tune RDMA paths, page cache, IO schedulers; validate end?to?end I/O profiles for LLM training/inference.

  • Reliability & Incident
  • Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA...

    Ready to Apply?

    Take the next step in your AI career. Submit your application to Larsen & Toubro today.

    Submit Application