Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
Deloitte
Deloitte
Lead the architectural design, engineering, and operational management of production NVIDIA GPU infrastructure across public cloud environments. This role involves integrating these capabilities with private AI estates, encompassing GPU landing zones, cluster fabrics, cloud-native orchestration, and ensuring performance, security, reliability, and AI infrastructure FinOps.
We are seeking a visionary leader to drive multi-year AI infrastructure transformations. This position offers the opportunity to shape strategic initiatives, lead commercial negotiations, and manage quality, risk, and economics for cutting-edge AI technologies.
Architect robust GPU landing zones, managing accounts, subscriptions, and projects. Design network topology, private connectivity, identity, encryption, policy, and observability frameworks. Select optimal NVIDIA GPU instances and cluster patterns for diverse AI workloads like distributed training and inference. Engineer high-performance cloud GPU clusters using Kubernetes or HPC schedulers, optimizing placement, topology, and network adapters. Design scalable storage solutions and implement efficient data management strategies for ingestion, checkpointing, and cross-region movement. Build seamless hybrid connectivity for workload portability between private and public cloud GPU environments.
Implement infrastructure-as-code using Terraform, establish CI/CD pipelines, and manage GitOps workflows. Automate autoscaling, quotas, reservations, and environment promotions. Integrate cloud ML services where appropriate while maintaining control over custom NVIDIA-based workloads. Establish comprehensive observability for GPU availability, utilization, latency, throughput, reliability, and cost. Drive AI infrastructure FinOps, focusing on commitments, spot usage, idle detection, rightsizing, and showback/chargeback models. Engineer robust security measures for images, drivers, secrets, endpoints, and the software supply chain. Lead market proposition, account strategy, pipeline development, and strategic alliances.
Build and cultivate a nationally recognized team of specialists in AI data centers, GPUs, networking, storage, cloud, and platforms. Foster a culture of innovation and excellence within the team.
A minimum of 18 years of experience in infrastructure, cloud, data center, HPC, or platform engineering is required, with a significant track record of leading production AI/GPU estates. Proven expertise in business development, executive advisory, alliance building, and large-program governance is essential.
Deep understanding of NVIDIA GPU architecture and systems, including DGX/HGX or equivalent certified platforms. Experience in designing GPU clusters encompassing compute, high-speed networking, storage, and management planes. Familiarity with AI workload characteristics, including distributed training, fine-tuning, RAG, batch inference, real-time inference, and HPC. Proficiency with Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation, and multi-tenancy is crucial.
Solid grasp of Linux, containers, the CUDA ecosystem, NCCL, drivers, firmware, and fundamental GPU observability. Expertise in security, resilience, capacity planning, performance optimization, automation, and day-2 operations for production AI infrastructure. Demonstrated deep expertise in at least one major cloud provider (AWS, Microsoft Azure, or Google Cloud) with working awareness of others. Experience with cloud GPU capacity management, high-performance networking, managed Kubernetes/HPC, IaC, and cloud cost optimization is highly valued.
Deloitte
Engineering