Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering (Bengaluru, IN)

Deloitte

18+ yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : today
Actively hiring

Job description

Join a dynamic team driving functional excellence and innovation in hybrid cloud engineering. This role focuses on architecting and deploying cutting-edge AI infrastructure solutions.

As a leader in this space, you will shape the future of AI data centres and build a nationally recognized team of specialists. Your expertise will be crucial in delivering scalable, high-performance AI infrastructure.

This position offers the opportunity to influence multi-year transformation programs and lead commercial negotiations, ensuring governance, delivery quality, and sound business economics.

Responsibilities

Take ownership of end-to-end architecture for NVIDIA GPU compute, management nodes, control plane, high-speed fabrics, storage, backup, and security.

Design and deploy GPU clusters, selecting optimal platforms for various AI workloads like training, fine-tuning, and inference.

Architect robust InfiniBand and Ethernet networks, focusing on topology, congestion control, and performance engineering.

Define high-performance file and object storage architectures to meet demanding data lifecycle requirements.

Engineer AI data-centre requirements, including rack layout, power, cooling, cabling, and DCIM integration.

Implement and manage Kubernetes/OpenShift or Slurm-based platforms with advanced GPU scheduling and workload isolation.

Integrate and operationalize NVIDIA AI Enterprise suite and related services.

Establish automated provisioning, lifecycle management, validation, and performance benchmarking processes.

Operate the platform using SRE principles, ensuring observability, incident management, capacity planning, and security.

Optimize GPU utilization, job throughput, and overall cost efficiency.

Drive market proposition, account strategy, pipeline development, and executive relationships for AI infrastructure services.

Build and mentor a high-performing team of AI data-centre, GPU, network, storage, cloud, and platform engineering specialists.

Qualifications

A minimum of 18 years of experience in infrastructure, cloud, data-centre, HPC, or platform engineering, with significant leadership in production AI/GPU environments.

Proven success in business development, executive advisory, strategic alliances, and governing large-scale technology programs.

Deep expertise in NVIDIA GPU architectures and systems, including DGX/HGX or equivalent certified platforms.

Demonstrated experience in designing GPU clusters across compute, networking, storage, and control planes.

Strong understanding of AI workload characteristics, including distributed training, RAG, and various inference types.

Hands-on architecture and leadership experience with Kubernetes/OpenShift and/or Slurm, covering GPU scheduling and multi-tenancy.

Proficiency in Linux, container ecosystems, CUDA, NCCL, GPU drivers, firmware, and observability.

Experience managing production AI infrastructure, focusing on security, resilience, performance, automation, and operations.

Solid understanding of high-density data-centre environments, including power, cooling, and physical infrastructure.

Experience designing and operating high-performance storage and InfiniBand/Ethernet fabrics for AI and HPC workloads.

Excellent commercial acumen, stakeholder management, and executive communication skills.

Experience with AI factories, GPU clouds, HPC centres, enterprise AI data centres, GCCs, OEMs, or colocation facilities is advantageous.

Familiarity with NVIDIA DGX/HGX, NVIDIA AI Enterprise, Spectrum-X, InfiniBand, BasePOD, SuperPOD, or NVIDIA validated reference architectures is a plus.

Experience with NVIDIA-certified infrastructure and production-scale AI platform deployments is preferred.

Relevant certifications such as CKA/CKS, Red Hat, NVIDIA, Linux, networking, storage, or data-centre certifications are beneficial.

Experience building AI infrastructure practices, centres of excellence, or nationally scaled specialist engineering teams is desirable.

Essential Skills

NVIDIA GPU architecturesKubernetes/OpenShiftSlurmLinuxContainerizationCUDA ecosystemNCCLGPU driver managementFirmware managementHigh-density data-centre designInfiniBandEthernet fabricsBusiness developmentExecutive advisoryStrategic alliancesStakeholder managementExecutive communication

Good to Have

CKA/CKSRed Hat certificationsNVIDIA certificationsLinux certificationsNetworking certificationsStorage certificationsData-centre certifications

Highlights

  • Actively hiring

More Details

RoleDirector | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering (Bengaluru, IN)
DepartmentEngineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

IT Consulting

Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering (Bengaluru, IN) at Deloitte | SkillMX | SkillMX