Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering (Bengaluru, IN)

Deloitte

18+ yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : today
Actively hiring

Job description

Lead the architecture, engineering, and operation of cutting-edge hybrid cloud GPU infrastructure. This role involves designing robust GPU landing zones, managing cluster fabrics, and implementing cloud-native orchestration for AI workloads. You will focus on ensuring the performance, security, reliability, and cost-effectiveness of production NVIDIA GPU environments, seamlessly connecting public cloud capabilities with private AI estates.

This position offers a unique opportunity to shape multi-year AI infrastructure transformations and drive commercial strategies, while building and mentoring a high-performing team of specialists in AI data centers, GPU technology, networking, storage, cloud platforms, and AI infrastructure management.

Responsibilities

Architect and implement secure, scalable GPU landing zones, including account structures, network topology, and identity management. Select optimal NVIDIA GPU instances and cluster patterns for diverse AI tasks like distributed training and low-latency inference. Engineer cloud GPU clusters using Kubernetes or HPC schedulers, optimizing for high-performance networking and storage solutions. Develop hybrid connectivity for seamless workload portability between private and public cloud environments. Implement infrastructure-as-code, CI/CD pipelines, and GitOps for efficient deployment and management. Integrate cloud ML services while maintaining control over custom NVIDIA-based workloads. Establish comprehensive observability for GPU utilization, performance, reliability, and costs. Drive AI infrastructure FinOps strategies, focusing on cost optimization and showback/chargeback mechanisms. Engineer robust security measures for images, models, secrets, and data. Own market proposition, strategic alliances, and executive relationships within the AI infrastructure domain.

Qualifications

Requires extensive experience, typically 18+ years, in infrastructure, cloud, data-center, HPC, or platform engineering, with significant leadership in production AI/GPU environments. Demonstrated success in business development, executive advisory, alliance management, and large-scale program governance is essential.

Expertise in NVIDIA GPU architecture and systems (e.g., DGX/HGX) is critical, along with deep knowledge of GPU cluster design across compute, networking, storage, and management planes. Understanding AI workload characteristics for training, inference, and HPC is necessary. Proficiency in Kubernetes/OpenShift or Slurm, including GPU scheduling and multi-tenancy, is required. Strong Linux, containerization, CUDA ecosystem, NCCL, and driver fundamentals are a must. Experience with security, resilience, capacity planning, performance optimization, automation, and day-2 operations for production AI infrastructure is crucial.

A deep understanding of at least one major cloud provider (AWS, Azure, GCP) and working knowledge of others is expected. Experience with cloud GPU capacity, high-performance networking, managed Kubernetes/HPC, Infrastructure-as-Code (IaC), and cloud cost optimization is highly valued.

Essential Skills

NVIDIA GPU architectureGPU cluster designAI workload characteristicsKubernetesOpenShiftSlurmLinuxcontainersCUDA ecosystemNCCLdriversfirmwareGPU observabilitysecurityresiliencecapacity planningperformance optimizationautomationAWSMicrosoft AzureGoogle Cloudcloud GPU capacityhigh-performance networkingmanaged KubernetesHPCIaCcloud cost optimizationTerraformCI/CDGitOpsautoscalingquota automationreservationscapacity blocksenvironment promotioncloud ML servicesAI infrastructure FinOpssoftware supply chain

Good to Have

hyperscalersGPU cloud providersGCC cloud platformenterprise cloud AI estatescloud certificationNVIDIA certificationKubernetes certificationTerraform certificationnetwork certificationFinOps certificationhybrid AIsovereign cloudreserved GPU capacityproduction-scale inference

Highlights

  • Actively hiring

More Details

RoleDirector | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering (Bengaluru, IN)
DepartmentEngineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

IT Consulting

Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering (Bengaluru, IN) at Deloitte | SkillMX | SkillMX