Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

Deloitte

18+ yrs Bengaluru Full Time Hybrid (office + remote)
Deloitte logo
Posted : 1 week ago
Actively hiring

Job description

This pivotal role involves leading the end-to-end architecture and deployment of advanced hybrid cloud solutions, with a specific focus on NVIDIA GPU compute.

The position requires a visionary leader to drive functional excellence and spearhead innovation in AI infrastructure.

This includes architecting high-performance computing environments designed for cutting-edge AI workloads, from training to real-time inference.

Responsibilities

Architect and implement comprehensive hybrid cloud solutions encompassing NVIDIA GPU compute, networking, storage, and management planes.

Design and deploy GPU clusters optimized for diverse AI workloads, including distributed training and inference.

Engineer robust high-speed network architectures and define resilient file and object storage strategies.

Develop AI data-centre specifications covering rack layout, power, cooling, and security.

Implement and manage Kubernetes/OpenShift or Slurm-based platforms with advanced GPU scheduling capabilities.

Integrate and operationalize NVIDIA AI Enterprise suite and related AI services.

Establish automated lifecycle management for firmware, drivers, and images, including provisioning and testing.

Operate the platform adhering to SRE principles, ensuring observability, incident management, and capacity planning.

Optimize GPU utilization, job throughput, and overall system efficiency for cost-effectiveness and performance.

Drive market strategy, business development, and executive relationships for AI infrastructure services.

Shape long-term AI infrastructure transformation programs, managing commercial negotiations and delivery quality.

Build and mentor a high-performing team of AI data-centre, GPU, network, storage, and platform engineering specialists.

Qualifications

A minimum of 18 years of progressive experience in infrastructure, cloud, data-centre, HPC, or platform engineering is essential, with significant leadership in production AI/GPU environments.

Demonstrated success in business development, executive advisory, strategic alliances, and governing large-scale technology programs is required.

Deep expertise in NVIDIA GPU architectures and systems, including DGX/HGX or equivalent certified platforms, is a must.

Proven track record in designing GPU clusters, covering compute, high-speed networking, storage, and control/management planes.

Solid understanding of AI workload characteristics, such as distributed training, fine-tuning, RAG, and various inference types.

Hands-on experience architecting and leading with Kubernetes/OpenShift and/or Slurm, including GPU scheduling and multi-tenancy.

Strong foundational knowledge of Linux, containers, CUDA ecosystem, NCCL, GPU drivers, firmware, and observability tools.

Extensive experience with production AI infrastructure, focusing on security, resilience, capacity management, performance, automation, and operational maturity.

Thorough understanding of high-density data-centre requirements, including power, cooling, networking, and facility dependencies.

Experience in designing and operating high-performance storage and InfiniBand/Ethernet fabrics for AI and HPC workloads.

Exceptional commercial acumen, stakeholder management skills, and executive communication capabilities.

Exposure to various AI infrastructure environments like AI factories, GPU clouds, HPC centres, and enterprise data centres.

Familiarity with NVIDIA DGX/HGX, NVIDIA AI Enterprise, Spectrum-X, InfiniBand, BasePOD, SuperPOD, or NVIDIA validated reference architectures.

Experience with NVIDIA-certified infrastructure and production-scale AI platform deployments is crucial.

Relevant certifications such as CKA/CKS, Red Hat, NVIDIA, Linux, networking, storage, or data-centre certifications are highly valued.

Experience building AI infrastructure practices, centres of excellence, or scaled specialist engineering teams is advantageous.

Candidates should demonstrate an ability to deliver production-ready, secure, resilient, and scalable NVIDIA GPU infrastructure, leading to high GPU utilization, predictable performance, and efficient workload onboarding. Deliverables include reliable capacity expansion, automated provisioning, measurable improvements in workload throughput and energy efficiency, strong platform reliability, enhanced security, and high client satisfaction. A strong market positioning and a sustainable AI infrastructure pipeline are also key outcomes. The ultimate goal is a high-performing, nationally recognized AI infrastructure engineering organization.

Essential Skills

NVIDIA GPU architecturesDGX/HGXKubernetesOpenShiftSlurmLinuxCUDANCCLInfiniBandEthernetHigh-speed networkingStorage architectureData-centre designBusiness developmentExecutive advisoryStrategic alliancesAI workload characteristicsContainerizationGPU schedulingPerformance optimizationAutomationIncident managementCapacity planningDisaster recoverySecurity controlsStakeholder managementExecutive communicationCloud infrastructureHPC environmentsNVIDIA AI EnterpriseSpectrum-XBasePODSuperPODNVIDIA validated reference architecturesNVIDIA-certified infrastructureAI platform deploymentsCKA/CKSRed HatNVIDIA certificationsLinux certificationsNetworking certificationsStorage certificationsData-centre certifications

Highlights

  • Actively hiring

More Details

RoleDirector | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
IndustryEngineering
DepartmentDirector
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Deloitte logo

Deloitte

Engineering