Senior Software Engineer

NVIDIA

10+ yrs Bengaluru Full Time Remote
NVIDIA logo
Posted : today
Actively hiring

Job description

Join a forward-thinking team as a Senior Software Engineer, focusing on infrastructure expertise to shape the future of our enterprise Observability, Automation, and AI-driven Reliability Platform. This role is pivotal in developing highly scalable distributed systems and platform services across Storage, Compute, Network, VMware, and OpenShift infrastructure. You will be instrumental in shifting infrastructure operations from reactive monitoring to proactive, predictive, and AI-driven autonomous management.

This position offers an exciting opportunity to innovate and implement cutting-edge solutions for large-scale enterprise environments. You will contribute to a platform that enhances reliability and efficiency through advanced automation and AI.

Responsibilities

Design, build, and operate robust distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at scale. Develop reusable platform services, APIs, and automation frameworks to empower self-service and streamline infrastructure operations. Construct scalable telemetry and event-processing systems capable of handling billions of infrastructure signals across metrics, logs, traces, and alerts. Engineer intelligent, AI-native reliability capabilities, including agentic workflows for anomaly detection, root-cause analysis, and closed-loop remediation. Drive technical architecture and engineering direction across critical infrastructure domains, tackling complex, cross-team challenges. Focus on production readiness, emphasizing software quality, scalability, security, performance, and maintainability.

Qualifications

Possess a Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience, coupled with over 10 years in software engineering, SRE, infrastructure, or distributed systems, including demonstrated technical leadership. Exhibit strong software engineering skills in Go, Python, or similar languages, with a proven track record in designing and building production-grade distributed systems, platform services, APIs, and automation. Demonstrate experience managing complex software/platform initiatives across multiple teams or infrastructure domains from conception to measurable impact. Showcase a deep understanding of distributed systems, event-driven architectures, microservices, APIs, and high-throughput data processing, leveraging technologies like Kafka, NATS, or gRPC. Possess extensive experience with modern observability and telemetry platforms such as OpenTelemetry, Prometheus, VictoriaMetrics, Vector, Loki, Grafana, or ClickHouse. Have robust SRE and infrastructure knowledge across Kubernetes/OpenShift, VMware, bare-metal, storage, and networking, utilizing automation tools like Terraform or Ansible. Effectively solve ambiguous problems, influence technical direction, mentor engineers, establish engineering standards, and deliver tangible operational and business outcomes.

Essential Skills

GoPythonDistributed SystemsPlatform ServicesAPIsAutomationKafkaNATSgRPCOpenTelemetryPrometheusVictoriaMetricsVectorLokiGrafanaClickHouseKubernetesOpenShiftVMwareBare-metalStorageNetworkingTerraformAnsible

Good to Have

Generative AIAIOpsLLMsAgentic AILangChainLlamaIndexAutoGenFastAPI

Highlights

  • Actively hiring

More Details

RoleSenior Software Engineer
DepartmentSoftware Development
Employment TypeFull Time, Remote

About the Company

Nvidia logo

Nvidia

IT Consulting

Senior Software Engineer at NVIDIA | SkillMX | SkillMX