SRE Reliability Enginner

NTT DATA

5+ yrs Bengaluru Full Time Hybrid (office + remote)
NTT DATA logo
Posted : today
Actively hiring

Job description

Join NTT DATA as a Site Reliability Engineer (SRE) and contribute to the stability and performance of critical production systems. This role emphasizes Kubernetes, observability, and Java-based applications. We seek proactive individuals passionate about innovation and growth within a dynamic organization.

Leverage your expertise in production support, observability tools like Datadog and Prometheus, and container orchestration with Kubernetes. A deep understanding of Java applications and SQL is vital for effective troubleshooting across all system layers. Your work will directly impact the availability, reliability, scalability, and performance of our business-critical services.

Responsibilities

Own the end-to-end reliability and operational health of production applications. Provide L2/L3 support, managing incident triage, troubleshooting, and communication. Monitor and manage Kubernetes deployments, including scaling and resource utilization. Implement and refine observability strategies using Datadog/Prometheus, focusing on metrics, logs, and alerts. Define and track SLIs, SLOs, and error budgets. Troubleshoot complex issues across Java applications, microservices, Kubernetes, and databases. Analyze JVM performance and SQL queries for root cause analysis. Drive permanent remediation for recurring issues through automation and engineering enhancements.

Automate operational tasks, support production releases, and collaborate closely with Development, Infrastructure, and Business teams. Participate in on-call rotations to ensure continuous service availability. Contribute to building robust monitoring dashboards and operational runbooks to enhance early detection and reduce recovery times.

Qualifications

A minimum of 5 years of IT experience, with substantial focus on SRE, Production Support, Application Support, or DevOps. Proven hands-on experience with Kubernetes and containerized environments. Proficiency in using Datadog and/or Prometheus for monitoring and observability. Strong background supporting Java/J2EE or Java-based microservices in production. Solid understanding of JVM troubleshooting, application logging, memory management, and performance tuning. Expertise in SQL for diagnosing relational database and application data issues. Experience managing P1/P2 incidents, including RCA and problem management. Familiarity with REST APIs, microservices, and distributed systems.

Proficiency in Linux/Unix environments and shell scripting is essential. Understanding of CI/CD pipelines, release management, and deployment processes. Excellent analytical and troubleshooting skills are required to diagnose issues across various technology layers. Experience with cloud platforms like AWS or Azure, Docker, Helm, and CI/CD tooling such as GitLab/GitHub Actions is preferred.

Essential Skills

KubernetesObservabilityJavaProduction SupportSQLDatadogPrometheusMicroservicesREST APIsLinux/UnixShell ScriptingCI/CDRelease ManagementTroubleshootingIncident ManagementRCA

Good to Have

AWSAzureDockerHelmGitLab ActionsGitHub ActionsELKOpenSearchSplunkAnsibleTerraformLoad BalancingNetworkingDNSCertificatesApplication SecurityITILChange Management

Highlights

  • Actively hiring

More Details

RoleSRE Reliability Enginner
DepartmentSite Reliability Engineer, SRE, DevOps Engineer, Production Support Engineer, Application Support Engineer
Employment TypeFull Time, Hybrid (office + remote)

About the Company

NTT DATA logo

NTT DATA

IT Consulting

SRE Reliability Enginner at NTT DATA | SkillMX | SkillMX