SRE Reliability Engineer

NTT DATA

5+ yrs Bengaluru Full Time Hybrid (office + remote)
NTT DATA logo
Posted : today
Actively hiring

Job description

We are seeking a skilled Site Reliability Engineer (SRE) with expertise in Kubernetes, production support, observability, and Java applications. This role is crucial for maintaining the availability, reliability, scalability, and performance of our vital production systems. You will leverage your strong application troubleshooting capabilities alongside SRE and DevOps best practices, utilizing tools like Datadog and Prometheus for monitoring and Kubernetes for container orchestration. A solid grasp of Java applications and SQL databases is essential for diagnosing complex issues across application, infrastructure, and data layers.

Responsibilities

Key responsibilities include owning the operational health of production applications, providing L2/L3 support with incident management, and monitoring applications on Kubernetes. You will implement and maintain monitoring solutions, define service-level indicators (SLIs) and objectives (SLOs), and troubleshoot production issues across various technology stacks. Analyzing Java application logs, diagnosing performance bottlenecks, and using SQL for incident investigation are also core duties. This role involves participating in root-cause analysis, driving automation for operational efficiency, and collaborating with cross-functional teams to ensure production readiness. On-call duties are also part of the role.

Qualifications

We require a minimum of 5 years of IT experience, with substantial background in SRE, Production Support, Application Support, or DevOps. Essential skills include hands-on experience with Kubernetes, Datadog/Prometheus for observability, and supporting Java/J2EE or Java-based microservices in production. Proficiency in JVM troubleshooting, application logging, memory management, and SQL for database troubleshooting is necessary. Experience managing critical production incidents and a good understanding of REST APIs, microservices, and Linux/Unix environments are also expected. Familiarity with CI/CD pipelines and strong analytical troubleshooting abilities are key.

Essential Skills

KubernetesObservabilityJavaProduction SupportSQLDatadogPrometheusMicroservicesREST APIsLinux/UnixShell ScriptingCI/CDTroubleshootingIncident ManagementRoot Cause Analysis

Good to Have

AWSAzureDockerHelmGitLab ActionsGitHub ActionsELKOpenSearchSplunkAnsibleTerraformITIL

Highlights

  • Actively hiring

More Details

RoleSRE Reliability Engineer
DepartmentDevOps / Cloud
Employment TypeFull Time, Hybrid (office + remote)

About the Company

NTT DATA logo

NTT DATA

IT Consulting

SRE Reliability Engineer at NTT DATA | SkillMX | SkillMX