SRE Consultant
NTT DATA
NTT DATA
NTT DATA is seeking an innovative SRE Consultant to join our dynamic team. This role is integral to supporting and enhancing our application and platform services, ensuring robust incident management, root cause analysis, and efficient change execution.
We are looking for a proactive individual to provide on-call production support for mission-critical applications, resolving issues within defined recovery time objectives. The role involves leveraging diagnostic tools to maintain and restore system services and data, while also bringing engineering and operational management expertise to areas like AI and Cloud.
Embrace a Site Reliability Engineering (SRE) mindset by developing automated solutions for routine production support tasks. This position requires strong collaboration with cross-functional support teams and stakeholders, effective escalation of issues, and the ability to identify system bottlenecks for continuous improvement. A commitment to driving stability and operational excellence is key.
Key responsibilities include supporting functions and driving the execution of multiple Application/Platform support services. This encompasses incident triage, root cause analysis, change evaluation, execution, and validation, along with deployment management, business continuity planning (BCP), and risk/vulnerability management.
Provide essential on-call production support for mission-critical applications, ensuring swift resolution of issues within Recovery Time Objectives (RTO). Utilize diagnostic tools to maintain, troubleshoot, and restore system services and data, bringing engineering and operational management skills to domains such as AI and Cloud.
Champion continuous improvement to enhance stability and operational excellence. Apply an SRE mindset by building automated solutions for repetitive production support activities. Effectively understand and escalate issues, identify system bottlenecks, and propose process improvements. Collaborate and communicate effectively with multiple support teams and stakeholders.
A minimum of 12 years of production support experience is essential, coupled with proven success in implementing SRE processes within support teams. Experience supporting applications built on the Python stack is required.
Demonstrated knowledge of technology infrastructure, including data center, network, storage, and compute domains, is necessary. Hands-on experience with Google Cloud Platform (GCP), ideally with certification, is expected. Proficiency in configuring, managing, and optimizing GPU-based infrastructure for AI/ML workloads is a significant advantage.
Strong expertise in diagnosing and resolving capacity, performance, and scalability issues related to AI and machine learning workloads is crucial. Proven experience troubleshooting, supporting, and resolving production issues in Generative AI (GenAI) applications and services is also required. Familiarity with Python/Shell scripting, supporting Data Science/AI/ML platforms, and strong user/administrator experience with Linux/Unix command line are necessary.
Exposure to big data technologies like Hadoop, Spark, and Elasticsearch, along with extensive experience with monitoring systems such as Splunk and App Dynamics, is expected. The candidate should have experience serving as a liaison between infrastructure operations and business projects. Excellent communication, problem-solving, and analytical skills are vital. Willingness to work in 24x7 shift models is required.
NTT DATA
IT Consulting