Systems Operations Engineer
NTT DATA
NTT DATA
Join NTT DATA, a global leader in business and technology services, and contribute to our mission of accelerating client success and positive societal impact through responsible innovation. We are seeking a skilled Systems Operations Engineer to enhance our Hyderabad-based team. This role is crucial for maintaining and improving the operational health of critical applications and platforms.
As a key member of our team, you will leverage cutting-edge technologies, including AI, cloud platforms, and large language models, to drive operational excellence. You'll be instrumental in implementing SRE principles and developing automated solutions to streamline production support activities, ensuring high availability and performance of our systems.
Take ownership of critical support functions, including incident triage, root cause analysis, and change management for multiple application and platform services. Ensure the timely resolution of production issues within defined recovery time objectives (RTO) by providing on-call support for mission-critical applications.
Utilize diagnostic tools to maintain, troubleshoot, and restore system services and data. Drive continuous improvement initiatives to enhance system stability and operational excellence, adopting an SRE mindset to build automated solutions for routine production support tasks. Effectively communicate and collaborate with various support teams and stakeholders to ensure seamless operations.
A minimum of 12 years of experience in production support is required, with a proven track record of implementing SRE processes within support teams. You should possess hands-on experience supporting applications built on a Python stack and have a strong understanding of technology infrastructure encompassing data center, network, storage, and compute domains.
Demonstrated expertise in Google Cloud Platform (GCP), ideally with certification, is essential. Experience configuring, managing, and optimizing GPU-based infrastructure for AI/ML workloads is highly valued. You must have proven experience in diagnosing and resolving capacity, performance, and scalability issues related to AI and machine learning applications, as well as troubleshooting production issues in Generative AI (GenAI) services.
Proficiency in Python and/or Shell scripting is necessary, along with experience supporting Data Science, AI, or ML platforms. Strong user/administrator experience with Linux/Unix command line is expected. Exposure to big data technologies like Hadoop, Spark, and Elasticsearch, alongside experience with monitoring systems such as Splunk and App Dynamics, is crucial. The ability to serve as a liaison between technology operations and business projects, coupled with exceptional problem-solving and analytical skills, is paramount. Willingness to work in 24x7 shift models is also required.
NTT DATA
IT Consulting