AppOLens CoE Engineer
NTT DATA
NTT DATA
Seeking a Lead Integration & Observability Specialist to spearhead enterprise observability and reliability solutions, while providing support for cloud-based integration platforms on AWS and Azure. This pivotal role emphasizes monitoring, automation, and ensuring the operational readiness of diverse applications, APIs, data pipelines, and messaging systems.
This is a hands-on technical leadership opportunity, involving mentoring and taking ownership of solutions within a Windows-based server and .NET application environment. While prior experience in these areas is advantageous, a strong aptitude for rapid learning and adaptability across various technologies, platforms, and application landscapes is highly valued.
Lead the deployment of enterprise observability across applications, APIs, services, batch jobs, and data pipelines. Standardize the design of monitoring, alerting, logging, metrics, and health checks for distributed systems. Integrate observability platforms with incident management and automation tools to enable proactive issue detection and resolution.
Support the reliability and availability of AWS/Azure integration platforms. Conduct advanced troubleshooting using logs, metrics, and traces to resolve production issues. Define operational readiness standards and non-functional requirements. Mentor engineers on best practices in observability and platform utilization.
Collaborate with product, support, and operations teams to enhance service stability and delivery. Work effectively across a range of application environments, including Windows servers, .NET applications, cloud platforms, and integration/messaging systems.
A minimum of 7 years of overall IT experience is required, with at least 5 years dedicated to Observability, Monitoring, or Reliability Engineering. Demonstrable hands-on expertise with enterprise observability tools like IBM Instana, Dynatrace, AppDynamics, Prometheus, or Grafana is essential. Proficiency in monitoring and alerting design, log management and analysis, metrics and distributed tracing, and understanding SLO/SLI concepts are critical.
Experience monitoring AWS/Azure workloads, strong troubleshooting and incident analysis skills, and a background in defining operational and non-functional requirements are necessary. Proven technical leadership and mentoring capabilities, coupled with experience in automation and ITSM integration (including ServiceNow workflows), are key.
Exposure to CI/CD, release management, cloud integration, and messaging systems is required. Preferred skills include experience with Windows server environments, supporting .NET applications, IIS, Windows services, or related Microsoft technologies. The ability to learn new tools and platforms quickly and work across diverse technology stacks is highly desirable.
NTT DATA
IT Consulting