Senior Site Reliability Engineer

UnitedHealth Group

5+ yrs Gurugram Full Time Hybrid (office + remote)
UnitedHealth Group logo
Posted : today
Actively hiring

Job description

Join Optum, a global leader in healthcare technology dedicated to improving lives through innovative solutions. We are seeking a skilled Senior Site Reliability Engineer (SRE) to enhance the reliability, performance, and scalability of our critical systems. This role offers the opportunity to contribute to impactful projects, including CI/CD, performance testing, and cloud migrations, while collaborating with teams in both India and the United States.

At Optum, we foster an inclusive culture that values collaboration, continuous learning, and career development. You will be part of a dynamic team working to optimize health outcomes worldwide. We believe in Caring, Connecting, and Growing together to make a significant difference in the communities we serve.

Responsibilities

Ensure the high availability, performance, and scalability of essential systems through robust site reliability engineering practices. Drive the advancement of observability systems by building scalable solutions with OpenTelemetry and other modern tools, focusing on improved monitoring, tracing, and logging. Lead and contribute to key projects, including performance testing, CI/CD pipeline enhancements, and migrating infrastructure and applications from on-premise to cloud environments. Actively participate in incident response, troubleshoot issues, and conduct post-mortem analyses to prevent recurrence. Develop and maintain automation tools to streamline processes and boost system reliability. Collaborate effectively with global engineering teams and stakeholders across different time zones to ensure project alignment and success. Identify and implement opportunities for continuous improvement in system reliability and operational efficiency. Leverage AI-powered tools to enhance observability, incident response, and overall operational effectiveness.

Qualifications

Possess a Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. Bring a minimum of 5 years of experience in Site Reliability Engineering, DevOps, or a comparable technical role. Demonstrate expertise in architecting and implementing observability platforms using tools like OpenTelemetry, Datadog, Splunk, or Grafana. Have hands-on experience with containerization and orchestration technologies such as Docker and Kubernetes. Proficient in CI/CD tools including Jenkins and GitHub Actions, with experience in building automation pipelines. Possess solid knowledge of public cloud platforms, particularly Azure, and experience with on-premise to cloud migrations. Exhibit deep understanding of systems architecture, cloud infrastructure, networking, and automation tools. Showcase proven scripting and programming skills in languages like Python, Go, Powershell, or Bash, and experience with Infrastructure-as-Code tools such as Terraform and Ansible. Demonstrate excellent problem-solving skills, with practical experience in incident management and root cause analysis. Exhibit strong communication and collaboration skills, adept at working within distributed, cross-timezone teams.

Essential Skills

Site Reliability EngineeringDevOpsOpenTelemetryDatadogSplunkGrafanaDockerKubernetesJenkinsGitHub ActionsAzureTerraformAnsiblePythonGoPowershellBashIncident ManagementTroubleshootingRoot Cause Analysis

Good to Have

PaymentsFintechHealthcareCloud Security

Highlights

  • Actively hiring

More Details

RoleSenior Site Reliability Engineer
IndustryInformation Technology & Services
DepartmentDevOps / Cloud, Cloud Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

UnitedHealth logo

UnitedHealth

Information Technology & Services

Senior Site Reliability Engineer at UnitedHealth Group | SkillMX | SkillMX