Senior Site Reliability Engineer
UnitedHealth Group
UnitedHealth Group
Join Optum, a global leader dedicated to improving health outcomes through technology. We connect individuals with the care, pharmacy benefits, data, and resources necessary for healthier lives. Our inclusive culture fosters collaboration with talented peers, offers comprehensive benefits, and provides ample career development opportunities. Contribute to advancing global health optimization and experience a culture of Caring, Connecting, and Growing together.
This Senior Site Reliability Engineer role is crucial for ensuring the operational excellence of critical business applications and infrastructure. You'll be instrumental in maintaining high availability and performance, while driving innovation in SRE practices.
Oversee critical applications and infrastructure, leading P1/P2 incident management for swift resolution and effective communication. Troubleshoot complex production issues, collaborating with cross-functional teams to restore services within defined SLAs and SLOs.
Enhance SRE practices, operational excellence, and service reliability through continuous improvement. Conduct root cause analysis, implement preventive measures, and manage incident backlogs. Develop and maintain robust monitoring, alerting, and observability solutions.
Implement automation for operational tasks, deployments, and compliance. Design and deploy AI-powered solutions to tackle complex business challenges and improve incident detection, operational efficiency, and decision-making. Explore innovative reliability engineering practices.
Participate in on-call rotations, supporting critical incidents and major events. Identify opportunities to reduce toil via automation and process improvements. Support disaster recovery and high availability initiatives.
Mentor junior engineers, fostering a culture of continuous learning and operational excellence. Ensure responsiveness during critical business events and production issues. Adhere to all company policies and directives regarding work arrangements.
Demonstrated hands-on experience with Terraform, GitHub Actions, CI/CD pipelines, and Infrastructure as Code.
Proficiency with cloud platforms such as Azure and AWS, alongside experience in monitoring and observability tools like Splunk, Datadog, or Grafana.
Proven track record in handling critical P1/P2 incidents within large-scale enterprise environments, with solid experience in Site Reliability Engineering, Production Support, or Operations Engineering.
Strong understanding of SLA, SLO, SLI, and incident management processes. Excellent scripting and automation skills using Python, PowerShell, or similar languages.
Exceptional communication, problem-solving, and stakeholder management abilities are essential for this role.
UnitedHealth
Healthcare