Principal Site Reliability Engineer
UnitedHealth Group
UnitedHealth Group
Join Optum, a global leader in health optimization, where technology drives better health outcomes for millions. Our inclusive culture fosters collaboration with talented peers, offering comprehensive benefits and career growth. You'll contribute to advancing health globally, embodying our commitment to Caring, Connecting, and Growing together.
This Principal Site Reliability Engineer role is pivotal in ensuring the seamless operation of critical business applications and infrastructure. You will be instrumental in maintaining service reliability, driving innovation, and fostering a culture of operational excellence.
We are dedicated to creating a healthier world for everyone by making the health system work better. Our commitment extends to addressing health disparities and enabling equitable care for all, regardless of background or circumstances.
Provide operational support for critical business applications and infrastructure. Lead and participate in P1/P2 incident management, ensuring timely resolution and communication. Troubleshoot complex production issues and coordinate cross-functional teams. Ensure timely resolution of incidents and service requests within defined SLAs and SLOs. Contribute to and continuously improve SRE practices, operational excellence, and service reliability. Drive root cause analysis and implement preventive measures. Develop and maintain monitoring, alerting, and observability solutions. Implement automation for operational tasks, deployments, monitoring, and compliance. Design, develop, and deploy AI-powered solutions to address complex business challenges. Explore innovative solutions to enhance reliability engineering practices. Participate in on-call rotations and provide support during critical incidents. Identify opportunities to reduce toil through automation and process improvements. Support disaster recovery (DR) and high availability (HA) initiatives. Mentor junior engineers and promote a culture of continuous learning. Comply with all applicable Company policies and procedures.
Hands-on experience with Terraform, GitHub Actions, CI/CD pipelines, and Infrastructure as Code. Proficiency with cloud platforms such as Azure and AWS. Experience utilizing monitoring and observability tools including Splunk, Datadog, or Grafana. Proven track record in handling critical P1/P2 incidents within large-scale enterprise environments. Solid experience in Site Reliability Engineering, Production Support, or Operations Engineering. Strong understanding of SLA, SLO, SLI, and incident management processes. Demonstrated scripting and automation skills using Python, PowerShell, or similar languages. Excellent communication, problem-solving, and stakeholder management abilities.
UnitedHealth
Healthcare