Principal Site Reliability Engineer

UnitedHealth Group

Fresher Hyderabad Full Time Hybrid (office + remote)
UnitedHealth Group logo
Posted : today
Actively hiring

Job description

Join Optum, a global leader in health optimization, where technology drives better health outcomes for millions. Our inclusive culture fosters collaboration with talented peers, offering comprehensive benefits and career growth. You'll contribute to advancing health globally, embodying our commitment to Caring, Connecting, and Growing together.

This Principal Site Reliability Engineer role is pivotal in ensuring the seamless operation of critical business applications and infrastructure. You will be instrumental in maintaining service reliability, driving innovation, and fostering a culture of operational excellence.

We are dedicated to creating a healthier world for everyone by making the health system work better. Our commitment extends to addressing health disparities and enabling equitable care for all, regardless of background or circumstances.

Responsibilities

Provide operational support for critical business applications and infrastructure. Lead and participate in P1/P2 incident management, ensuring timely resolution and communication. Troubleshoot complex production issues and coordinate cross-functional teams. Ensure timely resolution of incidents and service requests within defined SLAs and SLOs. Contribute to and continuously improve SRE practices, operational excellence, and service reliability. Drive root cause analysis and implement preventive measures. Develop and maintain monitoring, alerting, and observability solutions. Implement automation for operational tasks, deployments, monitoring, and compliance. Design, develop, and deploy AI-powered solutions to address complex business challenges. Explore innovative solutions to enhance reliability engineering practices. Participate in on-call rotations and provide support during critical incidents. Identify opportunities to reduce toil through automation and process improvements. Support disaster recovery (DR) and high availability (HA) initiatives. Mentor junior engineers and promote a culture of continuous learning. Comply with all applicable Company policies and procedures.

Qualifications

Hands-on experience with Terraform, GitHub Actions, CI/CD pipelines, and Infrastructure as Code. Proficiency with cloud platforms such as Azure and AWS. Experience utilizing monitoring and observability tools including Splunk, Datadog, or Grafana. Proven track record in handling critical P1/P2 incidents within large-scale enterprise environments. Solid experience in Site Reliability Engineering, Production Support, or Operations Engineering. Strong understanding of SLA, SLO, SLI, and incident management processes. Demonstrated scripting and automation skills using Python, PowerShell, or similar languages. Excellent communication, problem-solving, and stakeholder management abilities.

Essential Skills

TerraformGitHub ActionsCI/CDInfrastructure as CodeAzureAWSSplunkDatadogGrafanaIncident ManagementSREProduction SupportOperations EngineeringSLASLOSLIPythonPowerShellCommunicationProblem-SolvingStakeholder Management

Good to Have

KubernetesDevOpsAI/MLMicroservicesCloud-Native Architectures

Highlights

  • Actively hiring

More Details

RolePrincipal Site Reliability Engineer
IndustryHealthcare
DepartmentEngineering, Site Reliability Engineering
Employment TypeFull Time, Hybrid (office + remote)

About the Company

UnitedHealth logo

UnitedHealth

Healthcare

Principal Site Reliability Engineer at UnitedHealth Group | SkillMX | SkillMX