Senior Site Reliability Engineer

UnitedHealth Group

5+ yrs Hyderabad Full Time Hybrid (office + remote)
UnitedHealth Group logo
Posted : today
Actively hiring

Job description

Join Optum, a global leader dedicated to improving health outcomes through technology. We foster an inclusive culture where you can collaborate with talented peers, access comprehensive benefits, and pursue career development. Contribute to advancing health optimization worldwide by connecting individuals with the care, pharmacy benefits, data, and resources they need.

This role focuses on enhancing the health system's effectiveness for everyone. We are committed to addressing health disparities and promoting equitable care, making it an enterprise priority to help all individuals live their healthiest lives, regardless of background or circumstances.

Responsibilities

Key responsibilities include designing, deploying, and managing Azure infrastructure and services, optimizing cloud resource utilization. You will leverage Infrastructure as Code (IaC) tools like Terraform and Ansible to ensure scalable, secure, and resilient infrastructure.

Expertise in Kubernetes is essential for managing containerized applications and implementing robust orchestration solutions. You will also focus on AI/DevOps/MLOps, building automated CI/CD pipelines and developing tools for data engineering and data science teams.

Implement automation for operational processes and create self-service capabilities for Data Scientists. Ensure platform security and performance through best practices, monitoring, and optimization. Collaboration with cross-functional teams and mentorship of junior members are vital. Implement comprehensive monitoring and logging solutions, troubleshoot performance and security issues, and drive continuous improvement through industry trend adoption.

Qualifications

Required qualifications include a Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience. A minimum of 5 years in DevOps, Site Reliability Engineering (SRE), or a similar role is necessary.

Proficiency is expected in Azure cloud platforms, including core services like Storage, Networking, security, App Services, and AKS. Experience with IaC tools (Terraform, CloudFormation), CI/CD tools (preferably GitHub Actions), and scripting languages (Python, Bash, Ruby) is crucial.

Solid experience with Data and AI platforms (Databricks, Snowflake), orchestrating tools (Airflow, Data Factory), and monitoring tools (Splunk, Grafana, Prometheus) is required. Proven expertise in Kubernetes, containerization technologies (Docker, Podman), and strong problem-solving, analytical, communication, and collaboration skills are essential for success in this role.

Essential Skills

KubernetesDockerContainerizationCI/CDAzure CloudInfrastructure as CodeGrafanaPrometheusAKSTerraformAnsiblePythonBashRubySplunkGitHub ActionsDatabricksSnowflakeAirflowData Factory

Good to Have

AWSGCPFirewallsLoad BalancersDDOS solutions

Highlights

  • Actively hiring

More Details

RoleSenior Site Reliability Engineer
Employment TypeFull Time, Hybrid (office + remote)

About the Company

UnitedHealth logo

UnitedHealth

IT Consulting

Senior Site Reliability Engineer at UnitedHealth Group | SkillMX | SkillMX