Senior Site Reliability Engineer
UnitedHealth Group
UnitedHealth Group
Join Optum, a global leader dedicated to improving health outcomes through technology. We foster an inclusive culture where you can collaborate with talented peers, access comprehensive benefits, and pursue career development. Contribute to advancing health optimization worldwide by connecting individuals with the care, pharmacy benefits, data, and resources they need.
This role focuses on enhancing the health system's effectiveness for everyone. We are committed to addressing health disparities and promoting equitable care, making it an enterprise priority to help all individuals live their healthiest lives, regardless of background or circumstances.
Key responsibilities include designing, deploying, and managing Azure infrastructure and services, optimizing cloud resource utilization. You will leverage Infrastructure as Code (IaC) tools like Terraform and Ansible to ensure scalable, secure, and resilient infrastructure.
Expertise in Kubernetes is essential for managing containerized applications and implementing robust orchestration solutions. You will also focus on AI/DevOps/MLOps, building automated CI/CD pipelines and developing tools for data engineering and data science teams.
Implement automation for operational processes and create self-service capabilities for Data Scientists. Ensure platform security and performance through best practices, monitoring, and optimization. Collaboration with cross-functional teams and mentorship of junior members are vital. Implement comprehensive monitoring and logging solutions, troubleshoot performance and security issues, and drive continuous improvement through industry trend adoption.
Required qualifications include a Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience. A minimum of 5 years in DevOps, Site Reliability Engineering (SRE), or a similar role is necessary.
Proficiency is expected in Azure cloud platforms, including core services like Storage, Networking, security, App Services, and AKS. Experience with IaC tools (Terraform, CloudFormation), CI/CD tools (preferably GitHub Actions), and scripting languages (Python, Bash, Ruby) is crucial.
Solid experience with Data and AI platforms (Databricks, Snowflake), orchestrating tools (Airflow, Data Factory), and monitoring tools (Splunk, Grafana, Prometheus) is required. Proven expertise in Kubernetes, containerization technologies (Docker, Podman), and strong problem-solving, analytical, communication, and collaboration skills are essential for success in this role.
UnitedHealth
IT Consulting