Senior Site Reliability Engineer

Oracle

3–5 yrs Bengaluru Full Time Hybrid (office + remote)
Oracle logo
Posted : today
Actively hiring

Job description

As a Senior Site Reliability Engineer, you will proactively design and architect robust infrastructure and services, ensuring optimal reliability and functionality. This role involves forecasting demands, responding to capacity needs, and collaborating closely with software development teams to build scalable and dependable systems. You will conduct thorough data collection to maintain and enhance operations and reliability.

Your expertise will be crucial in performing incident response and maintenance tasks, providing comprehensive health and performance reporting, and identifying key opportunities for automation. You will effectively communicate about services, clearly articulating the potential impact of changes. This position also requires providing essential support for technology and documenting incidents, while also experimenting with new tools to assess their impact and staying abreast of site reliability trends.

Responsibilities

Key responsibilities include designing and architecting infrastructure for reliability and functionality, forecasting capacity demands, and collaborating with development teams on scalable solutions. You will perform data collection, triage, and root cause analysis for incident and service lifecycle management, monitoring service performance and documenting conditions.

Identify and drive automation opportunities, developing scripts to enhance monitoring, analysis, and remediation. Communicate service attributes and the impact of changes, providing operational support and escalating issues as needed. Participate in on-call duties and resolve technical issues across various services to meet service level objectives. Experiment with new tools and technologies to improve infrastructure performance and reliability, while identifying and implementing efficiency improvements.

Qualifications

A minimum of 3 to 5 years of experience in site reliability engineering is required. Proficiency in designing and architecting for reliability, capacity planning, and incident response is essential. You should possess strong skills in automation, troubleshooting complex issues, and performance monitoring.

Experience in conducting root cause analysis, documenting incidents, and experimenting with new tools is expected. The ability to communicate technical information clearly and effectively is vital. This role requires a proactive approach to continuous learning and improvement, with a strong understanding of industry trends and best practices. Fluency in English for reading, writing, and speaking is mandatory.

Essential Skills

site reliability engineeringinfrastructure designcapacity planningincident responseautomationtroubleshootingperformance monitoringroot cause analysissecurity standards

Good to Have

scriptingnew tool experimentation

Highlights

  • Actively hiring

More Details

RoleSenior Site Reliability Engineer
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Oracle logo

Oracle

IT Consulting

Senior Site Reliability Engineer at Oracle | SkillMX | SkillMX