Senior Site Reliability Engineer

Oracle

3–5 yrs Bengaluru Full Time Work from office
Oracle logo
Posted : today
Actively hiring

Job description

Join our team as a Senior Site Reliability Engineer and play a pivotal role in architecting and maintaining robust, scalable, and highly available infrastructure. You will be instrumental in ensuring the reliability and functionality of our services through proactive design, capacity planning, and performance optimization. This role offers an exciting opportunity to collaborate with development teams, implement innovative automation solutions, and contribute to the continuous improvement of our site reliability practices.

This position is focused on individual contribution and requires a proactive approach to infrastructure design and incident management. You will be a key player in forecasting demand, responding to capacity needs, and performing essential data collection for operational optimization. Your expertise will be crucial in handling incident response and maintenance tasks, providing comprehensive health and performance reporting, and identifying opportunities for automation. This role also involves communicating effectively about services and potential impacts of changes, providing critical technology support, and documenting incidents. We encourage experimentation with new tools and trends to continuously enhance our site reliability.

Responsibilities

Key Responsibilities include designing and architecting infrastructure for optimal reliability and functionality, forecasting infrastructure demands, and responding to capacity needs to ensure sufficient resources. You will collaborate with software development teams to build scalable infrastructures and features.

Your role will involve performing data collection, triage, and technical analysis to optimize operations and infrastructure reliability. You will independently monitor services, document their performance, and leverage your knowledge for incident response, root cause analysis, and maintenance tasks, including software installs, upgrades, and security updates. Providing health and performance reports and taking action based on data trends is essential.

Identify and implement automation opportunities to enhance efficiency and reduce manual effort. Develop automation tools or scripts for monitoring, analysis, and issue remediation. Conduct thorough testing to ensure automation effectiveness and expected results.

Communicate service attributes, requirements, and the impact of changes to technical teams. Provide operational support for technology, escalating issues as needed, and participate in on-call rotations. Troubleshoot and resolve technical issues across various services, aiming to meet service level objectives (SLOs) and conduct post-mortem analyses to prevent recurrence.

Experiment with new tools and technologies to improve infrastructure performance and reliability while adhering to security standards. Identify and implement performance bottleneck improvements and optimize deployments for efficiency, speed, and scalability. Develop a deep understanding of site reliability trends and share knowledge with team members and management to foster a culture of continuous learning and improvement.

Qualifications

We are seeking a Senior Site Reliability Engineer with a minimum of 3 to 5 years of experience in the field. Proficiency in Linux, Oracle AI Database, Oracle Exadata, Query Language, and Script Programming is required.

This role demands strong problem-solving skills, the ability to independently manage work, monitor timelines, and adapt to changing resource or timeline shifts. Effective collaboration and partnership across teams are crucial for aligning expectations and achieving shared objectives. Building and maintaining a comprehensive understanding of business and stakeholder needs is key to fostering effective partnerships.

You will be expected to independently identify and address standard and non-standard issues, escalating complex problems as appropriate. Analyzing data from multiple sources to troubleshoot errors and contributing to knowledge sharing are vital aspects of this position. Continuous learning and skill development, staying current with industry trends, and seeking feedback are essential for success.

Furthermore, you will be tasked with developing ideas and recommending updates to enhance process efficiency and effectiveness. Seeking input from team members on alternative approaches and methods for improving work is highly encouraged. This role requires proactive engagement in continuous improvement initiatives within the team.

Essential Skills

LinuxOracle AI DatabaseOracle ExadataQuery LanguageScript Programming

Highlights

  • Actively hiring

More Details

RoleSenior Site Reliability Engineer
Employment TypeFull Time, Work from office

About the Company

Oracle logo

Oracle

IT Consulting

Senior Site Reliability Engineer at Oracle | SkillMX | SkillMX