Senior Core Infrastructure Engineer
Oracle
Oracle
Join our team as a Senior Core Infrastructure Engineer and play a pivotal role in designing, implementing, and optimizing distributed systems with a strong focus on scalability, resilience, and operational excellence. This role involves developing features, conducting load and performance testing, and utilizing advanced data plane platforms for high-volume data processing. You will also be instrumental in building robust fault-tolerant systems, applying recovery-oriented principles, and ensuring proactive issue detection and mitigation through comprehensive testing and monitoring.
Your responsibilities will extend to authoring runbooks, participating in incident response and root cause analysis, and implementing advanced security controls. You'll also develop automation for troubleshooting and maintenance, ensuring all changes adhere to strict compliance and documentation standards. This is an opportunity to contribute to cutting-edge cloud solutions and impact billions of lives.
- Design and implement components for distributed systems, focusing on horizontal and vertical scalability, and leveraging distributed state management tools. - Optimize code and systems for large-scale data processing and retrieval. - Implement and review scalability requirements for assigned components. - Develop and execute performance and load tests. - Build fault-tolerant components with redundancy, replication, and automatic failover mechanisms. - Apply recovery-oriented computing principles to handle service disruptions effectively. - Implement retry mechanisms, circuit breakers, and timeouts to manage network unreliability. - Configure tests and alarms to proactively detect and resolve issues. - Draft and execute runbooks and operational procedures for incident recovery. - Create and customize dashboards, telemetry systems, and alerting for component health monitoring. - Implement functional requirements and testing for assigned features. - Conduct fault-injection and brown-out tests to evaluate system correctness. - Implement standard data replication and synchronization techniques. - Diagnose, debug, and resolve issues in system components. - Design and implement automation scripts and tooling for troubleshooting. - Participate in operational support rotations, incident responses, and root cause investigations. - Apply advanced security measures, including encryption and access controls. - Develop and implement remediation plans for continuous security improvement. - Ensure cloud infrastructure compliance with industry standards and maintain up-to-date documentation. - Maintain automation scripts and tools, including Infrastructure as Code (IaC). - Adhere to change management plans for patching, updating, and rolling back applications.
- Bachelor's degree or equivalent practical experience. - 3-5+ years of experience in designing, implementing, and optimizing distributed systems. - Proven experience with scalability, resiliency, and operability principles. - Strong understanding of data plane platforms and distributed state tools. - Experience with fault tolerance, redundancy, replication, and automatic failover. - Proficiency in applying recovery-oriented computing principles. - Familiarity with implementing retries, circuit breakers, and timeouts. - Experience with proactive issue detection using tests, alarms, dashboards, and telemetry. - Ability to author runbooks and participate in incident response and RCAs. - Experience developing automation and Infrastructure as Code (IaC). - Knowledge of advanced security controls such as encryption and access management. - Understanding of change management, compliance, and documentation standards. - Excellent problem-solving and analytical skills. - Ability to collaborate effectively across teams and build strong partnerships. - Commitment to continuous learning and staying current with industry trends. - Strong verbal and written communication skills in English.
Oracle
IT Consulting