Senior Core Infrastructure Engineer
Oracle
Oracle
Join a dynamic team focused on designing, implementing, and optimizing critical components within large-scale distributed systems. This role emphasizes scalability, resilience, and ease of operation. You will be instrumental in delivering new features, conducting rigorous load and performance tests, and leveraging advanced data platforms for high-volume data handling. Collaboration is key as you'll review peer implementations to ensure scalability compliance and contribute to building robust, fault-tolerant systems.
This position involves applying recovery-oriented principles, implementing retry mechanisms, circuit breakers, and timeouts to enhance system stability. You will proactively identify and resolve issues through comprehensive testing, alarms, dashboards, and telemetry. Additionally, you'll author runbooks, participate actively in incident response, and conduct root cause analyses. Implementing standard replication and synchronization, developing automation scripts for maintenance, and applying advanced security controls are also core aspects of this role, ensuring adherence to change, compliance, and documentation standards.
Key responsibilities include designing and developing scalable distributed system components, optimizing code for large-scale data processing, and ensuring scalability requirements are met. You will implement performance and load testing, and collaborate on building fault-tolerant systems with redundancy and automatic failover.
Enhance system reliability by applying recovery-oriented computing principles and implementing mechanisms like retries and circuit breakers. Proactively detect and address issues through tests, alarms, and dashboards. Support incident recovery by drafting and executing runbooks.
Diagnose and resolve system component issues, design automation scripts for troubleshooting, and participate in operational support and incident investigations. Apply advanced security measures, including encryption and access controls, to protect data and applications. Ensure cloud infrastructure compliance and maintain up-to-date documentation.
Track project timelines diligently, prioritizing and adjusting work as needed. Collaborate effectively across teams, aligning expectations and building strong partnerships by understanding business and stakeholder needs. Independently address standard and non-standard issues, escalating complex problems appropriately. Contribute to continuous learning by acquiring new skills and knowledge, and drive continuous improvement by recommending updates to enhance team processes and workflows.
We are seeking an experienced Core Infrastructure Engineer with 3 to 5+ years of experience in designing, implementing, and optimizing distributed systems. A strong understanding of scalability, resiliency, and operability is essential.
Proficiency in developing features, conducting load and performance tests, and utilizing data plane platforms for high-volume data processing is required. Experience in building fault-tolerant paths, applying recovery-oriented principles, and implementing retry mechanisms, circuit breakers, and timeouts is crucial.
Demonstrated ability to proactively detect and mitigate issues using tests, alarms, dashboards, and telemetry is expected. The role requires authoring runbooks, participating in incident response, and conducting root cause analyses. Experience with standard replication, synchronization, automation/IaC, and advanced security controls (encryption, access, remediation) is also necessary.
Candidates must possess strong problem-solving skills, the ability to track timelines with minimal supervision, and collaborate effectively across teams. A commitment to continuous learning and driving process improvements is highly valued. Fluency in English is mandatory. Please note, visa/work permit sponsorship is not available for this position.
Oracle
Technology