Principal Core Infrastructure Engineer
Oracle
Oracle
Join our team as a Principal Core Infrastructure Engineer, where you will lead the charge in architecting and developing highly scalable, elastic distributed systems. This role is pivotal in defining and enforcing scalability standards, optimizing performance for hyper-scale workloads, and leveraging data plane platforms for extensive data retrieval, storage, and processing.
You will design robust, fault-tolerant systems capable of in-service upgrades, employing strategies like redundancy and failover. We expect you to manage network unreliability through sophisticated load-shedding and throttling techniques, ensuring adherence to stringent Service Level Objectives (SLOs).
Key to this position is establishing comprehensive Key Performance Indicators (KPIs) and telemetry, building proactive monitoring dashboards and alerts, and implementing advanced validation techniques like fault injection. You will also be instrumental in diagnosing and resolving production issues, mentoring peers, and ensuring operational readiness. Implementing strong security controls, executing remediation plans, and maintaining compliance documentation are also core aspects of this role, alongside developing Infrastructure as Code (IaC) and automation for safe system updates and rollbacks.
Lead the design, development, and architecture of scalable distributed systems, focusing on horizontal and vertical scaling. Optimize code and systems for large-scale data processing and high-throughput hyper-scale environments.
Design fault-tolerant systems with built-in redundancy, replication, and automatic failover for in-service upgrades. Implement strategies to handle network partitions and unreliability, including load-shedding, throttling, and rate-limiting, ensuring compliance with SLOs.
Define and monitor system KPIs and telemetry through custom dashboards and alerts. Design and implement rigorous testing scenarios, including fault injection, to validate system correctness and durability through replication and synchronization techniques.
Proactively diagnose and resolve production issues, mentoring team members. Ensure operational readiness through robust design and implementation. Implement comprehensive security measures, execute remediation plans, and maintain compliance documentation.
Develop and maintain Infrastructure as Code (IaC) and automation for cloud infrastructure management. Create and follow change management plans for secure patching, updates, and rollbacks.
A strong background in designing and architecting scalable, elastic distributed systems is essential. Proven experience in optimizing code and data paths for high-throughput, hyper-scale workloads is required.
Expertise in building fault-tolerant, in-service-upgradable systems using redundancy, replication, and failover is expected. You should be adept at applying load-shedding, throttling, and rate-limiting to manage network unreliability while meeting SLOs.
Demonstrated ability in establishing KPIs and telemetry, building proactive dashboards and alerts, and designing complex validation scenarios (fault injection, brownouts) for correctness and durability is crucial. Experience in diagnosing and resolving production issues, mentoring peers, and ensuring operational readiness is a must.
Proficiency in implementing robust security controls, executing remediation, maintaining compliance documentation, and developing IaC and automation for safe patching, updates, and rollbacks within change management plans is also required.
Oracle
IT Consulting