Lead Principal Core Infrastructure Engineer
Oracle
Oracle
As a Lead Principal Core Infrastructure Engineer, you will architect and guide the development of highly scalable, interdependent distributed systems. This role involves identifying and resolving performance bottlenecks, defining scalability requirements in collaboration with stakeholders, and designing elastic, high-impact systems. You will drive innovation within data plane platforms and ensure the fault tolerance and upgradability of systems.
Your expertise will be crucial in optimizing resilience mechanisms, setting availability and durability standards, and establishing KPIs and advanced telemetry. You will also apply formal verification for complex features and develop robust replication strategies. Advising on complex production issues, setting operational readiness standards, and directing incident response are key aspects of this position. Furthermore, you will architect advanced security controls, drive remediation, and deliver enterprise-level automation for safe, automated system updates and rollbacks.
Architect and mentor teams on the design of highly scalable, interdependent distributed systems, ensuring optimal performance and elasticity.
Lead the identification and resolution of performance bottlenecks, collaborating with stakeholders to define system scalability requirements.
Drive innovation in data plane platforms and ensure systems meet non-functional scalability requirements, anticipating future business needs.
Design and oversee fault-tolerant, in-service-upgradable systems, optimizing resilience mechanisms like load-shedding and rate-limiting.
Define KPIs, telemetry, and advanced alerting mechanisms to ensure system health and reliability, applying formal verification techniques where necessary.
Architect advanced security measures, guide remediation efforts, and ensure compliance with industry standards and regulations.
Develop enterprise-level automation (IaC) and change management strategies for safe, automated patching, updates, and rollbacks.
Extensive experience, typically 6 to 10+ years, in leading the architecture and development of complex, distributed systems.
Demonstrated expertise in system design, scalability, performance optimization, and reliability engineering.
Proven ability to mentor teams and influence cross-functional leaders.
Experience with fault-tolerant designs, resilience mechanisms, and setting Service Level Objectives (SLOs).
Proficiency in defining KPIs, telemetry, and implementing advanced monitoring solutions.
Strong understanding of security controls, compliance, and Infrastructure as Code (IaC) principles.
Excellent problem-solving and incident management skills.
Ability to communicate effectively in English.
Oracle
IT Consulting