Principal Core Infrastructure Engineer
Oracle
Oracle
This role focuses on architecting and developing scalable, elastic distributed systems. You will be responsible for defining and enforcing scalability requirements, optimizing code and data paths for high-throughput workloads, and utilizing data plane platforms for large-scale data operations. The position involves designing fault-tolerant, in-service-upgradable systems that handle network unreliability while meeting Service Level Objectives (SLOs).
Key aspects include establishing performance metrics and telemetry, building proactive monitoring dashboards and alerts, and designing complex validation strategies. You will also diagnose and resolve production issues, mentor colleagues, and ensure operational readiness. Additionally, implementing robust security controls, managing compliance documentation, and developing Infrastructure as Code (IaC) for safe automated operations are critical components of this role.
Lead the development and architecture of scalable distributed system components, optimizing for high-throughput and large-scale data processing. Design fault-tolerant systems capable of in-service updates, managing network disruptions, and adhering to SLOs.
Define key performance indicators (KPIs) and telemetry to proactively monitor system health through custom dashboards and alerts. Implement strategies for system correctness and availability, including complex test scenarios and data replication techniques.
Proactively diagnose and resolve production issues, mentor peers, and ensure operational readiness. Implement robust security measures, execute remediation plans, and maintain compliance documentation. Develop IaC and automation for safe patching, updates, and rollbacks within change management frameworks.
We are seeking an engineer with 3 to 5+ years of experience in designing and building scalable, fault-tolerant distributed systems. A strong understanding of system scalability, reliability, and performance optimization is essential.
Proven ability to define and enforce scalability requirements, optimize code and data paths for high-throughput, hyper-scale workloads, and leverage data plane platforms is required. Experience in designing for redundancy, replication, failover, and handling network unreliability (load-shedding, throttling, rate-limiting) is crucial.
Demonstrated experience in establishing KPIs and telemetry, building dashboards and alerts, and designing complex validation scenarios (fault injection, brownouts) is necessary. The candidate should be adept at diagnosing and resolving production issues, mentoring peers, and ensuring operational readiness. Experience with implementing security controls, compliance documentation, and developing IaC and automation for change management is also required.
Oracle
IT Consulting