Senior Network Developer
Oracle
Oracle
Join the AI Infrastructure - Network Operations team at Oracle Cloud Infrastructure (OCI) to support and advance the RDMA/RoCE/InfiniBand network fabrics critical for OCI's leading AI and HPC customers. These high-performance networks are the backbone of OCI's AI, GPU, and HPC services, powering major generative AI industry players. If your AI workload runs on OCI, our team ensures the underlying RDMA network operates seamlessly.
As a Network Operations Engineer, you will be instrumental in the design, deployment, and ongoing operations of OCI's extensive global cloud computing environment. Your primary focus will be on the operation and support of RDMA/RoCE/InfiniBand network fabrics and systems, leveraging a deep understanding of networking principles and advanced automation skills to maintain a robust production environment. This role involves supporting hundreds of thousands of network devices across a global footprint, connecting millions of servers via dedicated backbone infrastructure and the internet.
Key contributions include designing, operating, validating, and scaling sophisticated network fabrics for large-scale cloud, AI, and data center environments.
You will develop automation scripts and tools to enhance network testing, operations, deployment, and troubleshooting efficiency.
Build and refine telemetry, dashboards, alerting, and monitoring systems to boost network health, reliability, and Service Level Objective (SLO) performance.
Formulate test strategies, lead pre-production validation, and conduct thorough root cause analysis (RCA) for network issues and changes.
Analyze network performance metrics such as capacity, latency, throughput, and packet loss to pinpoint issues and facilitate infrastructure growth.
Participate actively in incident response and operational support, resolving complex production and customer-facing issues.
Collaborate with engineering teams, vendors, and stakeholders on network architecture, deployment standards, and operational readiness.
Identify design and operational risks, implement effective mitigations, and drive continuous improvements in network processes and reliability.
Provide technical guidance and mentorship to fellow engineers, contributing to architecture, roadmap development, and engineering best practices.
Manage priorities and deliverables independently while fostering cross-team collaboration to achieve shared objectives.
Serve as a specialized Tier 2 escalation point for network incidents, leading root cause analysis, corrective actions, and long-term reliability enhancements.
We are seeking candidates with a minimum of 6 years of experience and a Bachelor's degree in Computer Science, Electrical Engineering, or a related field; a Master's degree is preferred.
Demonstrate strong expertise in large-scale network operations, design, and troubleshooting.
Proficiency in routing and switching technologies is essential, including BGP, OSPF, EVPN-VXLAN, MPLS, and data center networking.
Prior experience with RDMA/RoCE/InfiniBand technologies is highly advantageous.
Experience with network automation using Python, Ansible, APIs, or similar tools, including building AI agents for automation, is required.
A solid understanding of observability, monitoring, telemetry, and incident management principles is necessary.
Experience working within cloud infrastructure, hyperscale environments, or large-scale distributed systems is expected.
Proven ability to lead technical projects and influence outcomes across multiple teams is crucial.
Excellent written and verbal communication skills are essential for effective collaboration with engineering, operations, and leadership teams.
Oracle
Cloud Computing