Principal Systems Engineer
Oracle
Oracle
Join the AI Infra Operations team as a Principal Systems Engineer, focusing on GPU infrastructure within Oracle Cloud Infrastructure (OCI). This role is pivotal in leading the development and ongoing maintenance of automation and operational tools for GPU fleets across various regions, ensuring robust high availability.
You will be instrumental in collaborating with engineering and operations teams to architect and implement comprehensive automation, observability, and reliability solutions. The core objective is to continuously enhance GPU operations and support the advanced needs of AI infrastructure.
This is an individual contributor role within Technology Operations, requiring significant experience to drive operational excellence and technical leadership.
Develop and maintain critical automation and operational tooling for OCI's GPU infrastructure across multiple geographic regions.
Foster collaboration with software engineers, hardware teams, and operations partners to ensure the high availability of the GPU fleet.
Build and refine monitoring, alerting, and diagnostic systems for GPU fleet health, performance, capacity, and utilization, leveraging tools like Grafana.
Act as the senior escalation point for complex GPU host and repair challenges, providing expert-level troubleshooting.
Participate actively in incident response and root-cause analysis to resolve blockers impacting GPU capacity, availability, and regional deployments.
Drive continuous improvements in AI2 Ops processes, GPU fleet automation, and the readiness of OCI region builds.
Engage in on-call rotations to provide essential support for critical infrastructure issues.
Create and maintain comprehensive documentation, including operational procedures, automation workflows, troubleshooting guides, and runbooks.
Develop and enhance AI agents, overseeing their safe rollout, execution, and monitoring.
Mentor and guide junior engineers, imparting operational best practices, offering senior technical support, and championing ongoing improvement initiatives.
A Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience, is required.
Possess a minimum of 7 years in software operations or infrastructure automation, with advanced proficiency in Python and Bash scripting.
Demonstrate expert-level Linux administration skills, particularly with Ubuntu and Oracle Linux, in large-scale production environments.
Exhibit a strong understanding of distributed systems, encompassing peer-to-peer, node-to-node, and service-to-service communication paradigms.
Have substantial data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery strategies.
Showcase excellent problem-solving and troubleshooting capabilities.
Possess outstanding communication and teamwork skills to collaborate effectively.
Experience with observability tooling, including metrics, logging, dashboards, and alerting, is essential.
Familiarity with AI agents and associated tooling is expected.
Prior experience leading on-call operations and incident response is a strong asset.
Oracle
Technology