Join a dynamic team focused on Site Reliability Engineering (SRE), merging software and systems engineering to build and maintain highly scalable, distributed, and fault-tolerant systems. This role ensures Google Cloud's services maintain optimal reliability and uptime, meeting customer needs while driving continuous improvement. You will be instrumental in monitoring system capacity and performance.
Our SRE team excels at tackling the unique scaling challenges of Google Cloud through expert coding, algorithmic problem-solving, and large-scale system design. We foster a culture of curiosity, collaboration, and innovation in a supportive, blame-free environment. Opportunities abound for self-directed work on impactful projects, coupled with mentorship for professional growth.
Design, develop, and implement projects aimed at enhancing the reliability of mission-critical enterprise applications.
Collaborate with cross-functional teams to troubleshoot and resolve complex system issues.
Contribute to the development of automation tools and infrastructure to improve system efficiency and reduce manual intervention.
Participate in on-call rotations to ensure the availability and performance of production systems.
A Bachelor's degree in Computer Science, a related discipline, or equivalent practical experience is required.
Possess at least 1 year of hands-on experience in software development using one or more programming languages.
Demonstrate 1 year of experience working with data structures and algorithms.
In-person interviews are typically part of the selection process for this position.
IT Consulting