Senior Site Reliability Engineer
Microsoft
Microsoft
Join Microsoft's Azure Data engineering team as a Senior Site Reliability Engineer and shape the future of data analytics and AI.
This role focuses on building and maintaining mission-critical, AI-enabled operational databases, specifically contributing to Azure DocumentDB, an open-source PostgreSQL-based NoSQL database. You will leverage your expertise in software development and scalable system design to ensure unparalleled service stability, performance, and customer satisfaction.
Embrace Microsoft's culture of inclusivity and innovation, where diverse perspectives are actively sought and valued to better serve our global customer base.
Drive the design and implementation of comprehensive telemetry, alerting, and self-healing automation to enhance service health and reliability.
Participate in on-call rotations, owning and resolving service issues with a focus on communication, learning, and knowledge sharing.
Collaborate with customers and internal teams, engaging in deep technical discussions with product engineering and management to advance service capabilities.
Take ownership of service availability, performance, and supportability metrics, while also creating detailed technical documentation.
Possess a minimum of 6 years of experience in writing tools, scripting (Powershell, Python), and programming (C++, C#), with a proven track record of managing software in production environments.
Demonstrate 6+ years of experience in troubleshooting and debugging distributed services using telemetry analysis (KQL preferred), with expertise across network, hardware, and service layers, including code optimization.
Exhibit strong verbal and written communication skills, with a preferred understanding of distributed systems and networking concepts.
Microsoft Corporation
Information Technology & Services