Software Engineer II
Microsoft
Microsoft
Join our AI Knowledge engineering team to build cutting-edge platforms for next-generation observability, operational intelligence, and automation for Azure AI services. This role focuses on enhancing the reliability and efficiency of global-scale AI services, empowering engineers to understand service health, diagnose issues, automate workflows, and improve incident resolution.
You'll work at the intersection of AI, cloud infrastructure, and distributed systems, contributing to how large-scale AI services are developed and operated. Our team values a customer-focused, evidence-based approach, ownership, continuous learning, and a people-first philosophy to sustainably manage 24/7 cloud services.
Design, implement, test, and maintain critical components of monitoring, alerting, telemetry, diagnostics, and operational intelligence platforms for large-scale Azure services. Develop automation for improved incident detection, enrichment, triage, and resolution. Create dashboards and analytics for actionable insights into service health and reliability. Contribute to scalable logging, metrics, tracing, and diagnostics infrastructure. Build deployment pipelines and release engineering capabilities for safe and efficient rollouts.
Investigate production issues using telemetry and diagnostics, implementing durable fixes. Participate in DevOps and live site operations to ensure service quality and availability. Leverage AI-assisted tools to boost development, testing, and operational efficiency. Collaborate across teams to enhance service supportability and customer experience. Adhere to best practices in security, privacy, quality, reliability, accessibility, and inclusive product development.
A Bachelor's Degree in Computer Science, Engineering, or a related technical field, coupled with at least 3 years of professional software engineering experience, is required. Proficiency in modern programming languages such as C#, C++, Java, or Python is essential. You should have hands-on experience building, testing, deploying, and supporting production software, cloud services, distributed systems, infrastructure platforms, or operational tooling. Familiarity with telemetry, monitoring, alerting, logging, metrics, tracing, diagnostics, automation, or incident-management systems is crucial.
A strong understanding of distributed systems fundamentals, cloud platforms, or service operations is necessary. We seek individuals with excellent problem-solving, debugging, analytical, and communication skills, capable of independently owning features from design through production support. Effectiveness in collaborating with diverse groups, embracing feedback, demonstrating a bias for action, and navigating ambiguity are key attributes.
Microsoft Corporation
IT Consulting