Senior Site Reliability Engineer
BCG
BCG
Join Boston Consulting Group as a Senior Site Reliability Engineer and spearhead the engineering capabilities for critical reliability areas. This role demands a strong engineering mindset to enhance resilience, reduce operational inefficiencies, and embed reliability into all aspects of delivery and operations across infrastructure, cloud, observability, automation, identity, security, and network domains.
You will champion engineering quality and consistency within your scope, contributing to broader engineering standards and influencing how reliability is perceived and implemented across the organization. This position involves building reusable solutions, mentoring fellow engineers, and providing expert technical guidance to a diverse group of stakeholders.
Drive continuous improvement of reliability engineering systems, including automation, pipelines, observability, and operational tools. Design and deploy scalable engineering solutions to eliminate operational toil and integrate reliability into delivery workflows. Shape robust engineering standards, patterns, and reusable frameworks for the SRE practice.
Lead engineering responses to significant incidents, ensuring systemic remediation and facilitating post-incident learning. Mentor and develop junior engineers in reliability engineering, automation, observability, and SRE principles. Foster collaboration with engineering, platform, and operations teams to integrate reliability and governance through engineering controls.
Communicate engineering progress, risks, and strategic recommendations effectively to senior leadership. Contribute structured metrics on service health, performance, automation coverage, and improvement initiatives to monthly operational reviews.
Bring 5-8 years of experience in Site Reliability Engineering, Platform Engineering, or similar operational engineering fields. Possess strong hands-on expertise across multiple SRE domains like cloud, automation, observability, and CI/CD. Demonstrate a track record of designing and implementing large-scale automation and reliability solutions.
Exhibit deep knowledge of at least one major cloud platform (AWS or Azure), including its networking, identity, and observability features. Experience with Infrastructure-as-Code tools like Terraform and CI/CD pipelines is essential. Proficiency in scripting, particularly Python, is required.
Experience leading incident response and driving systemic improvements is crucial. Strong stakeholder management and technical communication abilities are necessary. Deep experience with enterprise observability platforms (e.g., Splunk, Datadog) and proven success in designing telemetry pipelines, ingestion controls, and managing observability costs are expected.
Demonstrated experience designing SLIs, SLOs, synthetic checks, alerts, and triggering automated operations from these signals is required. You should have experience driving SLO/SLI practices across teams. Extensive hands-on experience operating cloud infrastructure across at least two cloud providers (AWS, Azure, GCP, Alibaba Cloud) is a must.
Proven ability to design reusable IaC patterns and landing zone components is key. Strong grasp of cloud networking, account management, identity primitives, and policy enforcement across providers is needed. Experience driving cloud platform engineering standards and governance across multiple teams is highly valued.
Deep hands-on experience with identity platforms (e.g., Entra ID) and secrets management (e.g., HashiCorp Vault) is essential. Designing OIDC, workload identity, and dynamic credential patterns is expected. Experience driving Zero Trust and least-privilege adoption across teams is a plus.
Deep hands-on experience with security tooling integrated into CI/CD pipelines is required. Proven experience designing policy-as-code controls and secure-by-default patterns is important. Driving secure engineering adoption across teams is desirable.
Deep hands-on experience with hybrid and cloud network architectures is expected. Proven experience designing automated network controls through IaC is necessary. Driving Zero Trust segmentation and network observability adoption is a strong advantage.
BCG
Management Consulting