Senior Site Reliability Engineer
BCG
BCG
Join Boston Consulting Group as a Senior Site Reliability Engineer and drive the engineering capability for organizational reliability. This role focuses on enhancing infrastructure, cloud, observability, automation, identity, security, and network operations.
You will apply engineering principles to reduce operational burden, improve system resilience, and integrate reliability and governance into workflows. The position involves shaping reliability delivery, building reusable components, and mentoring engineers.
We seek a seasoned practitioner comfortable in diverse domains, adept at balancing delivery with mentorship, and skilled in communicating technical trade-offs to all audiences.
Manage and enhance reliability engineering systems, including automation, pipelines, and observability tools.
Design and implement scalable engineering solutions to eliminate operational toil and embed reliability into delivery processes.
Contribute to defining engineering standards, patterns, and reusable frameworks for the SRE practice.
Lead engineering responses to complex incidents, drive systemic fixes, and foster post-incident learning.
Mentor junior engineers in reliability engineering, automation, observability, and SRE principles.
Foster collaboration across engineering, platform, and operations teams to integrate reliability through engineering controls.
Communicate engineering status, risks, and recommendations to senior stakeholders and leadership.
Prepare structured metrics for operational reviews, focusing on service health, pipeline performance, and automation progress.
Possess 5-8 years of experience in Site Reliability Engineering, Platform Engineering, or similar operational engineering fields.
Demonstrate strong hands-on expertise across cloud platforms (AWS or Azure), automation, observability, and CI/CD.
Showcase experience designing and implementing large-scale automation and reliability solutions.
Exhibit deep knowledge of at least one major cloud platform, including its networking, identity, and observability features.
Proficient with Infrastructure-as-Code tools like Terraform and CI/CD pipelines.
Strong scripting skills, particularly in Python.
Experience leading incident response and driving continuous improvements.
Excellent stakeholder engagement and technical communication abilities.
Deep hands-on experience with enterprise observability platforms such as Splunk or Datadog.
Proven ability to design telemetry pipelines, ingestion controls, and manage observability costs.
Demonstrated experience designing signals (SLIs, SLOs, synthetic checks, alerts) and triggering operational automation from them.
Experience driving SLO/SLI practices across multiple teams.
Extensive hands-on experience operating cloud infrastructure across at least two of AWS, Azure, or GCP.
Proven success in designing reusable IaC patterns and landing zone components.
Solid understanding of cloud networking, account management, identity primitives, and policy enforcement.
Experience in establishing cloud platform engineering standards and governance.
Deep hands-on experience with identity platforms like Entra ID and secrets management tools such as HashiCorp Vault.
Proven ability to design OIDC, workload identity, and dynamic credential patterns.
Experience promoting Zero Trust and least-privilege principles.
Deep hands-on experience with security tooling integrated into CI/CD pipelines.
Proven success in designing policy-as-code controls and secure-by-default patterns.
Experience driving secure engineering practices across teams.
Deep hands-on experience with hybrid and cloud network architectures.
Proven ability to design automated network controls using IaC.
BCG
IT Consulting