Global IT Platform Engineer Director
BCG
BCG
Join Boston Consulting Group as a Senior Site Reliability Engineer and lead the engineering capability for critical reliability areas across our organization. This role is pivotal in enhancing resilience, reducing operational overhead, and embedding reliability principles into our delivery and operational workflows. You will apply engineering expertise across infrastructure, cloud, observability, automation, identity, and network operations.
We are seeking a seasoned practitioner adept at navigating multiple technical domains. Your ability to balance project delivery with mentorship, and clearly communicate complex engineering trade-offs to diverse audiences, will be key to success. You will contribute significantly to our broader engineering standards and influence how reliability is architected and maintained throughout BCG.
Drive continuous improvement of reliability engineering systems, including automation, pipelines, observability tools, and operational support mechanisms. Design and deploy engineering solutions that minimize repetitive operational tasks at scale and integrate reliability into development lifecycles.
Shape the firm's engineering standards, establish reusable patterns, and foster a culture of reliability within the SRE practice. Take the lead in engineering responses to significant incidents, spearheading systemic fixes and contributing to post-incident analysis. Mentor and guide junior engineers in reliability engineering, automation, observability, and SRE best practices.
Facilitate seamless collaboration between engineering, platform, and operations teams to embed reliability and governance through robust engineering controls. Provide clear updates on engineering progress, risks, and strategic recommendations to senior leadership and stakeholders. Contribute valuable insights to monthly operational reviews, presenting structured metrics on service health, performance, automation adoption, and improvement initiatives.
A minimum of 5 to 8 years of progressive experience in Site Reliability Engineering, Platform Engineering, or similar operational engineering fields is required. Demonstrated hands-on expertise across multiple SRE domains, including cloud environments, automation strategies, observability solutions, and CI/CD practices, is essential.
Proven success in designing and implementing large-scale automation and reliability solutions. In-depth knowledge of at least one major cloud platform (AWS or Azure), encompassing its networking, identity management, and observability features. Proficiency with Infrastructure-as-Code tools such as Terraform and experience building and managing CI/CD pipelines.
Strong scripting capabilities, particularly in Python. Experience leading incident response efforts and driving significant, systemic improvements. Excellent stakeholder management and technical communication skills are crucial. Deep familiarity with enterprise observability platforms like Splunk or Datadog, alongside proven experience in designing telemetry pipelines, ingestion controls, and cost management for observability.
Experience designing and implementing SLIs, SLOs, synthetic checks, and alerts, coupled with automation triggered by these signals. A track record of promoting SLO/SLI practices across multiple teams. Hands-on experience operating cloud infrastructure across at least two major cloud providers (AWS, Azure, GCP, or Alibaba Cloud). Expertise in designing reusable IaC patterns and landing zone architectures.
Solid understanding of cloud networking, account management, identity primitives, and policy enforcement across various cloud providers. Experience driving cloud platform engineering standards and governance initiatives. Profound hands-on experience with identity platforms (e.g., Entra ID) and secrets management solutions (e.g., HashiCorp Vault). Proven ability to design OIDC, workload identity, and dynamic credential patterns.
Experience in promoting Zero Trust and least-privilege adoption. Hands-on experience with security tools integrated into CI/CD pipelines. Proven ability to design policy-as-code controls and secure-by-default configurations. Experience driving secure engineering practices across teams. Deep expertise in hybrid and cloud network architectures, including designing automated network controls via IaC. Experience fostering Zero Trust segmentation and network observability.
BCG
Management Consulting