Engineering Division - Production Runtime Experience - Vice President - Bengaluru

Goldman Sachs

7+ yrs Bengaluru Full Time Hybrid (office + remote)
Goldman Sachs logo
Posted : 1 week ago
Actively hiring

Job description

Join our Enterprise Technology Operations (ETO) team as a Vice President in the Production Runtime Experience (PRX) group. This role focuses on leveraging advanced Machine Learning and Generative AI to enhance the operational excellence of our large-scale production management services. We are dedicated to reducing operational risks through cutting-edge engineering, automation, and data science.

The PRX team applies software engineering and ML techniques to streamline monitoring, alerting, automation, and workflows. We build robust solutions for managing the firm's extensive compute infrastructure and application estate. By combining classical ML with agentic AI, we deliver reliable, explainable, and cost-efficient operations at scale.

Responsibilities

Lead the launch and implementation of Generative AI agentic solutions designed to significantly reduce the risk and cost associated with managing complex, large-scale production environments. You will tackle diverse production runtime challenges by developing sophisticated AI agents capable of diagnosing issues, reasoning through problems, and executing actions within production systems to boost productivity and resolve support-related matters.

Key responsibilities include: - Designing and building tool-calling agents with retrieval, structured reasoning, and secure action execution capabilities, adhering to MCP protocols and implementing robust guardrails for safety and compliance. - Developing an evaluation framework for LLMs and implementing retrieval pipelines, prompt synthesis, response validation, and self-correction loops for production operations. - Integrating agents with observability, incident management, and deployment systems to automate diagnostics, runbook execution, remediation, and post-incident summarization with full traceability. - Collaborating with production engineers and application teams to translate pain points into AI roadmaps, define objective functions, and deliver auditable, business-aligned outcomes. - Ensuring safety, reliability, and governance by building validator models, adversarial prompts, and policy checks, and enforcing deterministic fallbacks and rollback strategies. - Optimizing cost and latency through advanced techniques like prompt engineering, context management, caching, model routing, and distillation. - Building and maintaining a RAG pipeline, curating domain knowledge, validating data quality, and establishing feedback loops for knowledge freshness. - Driving design reviews, experiment rigor, and high-quality engineering practices, while mentoring peers on agent architectures and safe deployment patterns.

Qualifications

We seek a candidate with a Bachelor's degree in a computational field such as Computer Science, Applied Mathematics, or Engineering, with a strong preference for a Master's or PhD. A minimum of 7 years of experience as an applied data scientist or machine learning engineer is essential.

Essential technical skills include: - Over 7 years of software development experience in languages like Python, C/C++, Go, or Java, with a preference for extensive experience in building and maintaining large-scale Python applications. - More than 3 years of hands-on experience in designing, architecting, testing, and launching production ML systems, encompassing model deployment, serving, evaluation, monitoring, data processing, and fine-tuning. - Practical experience with Large Language Models (LLMs), including API integration, prompt engineering, fine-tuning/adaptation, and developing applications using RAG and tool-using agents (vector retrieval, function calling, secure tool execution). - A thorough understanding of various LLMs, both commercial and open-source (e.g., OpenAI, Gemini, Llama, Qwen, Claude), and their respective capabilities. - A solid grasp of applied statistics, core ML concepts, algorithms, and data structures to engineer efficient and reliable solutions. - Demonstrated strong analytical problem-solving skills, a sense of ownership, and urgency, coupled with the ability to articulate complex ideas clearly and collaborate effectively across global teams to achieve measurable business impact.

Preferred qualifications include proficiency in building and operating on cloud infrastructure, ideally AWS, including containerized services, serverless computing, data services, orchestration tools, model serving platforms, and infrastructure-as-code solutions.

Essential Skills

PythonC/C++GoJavaMachine LearningLarge Language Models (LLMs)RAGAPI IntegrationPrompt EngineeringFine-tuningVector RetrievalFunction CallingApplied StatisticsData StructuresProblem-solvingCommunication

Good to Have

AWSECS/EKSLambdaS3DynamoDBRedshiftStep FunctionsSageMakerTerraformCloudFormation

Highlights

  • Actively hiring

More Details

RoleEngineering Division - Production Runtime Experience - Vice President - Bengaluru
DepartmentAI / Machine Learning
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Goldman Sachs logo

Goldman Sachs

IT Consulting