ML Ops Engineer, Chanakya
Sarvam AI
Sarvam AI
Sarvam is building the foundation for sovereign AI in India, creating a comprehensive AI platform spanning research, models, infrastructure, and applications. Our singular focus is to make AI truly effective for India's unique needs. We partner with leading enterprises and public institutions, supported by prominent investors like Lightspeed, Peak XV, and Khosla Ventures, and collaborate with major Indian brands such as Tata Capital, SBI Life, CRED, IDFC, and LIC.
This role as an MLOps Engineer is crucial for managing the entire lifecycle of AI models across all defense and strategic sector deployments. You will be responsible for serving infrastructure, monitoring, evaluation pipelines, and environment management, ensuring our systems are consistently operational, accurate, and auditable. You'll bridge the gap between field operations and product development, maintaining uncompromising standards where model failures represent operational risks.
Key responsibilities include designing and managing model serving infrastructure for both on-premise and cloud environments. You will develop and maintain robust CI/CD pipelines for seamless model updates, rollbacks, and deployment gating. Proactive monitoring of model performance in production—tracking latency, accuracy drift, throughput, and failure modes—is essential, along with building systems to identify issues before our clients do.
Furthermore, you will construct comprehensive evaluation infrastructure, including A/B testing and model comparison tools for both field and lab use. Managing containerized model serving in challenging environments like air-gapped or edge systems is also a core function. Collaboration with Data Scientists on evaluation pipelines, owning the underlying infrastructure, and creating detailed runbooks for field deployment engineers are critical. Incident response for model-layer failures and ensuring the operational health of all active deployments will be your domain.
We are seeking candidates with 3-5 years of experience in ML engineering or MLOps, specifically with at least one production LLM or ML system in continuous operation. Deep expertise in model serving technologies such as vLLM, TGI, or Triton Inference Server, along with experience in quantized model formats (GGUF, AWQ, GPTQ), is required.
Candidates should have experience fine-tuning and adapting models in constrained, on-premise, or air-gapped environments, including managing associated data pipelines and compute limitations. Proficiency in containerization using Docker and Kubernetes, or lightweight alternatives like K3s/K0s for constrained environments, is essential. Experience with heterogeneous hardware configurations is a plus.
Familiarity with monitoring and observability tools like Prometheus and Grafana, the ability to build custom evaluation dashboards, and strong Python fluency are necessary. Experience with fine-tuning workflows, model evaluation frameworks, and CI/CD tools for ML pipelines (e.g., GitHub Actions, ArgoCD, DVC) is also vital. We value individuals who can keep production ML systems running smoothly, proactively build systems to prevent failures, and create documentation that is actively used. Uptime, correctness, and operational reliability are paramount.
Sarvam AI
AI / Machine Learning