Senior Backend Engineer, Vision

Sarvam AI

5–6 yrs Bengaluru Full Time Work from office
Sarvam AI logo
Posted : yesterday
Actively hiring

Job description

Join Sarvam, a pioneering team dedicated to building India's sovereign AI platform. We are creating foundational AI models, infrastructure, and applications with a unique focus on making AI work for India. Our mission involves partnering with leading enterprises and public institutions, backed by top-tier venture capital firms and collaborating with major Indian brands.

This role is for a Senior Backend Engineer specializing in Vision, where you will take ownership of the architecture for our vision model serving harness. This critical system ensures frontier-grade extraction quality from our in-house models at a national scale, all while adhering to strict cost and latency budgets. You will navigate complex trade-offs between accuracy, latency, and cost per page, making key decisions on multi-pass inference, model routing, verification strategies, batching, and GPU utilization.

Furthermore, you will establish the benchmark for reliability. The pipelines you architect will process sensitive documents essential for KYC, loan underwriting, claims, and contracts. Durability, idempotency, and graceful degradation are not just goals but foundational requirements. The architecture you design will shape the future development of our team's initiatives.

Responsibilities

Own the complete end-to-end architecture for our OCR and extraction serving harness, encompassing the API layer, orchestration, inference, post-processing, and delivery mechanisms.

Design and implement a sophisticated accuracy harness, integrating multi-pass extraction, ensembling, cross-verification, schema-constrained decoding, confidence calibration, and targeted reruns, validating its effectiveness against evaluation datasets.

Architect robust and resumable document workflows using Temporal, managing fan-out across pages, partial failure recovery, exactly-once side effects, and long-running jobs.

Lead the inference serving layer in collaboration with infrastructure teams, defining batching strategies, managing GPU pools, implementing autoscaling based on real-time signals, and controlling queue depth and admission.

Proactively drive down latency, throughput, and unit economics through rigorous profiling, measurement, and defense of cost-per-page targets as volume increases.

Develop the foundational observability for the system, including distributed tracing, per-stage cost and quality metrics, Service Level Objectives (SLOs), alerting, and post-incident analysis.

Design for multi-tenancy, ensuring tenant isolation, rate limiting, and fair scheduling for enterprise customers with diverse load patterns.

Provide support for on-premises and constrained deployments, ensuring the entire harness functions within customer environments.

Set technical direction, elevate standards through thorough design reviews, and mentor junior engineers.

Qualifications

Possess 5-6+ years of experience in backend engineering, with significant hands-on experience operating high-throughput production systems and participating in on-call rotations.

Demonstrate deep proficiency in Go and/or Python, coupled with the judgment to apply them effectively.

Exhibit strong expertise in distributed systems design, including queues, workflow orchestration, idempotency, backpressure, retry and timeout semantics, consistency trade-offs, and graceful degradation.

Have production experience with Temporal or a comparable durable execution engine for critical workflows.

Show practical experience with Kubernetes in production, covering autoscaling, resource management, rollouts, and debugging under load; experience with GPU workload scheduling is highly advantageous.

Possess demonstrable experience serving ML or LLM inference in production environments, including batching, caching, model versioning, A/B rollouts, and latency budgeting.

Apply rigor to observability and reliability practices, with a track record of designing SLOs, managing incidents, and implementing corrective actions.

Apply hard-won cost intuition, having successfully reduced system costs without compromising quality.

Essential Skills

GoPythonTemporalRESTKubernetesPostgreSQLRedisObject StorageOpenTelemetryDistributed SystemsML Inference ServingObservabilityCost Optimization

Good to Have

OCRIDPDocument AITextractAzure DIGPU InferencevLLMTensorRT-LLMTritonSGLangRay ServeBFSIHealthcare ComplianceData ResidencyOn-prem DeploymentAir-gapped Deployment

Highlights

  • Actively hiring

More Details

RoleSenior Backend Engineer, Vision
IndustryAI / Machine Learning
DepartmentSoftware Development
Employment TypeFull Time, Work from office

About the Company

Sarvam AI logo

Sarvam AI

AI / Machine Learning

Senior Backend Engineer, Vision at Sarvam AI | SkillMX | SkillMX