Embedded Infrastructure Engineer, Chanakya
Sarvam AI
Sarvam AI
Join Sarvam, a pioneering force building India's sovereign AI platform. We are developing cutting-edge AI solutions across research, models, infrastructure, and applications, with a clear mission to make AI impactful for India. Partnering with leading enterprises and public institutions, Sarvam is backed by top venture capital firms and collaborates with major Indian brands.
As an Embedded Infrastructure Engineer, you will be instrumental in designing, constructing, and maintaining the robust data infrastructure essential for AI system deployments at client locations. You will collaborate closely with Embedded Data Scientists and Strategic Deployment Engineers to ensure the reliable and performant ingestion, storage, querying, and serving of terabyte-scale datasets to AI reasoning engines.
Your role involves building and operating sophisticated data platforms capable of handling substantial persistent data stores and high daily ingestion volumes across diverse data types like structured records, documents, imagery, audio, and geospatial data. You will be responsible for designing and maintaining critical components such as databases, object stores, ingestion pipelines, and processing layers.
Key contributions include making crucial decisions on storage architecture, indexing strategies, pipeline orchestration, and system performance, particularly within constrained, air-gapped, or operationally sensitive environments where standard cloud services may not be an option. You will assume full ownership of the reliability and performance of the infrastructure layer within your designated client accounts.
We are seeking individuals with 4–8 years of experience in data infrastructure, data engineering, platform engineering, or site reliability engineering, preferably within organizations managing significant data scales. A proven track record of managing multi-terabyte data stores, including personally building or operating systems handling over 10TB of persistent data with high-throughput ingestion, is essential.
Demonstrated expertise in at least two of the following systems is required: PostgreSQL, MongoDB, Elasticsearch, or ClickHouse, including deep knowledge of tuning, indexing, partitioning, and operational management. Experience building production data ingestion pipelines using frameworks like Apache Kafka, Apache Spark, Airflow, Flink, or dbt is also critical. Proficiency in Python and/or Go for production infrastructure tooling and automation, coupled with a solid understanding of storage systems (object storage, columnar formats) and familiarity with containerization (Docker, Kubernetes) and infrastructure-as-code (Terraform), are key qualifications.
Bonus points are awarded for experience with vector databases or embedding stores, as well as deploying and operating infrastructure in air-gapped, on-premise, or hybrid environments.
Sarvam AI
AI / Machine Learning