Data Engineer

Capgemini

Fresher Gurugram Full Time Work from office
Capgemini logo
Posted : today
Actively hiring

Job description

Join Capgemini and be empowered to shape your career within a global, collaborative community. We help leading organizations unlock technology's potential to build a more sustainable and inclusive world. As a Data Engineer, you will architect, build, and maintain enterprise data platforms critical for AI-driven products and analytics.

You'll collaborate closely with Data Scientists, Architects, AI Engineers, and business stakeholders. Your work will involve developing robust data pipelines, digital twin platforms, knowledge graphs, and AI-ready data architectures. This role is central to transforming raw data into actionable insights and intelligent solutions.

Responsibilities

Design and implement data source connectors for various enterprise applications and telemetry data. Build scalable ingestion, validation, and data contract frameworks for large-scale data onboarding. Develop capabilities to transform raw telemetry into incident-aware datasets and manage topology mapping.

Create and maintain Network Digital Twin and topology graph platforms. Build and manage Knowledge Graphs, including ontology frameworks and semantic data models. Develop reusable Feature Engineering Frameworks and Feature Stores for AI/ML applications.

Enable advanced data foundations for RAG, GraphRAG, Vector Databases, and AI Agent frameworks. Develop and optimize data pipelines for AI model lifecycle management. Implement robust metadata management, lineage tracking, observability, security, and governance frameworks.

Qualifications

Possess strong programming skills in Python and SQL, alongside a solid understanding of Data Modeling, ETL/ELT processes, and Data Architecture principles. Hands-on experience with Apache Spark, Spark Streaming, and Kafka is essential, as is expertise in building and managing Data Lakes.

Proficiency in AWS services such as S3, Glue, Lambda, and SageMaker is required. Knowledge of Machine Learning, Deep Learning, Predictive Analytics, and Statistical Modeling is crucial. Experience with Generative AI technologies, including LLMs, Prompt Engineering, RAG, and AI Agents, is highly valued.

Demonstrated expertise in Relational Databases, Graph Databases, and Vector Databases, coupled with experience in building APIs, Data Pipelines, and Batch/Real-Time Data Processing frameworks. A strong grasp of Data Quality, Metadata Management, Data Lineage, and Data Governance is necessary.

Essential Skills

PythonSQLData ModelingETLELTData ArchitectureApache SparkSpark StreamingKafkaData LakesAWSS3GlueLambdaSageMakerMachine LearningDeep LearningPredictive AnalyticsStatistical ModelingGenerative AILLMsPrompt EngineeringRAGAI AgentsRelational DatabasesGraph DatabasesVector DatabasesAPIsData PipelinesBatch ProcessingReal-Time Data ProcessingData QualityMetadata ManagementData LineageData GovernanceAWS GlueAzure Data FactorySageMaker PipelinesDigital TwinGraphRAGMLOpsLLMOpsDataOpsKnowledge ManagementAnalytical SkillsDebuggingProblem-Solving

Good to Have

JavaScriptReactAngular

Highlights

  • Actively hiring

More Details

RoleData Engineer
DepartmentData Engineering
Employment TypeFull Time, Work from office

About the Company

Capgemini logo

Capgemini

IT Consulting

Data Engineer at Capgemini | SkillMX | SkillMX