Data Engineer
Capgemini
Capgemini
Join Capgemini and be empowered to shape your career within a global, collaborative community. We help leading organizations unlock technology's potential to build a more sustainable and inclusive world. As a Data Engineer, you will architect, build, and maintain enterprise data platforms critical for AI-driven products and analytics.
You'll collaborate closely with Data Scientists, Architects, AI Engineers, and business stakeholders. Your work will involve developing robust data pipelines, digital twin platforms, knowledge graphs, and AI-ready data architectures. This role is central to transforming raw data into actionable insights and intelligent solutions.
Design and implement data source connectors for various enterprise applications and telemetry data. Build scalable ingestion, validation, and data contract frameworks for large-scale data onboarding. Develop capabilities to transform raw telemetry into incident-aware datasets and manage topology mapping.
Create and maintain Network Digital Twin and topology graph platforms. Build and manage Knowledge Graphs, including ontology frameworks and semantic data models. Develop reusable Feature Engineering Frameworks and Feature Stores for AI/ML applications.
Enable advanced data foundations for RAG, GraphRAG, Vector Databases, and AI Agent frameworks. Develop and optimize data pipelines for AI model lifecycle management. Implement robust metadata management, lineage tracking, observability, security, and governance frameworks.
Possess strong programming skills in Python and SQL, alongside a solid understanding of Data Modeling, ETL/ELT processes, and Data Architecture principles. Hands-on experience with Apache Spark, Spark Streaming, and Kafka is essential, as is expertise in building and managing Data Lakes.
Proficiency in AWS services such as S3, Glue, Lambda, and SageMaker is required. Knowledge of Machine Learning, Deep Learning, Predictive Analytics, and Statistical Modeling is crucial. Experience with Generative AI technologies, including LLMs, Prompt Engineering, RAG, and AI Agents, is highly valued.
Demonstrated expertise in Relational Databases, Graph Databases, and Vector Databases, coupled with experience in building APIs, Data Pipelines, and Batch/Real-Time Data Processing frameworks. A strong grasp of Data Quality, Metadata Management, Data Lineage, and Data Governance is necessary.
Capgemini
IT Consulting