Capgemini

Data Engineer

Capgemini • $86K — $127K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of data engineering experience, including 2+ years with Databricks
  • Strong proficiency in Python for data transformation and automation
  • Expert-level SQL skills for schema design and analytical workloads
  • Fluency in GCP services like BigQuery and Cloud Storage
  • Experience with privacy-preserving data exchange technologies
  • Knowledge of PII governance in regulated industries

Responsibilities

  • Design and operate the full medallion stack on Databricks
  • Implement Delta Live Tables and Lakeflow pipelines for data ingestion
  • Enforce schema evolution at the Bronze boundary to prevent downstream failures
  • Build unstructured processing tiers using Databricks AI Functions
  • Manage Unity Catalog for governance and access control
  • Establish data-quality frameworks to minimize manual checks
  • Oversee CI/CD and cost management for the Databricks environment

Benefits

  • Paid time off ranging from 12-25 days based on employee grade
  • Comprehensive medical, dental, and vision coverage
  • Retirement savings plans like 401(k) or RRSP
  • Life and disability insurance
  • Employee assistance programs
  • Additional benefits as per local policy
Full Job Description
Job Description What You Will Own Medallion Architecture - Bronze through Gold • Design, build, and operate the full medallion stack on Databricks: raw landing (Bronze), identity-resolved and feature-engineered Silver, and the governed serving layer (Gold) that the Helm agent layer queries on every conversational turn • Implement Delta Live Tables (DLT) and Lakeflow pipelines for batch and structured-streaming ingestion of transaction signals, behavioral events, and third-party identity data • Own schema enforcement and evolution at the Bronze boundary so upstream change surfaces as a controlled event, not a downstream failure • Build the unstructured processing tier (Bronze 14 Silver) using Databricks AI Functions or ai_query to convert PDFs, brand guidelines, compliance rule sets, and creative assets into governed Silver tables Unity Catalog & Governance • Own Unity Catalog as the governance layer: access control, lineage tracking, and per-client tenant isolation across Helm's financial institution client base • Implement PII governance controls at the pipeline layer - redacted ID egress, consent signal propagation, and guardrail alignment to GLBA, Fair Lending, and UDAAP • Build and maintain data-quality expectations frameworks: reconciliation, deduplication, and lineage checks that reduce manual investigation cycles Serving Layer & Platform Operations • Build and operate the Gold serving layer behind Helm's twelve analytical models: MTA, MMM, incrementality, LTV:CAC, reach/frequency, cohorts, cobrand overlap, channel viability, card activation, product holdings, scenario reallocation, and channel universe • Meet the serving contract the agentic layer depends on: fixed signal shapes, live query on every turn, predictable latency, and per-client isolation • Own CI/CD, infrastructure-as-code, cost management, and pipeline observability for the Databricks environment Required Qualifications • 5-7 years of data engineering experience with 2+ years of production Databricks work: Delta Lake, Unity Catalog, DLT/Lakeflow, structured streaming, and SQL warehouses • Strong Python for transformation logic, data-quality automation, and pipeline orchestration • Expert-level SQL; comfort designing schemas for both analytical and serving workloads • GCP fluency: BigQuery, Cloud Storage, Cloud Run, IAM - Helm runs on GCP and Databricks runs within it • Hands-on experience with at least one clean room or privacy-preserving data exchange technology • PII governance knowledge in a regulated industry context: data residency, consent frameworks, GLBA, Fair Lending, UDAAP Preferred Qualifications • Databricks AI Functions (ai_query) or Vertex AI used inside pipelines for document extraction and normalization • dbt or a comparable transformation and lineage framework • Identity resolution concepts at the pipeline layer: deterministic matching, RampID or UID2 linkage, match-rate monitoring • LiveRamp, AMC, GMP, or XMi clean room connector experience • Familiarity with LangSmith or LangGraph as a data consumer - understanding what the agentic layer needs from the serving layer The base compensation range for this role in the posted location is 86129 - 127189 Capgemini provides compensation range information in accordance with applicable national, state, provincial, and local pay transparency laws. The base compensation range listed for this position reflects the minimum and maximum target compensation Capgemini, in good faith, believes it may pay for the role at the time of this posting. This range may be subject to change as permitted by law. The actual compensation offered to any candidate may fall outside of the posted range and will be determined based on multiple factors legally permitted in the applicable jurisdiction. These may include, but are not limited to: Geographic location, Education and qualifications, Certifications and licenses, Relevant experience and skills, Seniority and performance, Market and business consideration, Internal pay equity. It is not typical for candidates to be hired at or near the top of the posted compensation range. In addition to base salary, this role may be eligible for additional compensation such as variable incentives, bonuses, or commissions, depending on the position and applicable laws. Capgemini offers a comprehensive, non-negotiable benefits package to all regular, full-time employees. In the U.S. and Canada, available benefits are determined by local policy and eligibility and may include: ul Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade, Company paid holidays, Personal Days, Sick Leave ul Medical, dental, and vision coverage (or provincial healthcare coordination in Canada) ul Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada) ul Life and disability insurance ul Employee assistance programs ul Other benefits as provided by local policy and eligibility

About Capgemini

Capgemini is a global leader in consulting, digital transformation, technology and engineering services. The company is headquartered in Paris, France and operates in over 50 countries. Capgemini provides a range of services including strategy and transformation, application services, technology services, and engineering services. The company serves clients in a variety of industries including automotive, consumer products, financial services, healthcare, and retail.
Learn more about Capgemini
Industry
Founded
1967
NASDAQ

Similar Jobs

More Jobs at Capgemini

  • Capgemini
    Data Engineer
    $86K — $127K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Capgemini
    Lead Salesforce Developer
    $96K — $150K *
    New York, NY 10025 (New York County)
    Enterprise Technology
    In-Person
  • Capgemini
    Senior Data Engineer - Databricks
    $133K — $153K *
    Chicago, IL 60629 (Cook County)
    Information Technology
    In-Person
  • Capgemini
    Powerplatform Developer
    $130K — $140K *
    Chicago, IL 60629 (Cook County)
    Enterprise Technology
    In-Person
  • Capgemini
    Cloud HPC Engineer
    $86K — $117K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar Data Engineer jobs: