Data Engineer Lead

Capital Group

• $213K — $342K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in data or software engineering, with a focus on leading complex data platforms.
  • Proficiency in Python and SQL, with strong software design understanding and distributed data processing expertise.
  • Hands-on experience with Databricks on AWS, including key tools and understanding security implications.
  • Experience with orchestrating production pipelines using Apache Airflow and building dependable transformations with dbt.
  • Strong background in data modeling and governance, specifically with dimensional models and data lineage.
  • Implemented data quality and CI/CD processes for data platforms using tools like Deequ and Terraform.
  • Familiarity with AI tools for preparing governed data and evaluating AI-generated outputs.

Responsibilities

  • Set the data engineering strategy and roadmap, influencing practices at Capital Group.
  • Design and build data pipelines using cutting-edge tools, promoting reusable solutions.
  • Analyze new datasets and incorporate them into the data platform.
  • Lead cross-team data initiatives, managing full project lifecycles and balancing risks.
  • Drive AI-first engineering practices by defining requirements and directing AI agents effectively.
  • Develop reusable workflows and skills for data engineering processes.
  • Ensure data reliability and quality through comprehensive testing strategies and automated controls.
  • Collaborate with investment teams to demonstrate the impact of data and AI on portfolio construction.

Benefits

  • Individual annual performance bonus opportunities.
  • 15% contribution to retirement plan from Capital on eligible earnings.
  • Engagement with a talent community for other job opportunities.
  • Comprehensive benefits package including health, wellness, and community initiatives.
Full Job Description
"I can succeed as a Data Engineer Lead at Capital Group."

As a Data Engineer Lead in Capital Solutions Group Technology (CSGT), you will set the technical direction for the data platform and data products that power portfolio construction, investment research, and monitoring capabilities for the Capital Solutions Group (CSG), the Capital Group investment unit that manages our multi-asset fund of funds. Our investment professionals rely on this data to construct and monitor portfolios, evaluate underlying funds, and make investment decisions. This is a senior, hands-on individual contributor role. You will write code, own architecture and delivery, and partner closely with investment group, product management, and technology leaders. You will be a leader in our AI-first engineering practices, where engineers direct AI agents through clear specifications, context, and verification. You will build the trusted data foundation that AI applications need to answer complex investment questions. You have an agile mindset and will lead the design, implementation, and delivery of large-scale, critical, and complex data architecture, storage, and pipelines. You are passionate about our mission, and committed to driving superior long-term investment results through the application of modern engineering and data management methods.

You will:
  • Set the data engineering strategy and roadmap for CSGT, including the Lakehouse architecture on Databricks and AWS. Make and explain decisions on scalability, security, reliability, and cost, and influence standards and practices across adjacent teams and the wider Capital Group technology organization, contributing to centers of excellence and firm-wide engineering practices.
  • Design and build ingestion, transformation, and serving pipelines with Databricks, PySpark, Delta Lake, dbt, and Airflow, and create reusable patterns and frameworks so the team delivers consistent, maintainable data products.
  • Analyze new structured and unstructured datasets at the business-capability level and fit them into the platform's data domains and subject areas.
  • Own delivery of complex, cross-team data initiatives from requirements through production support, including estimates, work breakdown, sequencing, dependencies, and cost. Surface risks early and balance immediate business needs with a durable platform.
  • Lead the continued evolution of AI-first engineering: turn business outcomes into well-defined specifications, engineer the business, architectural, and repository context agents need, and direct agents to plan, build, test, and document changes in small, reviewable increments.
  • Build reusable agent workflows, skills, and tool integrations for data engineering work such as profiling, source-to-target mapping, pipeline and test generation, schema-change analysis, and incident investigation. Define which actions agents can take on their own, which require human approval, and how agent activity is reviewed and traced.
  • Make data understandable and reliable for AI through curated Unity Catalog metadata, lineage, business definitions, semantic models, and access controls, and connect it to Databricks Genie and other AI applications used by investment professionals. Define evaluation datasets and acceptance criteria for agents and AI-generated SQL and code.
  • Define the testing strategy across all layers of the platform, including performance, stability, and availability, and review and approve quality metrics before release. Build data quality, reconciliation, freshness, observability, and recovery controls into automated testing and CI/CD, with security and policy checks embedded from the start.
  • Partner with investment professionals and product managers to develop a shared product vision and ownership of business outcomes, demonstrating how data and AI can scale research and portfolio construction.
  • Raise the engineering bar: lead design and code reviews, direct the day-to-day work of engineers on your initiatives, and guide them through the most complex data and performance issues. Teach engineers to inspect and challenge AI-generated work, share reusable patterns and context through internal and external forums, and help managers identify strengths and development needs.


"I am the person Capital Group is looking for."

Required qualifications:
  • You have 10+ years of experience in data or software engineering, including technical leadership of complex production data platforms delivered across multiple teams.
  • You have strong hands-on Python and SQL skills, sound software design judgment, and a deep understanding of distributed data processing, query performance, and automated testing.
  • You have production experience with Databricks on AWS, including PySpark, Delta Lake, Unity Catalog, Databricks Jobs, Databricks SQL, and Databricks Asset Bundles, and you understand the security, access, and cost implications of your designs.
  • You have orchestrated production pipelines with Apache Airflow (including Astronomer) and built tested transformations with dbt, with reliable retries, backfills, and dependency management.
  • You have strong data modeling and governance experience, including dimensional and time-series models, slowly changing dimensions, bi-temporal history, data contracts, lineage, and semantic metadata.
  • You have implemented data quality, observability, and CI/CD for data platforms using tools such as Deequ, dbt tests, Lakehouse Monitoring, Datadog, Terraform, and Harness.
  • You have experience preparing governed data for AI through natural-language-to-SQL tools such as Databricks Genie, semantic metadata, or other governed data-access patterns.
  • You use AI coding agents well beyond code completion: writing specifications, supplying context, running tests, and reviewing generated changes through source control.
  • You evaluate AI-generated output with representative test cases, regression tests, execution traces, and human review, and you can distinguish a plausible answer from a verified one.
  • You understand prompt injection, sensitive data handling, and least-privilege access, and can design approval boundaries and audit trails for agents working against enterprise systems.
  • You lead architecture discussions, influence without formal authority, develop other engineers, and explain technical choices and trade-offs clearly to investment professionals and technology leaders.
  • You are an agent of change with a sense of urgency: you question how work gets done and remove or automate what does not add value, while respecting what came before.
  • You have a bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.


Preferred qualifications:
  • You have experience with investment management data such as portfolios, positions, returns, exposures, benchmarks, and attribution, or with multi-asset portfolio construction.
  • You have experience building LLM applications or agent workflows that call tools and APIs, including context management, retrieval, state, and error handling.
  • You have familiarity with Model Context Protocol (MCP).
  • You have experience with PostgreSQL, SQL Server, or Lakebase, or with modernizing legacy data platforms onto a Lakehouse.


"I can apply in less than 4 minutes."

You've reviewed this job posting and you're ready to start the candidate journey with us. Apply now to move to the next step in our recruiting process. If this role isn't what you're looking for, check out our other opportunities and join our talent community.

Southern California Base Salary Range: $201,683-$322,693

New York Base Salary Range: $213,795-$342,072

In addition to a highly competitive base salary, per plan guidelines, restrictions and vesting requirements, you also will be eligible for an individual annual performance bonus, plus Capital's annual profitability bonus plus a retirement plan where Capital contributes 15% of your eligible earnings.

You can learn more about our compensation and benefits here.

Similar Jobs

More Jobs at Capital Group

More Information Technology Jobs

Find similar Data Engineer Lead jobs: