Data Engineer 4 (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI)
Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive, and iterative delivery environment?
In this role within the Marketing and Messaging engineering group, you will be at the forefront of driving a major technical transformation. Your team builds scalable platforms that deliver hyper-personalized, omnichannel messages and experiences across owned and paid Adtech channels to help millions of Americans achieve financial empowerment.
What You’ll Do- Architect & Build: Independently design, develop, test, and deploy resilient, scalable, cloud-first data solutions, lakehouse architectures, and real-time streaming data pipelines.
- Technical Leadership & Mentorship: Lead and influence cross-functional Agile teams (developers, data analysts, and data scientists). Enforce common data engineering design patterns to ensure code quality, maintainability, and operational efficiency.
- Modern Tech Stack: Leverage programming languages like Python, SQL, Scala, Spark, along with NoSQL databases, open-source RDBMS, and cloud data platforms such as Databricks and Snowflake.
- Governance & Security: Implement data security standards (encryption at rest/in transit, fine-grained access control) and uphold strict standards for data quality, governance, and permissible use.
- Cross-Functional Collaboration: Partner with digital product managers and software engineers to deliver robust solutions that meet business demands and enhance customer experiences.
- Ambassador & Innovator: Communicate technical concepts and data outcomes clearly to both technical and non-technical stakeholders. Stay on top of tech trends, participate in engineering communities, and leverage interactive AI tooling (e.g., Claude Code, GitHub Copilot) to accelerate delivery.
- Code Quality & Reliability: Perform unit tests, conduct peer code reviews, and optimize systems for performance, observability, and downstream consumers.
Basic Qualifications:
- Bachelor's Degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 4 years of experience in application development (Internship experience does not apply)
- At least 2 years of experience in distributed data
- At least 2 years of experience with SQL
- At least 2 years of experience with one programming language (Python, Java, or Scala)
- At least 2 years of experience in data pipeline design and development
- At least 1 year of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems
Preferred Qualifications
- Master’s Degree in Computer Science, Data Engineering, Information Systems, or a related technical field.
- 7+ years of experience in application development using languages such as Python, SQL, Spark, Scala, or Java.
- 4+ years of hands-on experience designing, deploying, and operating data workloads in a public cloud environment (AWS, Azure, or GCP).
- 4+ years of experience building distributed data/computing workloads using tools like EMR, Spark, AWS Glue, Databricks, MapReduce, Hadoop, or Hive.
- 4+ years of experience designing, implementing, and operating real-time and streaming data applications (e.g., Kafka).
- 4+ years of experience designing and managing data warehousing platforms (Snowflake or Redshift) and data modeling for data warehousing.
- 4+ years of experience with NoSQL implementations and semi-structured data (MongoDB, Cassandra, or DynamoDB).
- 4+ years of experience with UNIX/Linux environments, basic commands, and shell scripting.
- 2+ years of experience with data orchestration (e.g., Airflow, Dagster) or data observability tools (e.g., Monte Carlo, Splunk).
- 2+ years of experience developing user-centric reusable data products.
- 2+ years of experience in Agile engineering practices and delivery environments.
- Hands-on experience leveraging interactive AI tooling (e.g., Claude Code, GitHub Copilot) to accelerate software delivery.
At this time, Capital One will not sponsor a new applicant for employment authorization, or offer any immigration related support for this position (e.g. H1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN, E-3, and O-1, or any other forms of work authorization that require immigration support from an employer).
The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked.
McLean, VA: $197,300 - $225,100 for Data Engineer 4New York, NY: $215,200 - $245,600 for Data Engineer 4
Richmond, VA: $179,400 - $204,700 for Data Engineer 4
Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate’s offer letter.
This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan.
Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website. Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level.
This role is expected to accept applications for a minimum of 5 business days.