Data Engineer 4 (Risk Tech)
In this role at Risk Tech, you will work with our GRC team and partners across the company to build and deploy proprietary solutions for Risk management that are powered by state-of-the-art AI technology. Our products, enhanced with the transformative power of AI, are central to our business and deliver tremendous customer value.
What You’ll Do:
- Collaborate with and across Agile teams to design, develop, test, implement, and support technical solutions in full-stack development tools and technologies
- Influence a team of developers, data analyst and data scientists with deep experience in machine learning, distributed microservices, lakehouse architecture, and full stack systems
- Utilize programming languages such as Python and Spark, along with open-source relational and NoSQL databases, and cloud-based data warehousing platforms including Databricks and Snowflake
- Share your passion for staying on top of trends in data, experimenting with and learning new technologies, participating in internal & external technology communities, and mentoring other members of the data community
- Collaborate with product managers and software engineering to deliver robust cloud-first data solutions that drive powerful experiences to help millions of Americans achieve financial empowerment
- Independently design, build and deliver world-class cloud data solutions and applications with little or no support from supervisors or managers
- Architect and enforce common data engineering design patterns to ensure code quality, maintainability, and reusability across data platforms and pipelines.
- Serve as an ambassador for the data engineering team, communicating technical concepts and data outcomes clearly to both internal and external stakeholders to drive alignment and shared understanding.
- Design and build data pipelines and platforms with a focus on scalability, resilience, and operational efficiency, ensuring robust performance under increasing data volume and business demands.
- Implement data security standards, including encryption at rest/transit and fine-grained access control, ensuring compliance with data privacy regulations
Basic Qualifications:
- Bachelor's Degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 4 years of experience in application development (Internship experience does not apply)
- At least 2 years of experience in distributed data
- At least 2 years of experience with SQL
- At least 2 years of experience with one programming language (Python, Java, or Scala)
- At least 2 years of experience in data pipeline design and development
- At least 1 year of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems
Preferred Qualifications:
- 7+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java
- 4+ years of hands-on experience designing, deploying and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud)
- 4+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks
- 4+ years of experience designing, implementing, and operating real-time or streaming data pipelines
- 2+ years of experience working on data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster)
- 4+ years of experience working with unstructured or semistructured data using NoSQL databases (e.g., Mongo, Cassandra, DynamoDB)
- 4+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift)
- 2+ years of experience working in an Agile development environment
- 2+ years of experience developing user-centric reusable data products
- Experience building and maintaining pipelines for vector embeddings or managing vector search solutions (e.g., Pinecone, Milvus, Qdrant, pgvector, Databricks Vector Search)
- Experience designing data pipelines and context retrieval frameworks to support large language model (LLM) integrations using tools such as LangChain or LlamaIndex
- Experience preparing unstructured text, logs, or document datasets for Generative AI workflows, such text chunking and automated feature extraction
At this time, Capital One will not sponsor a new applicant for employment authorization, or offer any immigration related support for this position (e.g. H1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN, E-3, and O-1, or any other forms of work authorization that require immigration support from an employer).
The minimum and maximum full-time annual salaries for this role are listed below, by location. Please note that this salary information is solely for candidates hired to perform work within one of these locations, and refers to the amount Capital One is willing to pay at the time of this posting. Salaries for part-time roles will be prorated based upon the agreed upon number of hours to be regularly worked.
McLean, VA: $197,300 - $225,100 for Data Engineer 4Richmond, VA: $179,400 - $204,700 for Data Engineer 4
Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate’s offer letter.
This role is also eligible to earn performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI). Incentives could be discretionary or non discretionary depending on the plan.
Capital One offers a comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being. Learn more at the Capital One Careers website. Eligibility varies based on full or part-time status, exempt or non-exempt status, and management level.
This role is expected to accept applications for a minimum of 5 business days.