AI/ML Data Engineer to design and govern the data foundation behind secure AI-enabled applications for the U.S. Government. You will own AI-ready datasets, from ingestion and quality controls through privacy protection, lineage, and retrieval-quality measurement, so that every AI output can be traced back to trusted source data.
Key Responsibilities:- Design and maintain data ingestion, transformation, and processing pipelines (ETL/ELT) for AI training, evaluation, retrieval, and operations, including support for data migration and cleansing.
- Curate, validate, and version datasets, and maintain dataset inventories, metadata, lineage, provenance, and ingestion logs.
- Implement automated data-quality checks for duplication, schema changes, completeness, and freshness, and maintain dataset quality scorecards and drift reports.
- Design data models, vector stores, and embedding schemas for Retrieval-Augmented Generation (RAG) knowledge bases, and re-index content when sources change.
- Measure retrieval and model quality against established baselines using metrics such as precision/recall, MRR, NDCG, and context relevance.
- Prepare data-related deliverables, including AI model cards, ML and AI pipeline documentation, RAG/AI Pipeline Evaluation Reports, data dictionaries and embedding schema documentation, and responses to Government data calls.
- Build secure structured-data access for AI applications (e.g., natural-language-to-SQL with query validation and role-based authorization), and support dashboards and operational analytics.
- Monitor data and ML pipelines, troubleshoot failures, and support root-cause analysis, while keeping all data in FedRAMP-authorized cloud regions with FIPS-validated encryption.
- Support Responsible AI practices by preparing representative evaluation datasets, testing AI outputs for bias, accuracy, and hallucination, and documenting results to meet federal AI governance requirements.
- Secure the AI data path, from source datasets and embeddings to prompts and logs, and support AI risk testing such as data poisoning.
Requirements
- Bachelor's degree in Computer Science, Engineering, Mathematics, Information Systems, or a related field and 5+ years of relevant experience. Equivalent experience may substitute for the degree.
- 5+ years of hands-on data engineering or database development experience, including data modeling, SQL, and ETL/ELT pipeline development.
- Strong proficiency in Python and experience with data processing frameworks (e.g., pandas, Spark) and workflow orchestration tools (e.g., Airflow).
- 1+ year of experience preparing data for Generative AI or machine learning, such as embeddings, vector databases, RAG knowledge bases, or training and evaluation datasets.
- Experience implementing data quality, lineage, metadata management, and data governance controls.
- Experience protecting sensitive data (PII), including masking, minimization, and access controls.
- Experience with a major cloud data platform (AWS, Azure, or Google Cloud), Git, and CI/CD tools.
- Must be a U.S. citizen or lawful permanent resident (green card holder).
- Must reside in the Washington, DC metropolitan area and be able to work on-site at Government offices.
- Must be able to obtain and maintain a Public Trust background investigation.
Preferred Qualifications - Master's degree in a related field and 7+ years of data engineering experience, including support of federal agency programs.
- Experience evaluating retrieval quality and building RAG pipelines with vector stores (e.g., pgvector, OpenSearch).
- Experience with MLOps tooling, including model and dataset versioning, feature stores, or model registries.
- Generative AI, LLM security, or data certification (e.g., Databricks Generative AI Engineer, Snowflake SnowPro, AWS, Microsoft Azure, or Google Cloud).
- An active Public Trust or prior federal background investigation.
BENEFITS AND PERKS- Comprehensive Health Benefits (Medical, Dental and Vision)
- Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
- Retirement Plan with 4% match and discretionary match at year end
- Paid Time Off (PTO): 15 days of PTO accrued per year; 7 holidays+ 3 Floating holidays; 2 Innovation days (paid training days)
- Short Term and Long-Term Disability
- Paid Parental Leave
- Paid Jury Duty leave
- Life and AD&D Insurance
- Critical Illness Insurance
- Training and Development
- Wellness Incentives & Discount programs
- Employee Referral Program
- Annual Charity Donation Match
- Awards and Recognition