Overview
Position contingent on contract award - details below are subject to change based on final award.
You will architect and operate enterprise‑scale data pipelines and platforms supporting mission‑critical public health analytics. The role centers on Databricks (Delta Lake, Unity Catalog, job orchestration), Palantir Foundry (transforms, datasets, pipelines), and distributed processing with PySpark and SQL. You will own end‑to‑end pipelines—from schema design and medallion architecture through production operations—while optimizing transformations, partitioning, and query performance. You’ll modernize multilingual analytical codebases (Python, SQL), build shared libraries/templates, and guide platform onboarding/migrations (including Foundry) to accelerate adoption. Work spans cloud‑native orchestration (Azure Data Factory/Functions/Storage or AWS/GCP equivalents), CI/CD for data services, observability and data quality, and data security & governance. Strong stakeholder communication/presentation is essential to translate complex engineering choices into clear, outcome‑oriented recommendations for program leadership and CDC partners.
Responsibilities
- Design, build, and own scalable data pipelines and platform components (Databricks, Palantir Foundry).
- Engineer robust schemas, medallion patterns, partitioning, and performance tuning (PySpark/SQL).
- Implement cloud‑native orchestration and CI/CD for data services; automate testing/validation.
- Lead onboarding/migration to data platforms; develop reusable libraries and templates.
- Champion data quality, observability, and governance; communicate tradeoffs and decisions to stakeholders.
- Other duties as assigned
Qualifications
- 5+ years of experience with PySpark, SQL, distributed data processing and enterprise-scale data engineering
- 5+ years of experience designing, developing, owning, and maintaining end-to-end production data pipelines
- 5+ years of experience with data architecture, including schema design, Medallion architecture, partitioning strategies, and performance optimization.
- 5+ years of experience collaborating with business stakeholders, gathering requirements, and delivering technical presentations to technical and non-technical audiences
- 2+ years of hands-on experience with Databricks, including Delta Lake, Unity Catalog, Job orchestration
- 2+ years of experience with Microsoft Azure, including Azure Data Factory, Azure Functions, Azure Storage, or equivalent cloud platforms such as AWS or Google Cloud Platform (GCP)
- Experience with Palantir Foundry, including pipeline development, transformation logic, and data modeling
- Experience supervising others and leading deliverables in cross-functional environments
- Ability to obtain and maintain a Public Trust or Suitability/Fitness determination based on client requirements
- Bachelor's degree
- Successfully pass background and drug screening.
Final salary determination based on skill-set, qualifications, and approved funding.
Many of our jobs come with great benefits – Some offerings are dependent upon the role, work schedule, or location, and may include the following:
Paid Time Off
PTO / Vacation – 5.67 hours accrued per pay period / 136 hours accrued annually
Paid Holidays - 11
California residents receive an additional 24 hours of sick leave a year
Health & Wellness
Medical
Dental
Vision
Prescription
Employee Assistance Program
Short- & Long-Term Disability
Life and AD&D Insurance
Spending Account
Flexible Spending Account
Health Savings Account
Health Reimbursement Account
Dependent Care Spending Account
Commuter Benefits
Retirement
401k / 401a
Voluntary Benefits
Hospital Indemnity
Critical Illness
Accident Insurance
Pet Insurance
Legal Insurance
ID Theft Protection
Teleworking Permitted?Yes
Teleworking Details100% Remote - US
Estimated Salary/WageUSD $110,000.00/Yr. Up to USD $120,000.00/Yr.