Description
Five Rivers Analytics is seeking a versatile and driven Data Engineer with strong hands-on expertise in Databricks. This is a remote role based in South Carolina and the qualified candidate will be responsible for building, optimizing, and maintaining robust data pipelines, with a particular focus on reliably capturing source changes and preserving historical truth. The ideal candidate understands how to structure scalable data Lakehouse's using Delta Lake, write highly optimized SQL and Python, and implement Change Data Capture (CDC) and Slowly Changing Dimension (SCD) architectures. You will deliver secure, auditable, and high-performance data products to empower our downstream data consumers. This is in support of the Logistics Analytics Data Services (LADS) program for the Marine Corps.
To join our team of outstanding professionals, apply today!
Responsibilities
Data Engineering, CDC & Pipeline Development
- Pipeline Orchestration: Design, build, and maintain scalable batch and streaming ETL/ELT pipelines using Python, SQL, and built-in Databricks Workflows.
- Change Data Capture (CDC): Establish efficient pipelines to ingest incremental source changes (inserts, updates, and deletes) into the lakehouse using native Databricks capabilities.
- Lakehouse Architecture: Implement and manage data layouts using the Medallion Architecture (Bronze, Silver, Gold layers) to clean, enrich, and optimize structured and semi-structured datasets.
- Data Governance & Security: Administer schemas, catalogs, and secure access permissions using Unity Catalog to ensure strict data governance.
- Performance Tuning: Optimize SQL queries, caching, indexing, and cluster configurations to reduce execution times and platform compute costs.
Historical Dimensional Modeling & Cross-Functional Collaboration
- SCD & History Implementation: Design and implement robust Slowly Changing Dimension (SCD) patterns and historical truth tables in Delta Lake.
- Requirements & Collaboration: Work closely with dashboard developers, data analysts, and data scientists to deeply understand and capture their data requirements.
- Source System Mastery: Proactively investigate and learn about various source systems to understand how they work and how data is generated and extracted.
- Data Modeling: Create clean, optimized, and reusable semantic data models in Databricks SQL that serve as the "single source of truth" for the business.
Data Quality & Operations
- CI/CD & Git Integration: Implement software development best practices using Git integration in Databricks Repos for version control, code review, and automated deployments.
- Data Quality Monitoring: Implement automated data validation, logging, and alerting systems using native Databricks data quality tools to ensure high data integrity.
Qualifications
- Current and active Secret security clearance
- Bachelor’s degree in CS, IT, Engineering, Math, or related field, or equivalent experience.
- 8+ years' experience (Bachelor’s) or 6+ years (Master’s).
- Strong understanding of relational database principles, indexing, transaction management, and query optimization.
- Hands-on experience with SQL.
- Knowledge of database security controls and encryption.
- Experience in Agile/SAFe environments.
Preferred Qualification:
- Experience modeling data for use in Foundry (Ontologies).
Benefits InformationRegular - The company offers a comprehensive benefits program, including medical, dental, vision, life insurance, 401(k) and a range of other voluntary benefits. Paid Time Off (PTO) is offered to regular full-time and part-time employees.
Pay Range$140,000 - $160,000
Job ID2026-24960
Work TypeRemote