Type of Requisition:Regular
Clearance Level Must Currently Possess:None
Clearance Level Must Be Able to Obtain:None
Public Trust/Other Required:None
Job Family:Data Science and Data Engineering
Job Qualifications:Skills:Cloud Technology, Databricks Lakeflow, Databricks Platform, Databricks Unity Catalog, Data Lake
Certifications:None
Experience:7 + years of related experience
US Citizenship Required:No
Job Description:GDIT is seeking a Databricks Engineer to help build and operate modern data solutions supporting the NIA Data Enclave. In this role, you will design, develop, and optimize scalable data pipelines that enable researchers and analysts to securely work with sensitive federal and non-federal datasets.
You will work at the intersection of data engineering, cloud technology, and mission-focused analytics, partnering with data scientists, researchers, software engineers, architects, and platform teams to turn complex data into reliable, governed, and accessible analytical resources.
This is an opportunity to apply your Databricks and Spark expertise to a mission where data quality, security, scalability, and reliability matter.
How You'll Make an ImpactAs a Databricks Engineer, you will:
- Design, build, and optimize scalable ETL/ELT pipelines using Databricks, PySpark, SQL, and Delta Lake.
- Develop and maintain Databricks notebooks, jobs, and workflows that support high-volume analytical workloads.
- Build reliable data ingestion, transformation, validation, and integration processes.
- Help migrate and modernize existing data workloads for improved scalability, performance, and maintainability.
- Optimize Spark workloads through effective partitioning, caching, joins, file management, and other performance-tuning techniques.
- Implement automated testing, data quality checks, monitoring, logging, and operational processes.
- Support CI/CD and infrastructure automation for Databricks workloads using Git and tools such as Azure DevOps or GitHub Actions.
- Configure and optimize Databricks compute, clusters, runtimes, and job execution.
- Work within secure, role-based cloud environments and help implement appropriate data governance and access controls.
- Troubleshoot production issues, perform root-cause analysis, and continuously improve the reliability of data services.
- Collaborate with data scientists, researchers, analysts, and other engineers to deliver high-quality analytical datasets.
What You'll Need to Succeed- Bachelor's degree in computer science, software engineering, data engineering, or a related technical field.
- 5+ years of data engineering experience, including significant hands-on experience with Databricks.
- Strong experience with PySpark, SQL, Apache Spark, Delta Lake, and Databricks.
- Experience developing production-grade data pipelines and workflows.
- Experience working with cloud-based data platforms and storage such as AWS S3, Azure Data Lake Storage, or Google Cloud Storage.
- Experience with Git and CI/CD practices for deploying and managing data engineering workloads.
- Understanding of distributed data processing, data modeling, data quality, and pipeline performance optimization.
- Experience troubleshooting and supporting production data workloads.
- Understanding of cloud security concepts such as role-based access control, identity management, least-privilege access, and data protection.
- Strong communication skills and the ability to collaborate effectively with technical and mission-focused stakeholders.
Preferred Qualifications- Experience with Databricks Unity Catalog and enterprise data governance.
- Experience working in FISMA Moderate/High or other regulated environments.
- Experience with AWS, Azure, and/or GCP in a multi-cloud environment.
- Familiarity with CMS, federal, healthcare, biomedical, or other sensitive datasets.
- Experience with secure data enclaves, restricted-access environments, or federated data platforms.
- Experience with Terraform or other infrastructure-as-code technologies.
- Experience with Databricks Lakeflow Declarative Pipelines / Delta Live Tables.
- Experience implementing data lineage, metadata management, monitoring, and audit-ready logging.
- Familiarity with Databricks APIs, SDKs, or automation frameworks.
The likely salary range for this position is $140,250 - $189,750. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.
Scheduled Weekly Hours:
40
Travel Required:
10-25%
Telecommuting Options:
Hybrid
Work Location:
Any Location / Remote
Additional Work Locations:
USA MD Gaithersburg
Total Rewards at GDIT:
Our benefits package for all US-based employees includes a variety of medical plan options, some with Health Savings Accounts, dental plan options, a vision plan, and a 401(k) plan offering the ability to contribute both pre and post-tax dollars up to the IRS annual limits and receive a company match. To encourage work/life balance, GDIT offers employees full flex work weeks where possible and a variety of paid time off plans, including vacation, sick and personal time, holidays, paid parental, military, bereavement and jury duty leave. GDIT typically provides new employees with 15 days of paid leave per calendar year to be used for vacations, personal business, and illness and an additional 10 paid holidays per year. Paid leave and paid holidays are prorated based on the employee's date of hire. The GDIT Paid Family Leave program provides a total of up to 160 hours of paid leave in a rolling 12 month period for eligible employees. To ensure our employees are able to protect their income, other offerings such as short and long-term disability benefits, life, accidental death and dismemberment, personal accident, critical illness and business travel and accident insurance are provided or available. We regularly review our Total Rewards package to ensure our offerings are competitive and reflect what our employees have told us they value most.