Data Engineer W/ Databricks + GCP - Austin

Biorce

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience in Data Engineering or related roles
  • Hands-on experience with Databricks Lakehouse platform
  • Knowledge of Google Cloud Platform (GCP) data ecosystem
  • Strong proficiency in SQL and Python
  • Experience designing scalable batch and streaming data pipelines
  • familiarity with data modeling and Lakehouse optimization
  • Bachelor's or Master's degree in Computer Science, Engineering, or related field

Responsibilities

  • Design and maintain scalable ETL/ELT pipelines on Databricks
  • Model data in Delta Lake medallion architecture and build transformations in dbt
  • Build data ingestion workflows from various sources into the Lakehouse
  • Collaborate with data scientists for model training and feature generation
  • Ensure data quality and integrity through testing and monitoring
  • Develop and optimize SQL, Python, and PySpark transformations
  • Manage data storage and lifecycle strategies for efficiency

Benefits

  • Dynamic international team environment
  • Collaboration-focused workplace culture
  • Comprehensive health coverage for physical and mental well-being
  • Flexible hybrid work model
  • Company events for team bonding and recognition
  • Provision of MacBook for enhanced productivity
  • Pet-friendly office environment
Full Job Description
About the role

Following our successful expansion into the U.S. and continued growth across Europe, we are seeking a Data Engineer to help drive our AI and data engineering from our Austin TX (USA) office. Reporting directly to our Data / AI leadership, this person will play a critical role in driving the development of scalable, reliable, and efficient data pipelines across our cloud data platform.

This is an exciting opportunity to build and optimize the data backbone of Biorce's next-generation platform on a Databricks Lakehouse running on Google Cloud (GCP), combining Databricks-native tooling such as Delta Lake, Unity Catalog, Spark/PySpark, and Databricks Workflows with GCP services like BigQuery, Cloud Storage, and Pub/Sub, and dbt as the transformation layer, all in a high-impact, fast-iterating environment.

Who We're Looking For

We are looking for a skilled Data Engineer to join our growing AI and data team. Someone who can work closely with data scientists, AI engineers, and DevOps to design and operationalize robust data flows that fuel advanced analytics, machine learning, and regulatory-grade insights. This person should be able to shape the evolution of Biorce's data and AI architecture, with the Databricks Lakehouse at its core while ensuring scalable, reliable, compliant, and cost-efficient data operations.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines on Databricks (Spark/PySpark, Databricks Workflows, Delta Live Tables / declarative pipelines, Auto Loader), integrating with GCP services such as Dataflow, Pub/Sub, and BigQuery where appropriate.
  • Model data in a Delta Lake medallion architecture (bronze  silver  gold) and build transformations in dbt on Databricks SQL.
  • Build and orchestrate complex data ingestion workflows from diverse clinical, research, and third-party sources into the Lakehouse.
  • Collaborate with data scientists to enable seamless model training, feature generation, and inference data flows (including integration with MLflow).
  • Ensure data quality, integrity, and lineage across all systems through rigorous validation, testing, and monitoring, leveraging Unity Catalog for governance and lineage.
  • Develop and optimize SQL, Python, and PySpark transformations to ensure high performance and maintainability.
  • Manage data storage, partitioning, clustering, and lifecycle strategies (Delta optimization, liquid clustering) for efficiency and cost control.
  • Ensure compliance with SOC 2, ISO 27001, HIPAA, GDPR, and clinical data governance standards in all data operations, using Unity Catalog for access control and auditability.
  • Continuously improve internal frameworks for ingestion, metadata management, and data documentation.
  • Contribute to cross-functional discussions to shape the evolution of Biorce's Databricks-based data and AI architecture.


Must-Haves

  • 3+ years of professional experience in Data Engineering or related roles.

Hands-on experience with the Databricks Lakehouse platform: Delta Lake, Spark/PySpark, Databricks Workflows, and Databricks SQL (Unity Catalog a strong plus).

  • Working knowledge of the GCP data ecosystem: BigQuery, Cloud Storage, Pub/Sub, and related services (Dataflow, Composer, Cloud Functions).

Strong proficiency in SQL and Python for data transformation and automation.

  • Experience designing batch and streaming data pipelines with scalable and fault-tolerant architectures.
  • Familiarity with data modeling, schema design, and Lakehouse/warehouse optimization (medallion architecture, Delta tables).
  • Understanding of API-based ingestion, data normalization, and pipeline monitoring.
  • Exposure to version-controlled, modular pipeline development, such as Terraform, Databricks Asset Bundles, or GitOps.
  • Experience working collaboratively with data scientists and MLOps teams.
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field.


Nice-to-Haves

  • Production experience with dbt (dbt-Databricks adapter) for complex transformations.
  • Familiarity with Delta Live Tables, Auto Loader, Photon, or Databricks performance tuning.
  • Experience with clinical, biomedical, or healthcare datasets.
  • Familiarity with MLflow, Mosaic AI, Vertex AI, or ML metadata tracking.
  • Understanding of data governance and cataloging (Unity Catalog, Data Catalog, Looker, or similar).
  • Knowledge of Apache Beam or advanced Spark internals.
  • Exposure to infrastructure-as-code (Terraform) and containerized workflows (Kubernetes or Docker).
  • Experience implementing data validation frameworks such as Great Expectations, dbt tests, or TFX Data Validation.
  • Strong focus on reliability, observability, and continuous improvement of data systems.


Why Join Us?

  • A dynamic work environment with an international team, where collaboration and diversity thrive.
  • Work alongside top talent, united by a shared purpose and committed to making a real impact.
  • Comprehensive private health coverage to ensure your physical and mental well-being.
  • Hybrid work model offering flexibility to balance your professional and personal life.
  • Company events to celebrate achievements and enjoy time together.
  • Get equipped with a MacBook to enhance your productivity and work experience.
  • Our office is pet-friendly! You'll likely be greeted by a few wagging tails upon arrival.


Similar Jobs

More Jobs at Biorce

More Information Technology Jobs

Find similar Data Engineer W/ Databricks + GCP - Austin jobs: