About the roleFollowing our successful expansion into the U.S. and continued growth across Europe, we are seeking a Data Engineer to help drive our AI and data engineering from our Austin TX (USA) office. Reporting directly to our Data / AI leadership, this person will play a critical role in driving the development of scalable, reliable, and efficient data pipelines across our cloud data platform.
This is an exciting opportunity to build and optimize the data backbone of Biorce's next-generation platform on a Databricks Lakehouse running on Google Cloud (GCP), combining Databricks-native tooling such as Delta Lake, Unity Catalog, Spark/PySpark, and Databricks Workflows with GCP services like BigQuery, Cloud Storage, and Pub/Sub, and dbt as the transformation layer, all in a high-impact, fast-iterating environment.
Who We're Looking ForWe are looking for a skilled Data Engineer to join our growing AI and data team. Someone who can work closely with data scientists, AI engineers, and DevOps to design and operationalize robust data flows that fuel advanced analytics, machine learning, and regulatory-grade insights. This person should be able to shape the evolution of Biorce's data and AI architecture, with the Databricks Lakehouse at its core while ensuring scalable, reliable, compliant, and cost-efficient data operations.
Key Responsibilities- Design, develop, and maintain scalable ETL/ELT pipelines on Databricks (Spark/PySpark, Databricks Workflows, Delta Live Tables / declarative pipelines, Auto Loader), integrating with GCP services such as Dataflow, Pub/Sub, and BigQuery where appropriate.
- Model data in a Delta Lake medallion architecture (bronze silver gold) and build transformations in dbt on Databricks SQL.
- Build and orchestrate complex data ingestion workflows from diverse clinical, research, and third-party sources into the Lakehouse.
- Collaborate with data scientists to enable seamless model training, feature generation, and inference data flows (including integration with MLflow).
- Ensure data quality, integrity, and lineage across all systems through rigorous validation, testing, and monitoring, leveraging Unity Catalog for governance and lineage.
- Develop and optimize SQL, Python, and PySpark transformations to ensure high performance and maintainability.
- Manage data storage, partitioning, clustering, and lifecycle strategies (Delta optimization, liquid clustering) for efficiency and cost control.
- Ensure compliance with SOC 2, ISO 27001, HIPAA, GDPR, and clinical data governance standards in all data operations, using Unity Catalog for access control and auditability.
- Continuously improve internal frameworks for ingestion, metadata management, and data documentation.
- Contribute to cross-functional discussions to shape the evolution of Biorce's Databricks-based data and AI architecture.
Must-Haves- 3+ years of professional experience in Data Engineering or related roles.
Hands-on experience with the Databricks Lakehouse platform: Delta Lake, Spark/PySpark, Databricks Workflows, and Databricks SQL (Unity Catalog a strong plus).
- Working knowledge of the GCP data ecosystem: BigQuery, Cloud Storage, Pub/Sub, and related services (Dataflow, Composer, Cloud Functions).
Strong proficiency in SQL and Python for data transformation and automation.
- Experience designing batch and streaming data pipelines with scalable and fault-tolerant architectures.
- Familiarity with data modeling, schema design, and Lakehouse/warehouse optimization (medallion architecture, Delta tables).
- Understanding of API-based ingestion, data normalization, and pipeline monitoring.
- Exposure to version-controlled, modular pipeline development, such as Terraform, Databricks Asset Bundles, or GitOps.
- Experience working collaboratively with data scientists and MLOps teams.
- Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field.
Nice-to-Haves- Production experience with dbt (dbt-Databricks adapter) for complex transformations.
- Familiarity with Delta Live Tables, Auto Loader, Photon, or Databricks performance tuning.
- Experience with clinical, biomedical, or healthcare datasets.
- Familiarity with MLflow, Mosaic AI, Vertex AI, or ML metadata tracking.
- Understanding of data governance and cataloging (Unity Catalog, Data Catalog, Looker, or similar).
- Knowledge of Apache Beam or advanced Spark internals.
- Exposure to infrastructure-as-code (Terraform) and containerized workflows (Kubernetes or Docker).
- Experience implementing data validation frameworks such as Great Expectations, dbt tests, or TFX Data Validation.
- Strong focus on reliability, observability, and continuous improvement of data systems.
Why Join Us?- A dynamic work environment with an international team, where collaboration and diversity thrive.
- Work alongside top talent, united by a shared purpose and committed to making a real impact.
- Comprehensive private health coverage to ensure your physical and mental well-being.
- Hybrid work model offering flexibility to balance your professional and personal life.
- Company events to celebrate achievements and enjoy time together.
- Get equipped with a MacBook to enhance your productivity and work experience.
- Our office is pet-friendly! You'll likely be greeted by a few wagging tails upon arrival.