Senior Data Engineer, Bioinformatics, Cheminformatics, Materials

Lila Sciences

$144K — $240K *
Pharmaceuticals & Biotech
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2-6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science.
  • Strong Python skills, including production-quality code.
  • Proficient in SQL, especially with Postgres.
  • Experience with ETL pipelines, data modeling, and reusable transformations.
  • Solid understanding of data science principles and tools like pandas and NumPy.
  • Ability to clean and validate noisy scientific measurements.
  • Experience with workflow orchestration tools like Flyte or Airflow.

Responsibilities

  • Design and build ETL pipelines for transforming lab outputs into scientific data.
  • Model and structure heterogeneous data from various scientific instruments.
  • Develop validation checks and data quality workflows.
  • Create reusable data analysis functions for research workflows.
  • Enhance automation and observability in data flows from instruments to results.
  • Establish trustworthy canonical datasets for scientists and researchers.
  • Utilize AI coding tools to streamline development processes.

Benefits

  • Comprehensive medical, dental, and vision coverage.
  • Employer-paid life and disability insurance.
  • Flexible time off with generous holidays.
  • Paid parental leave and educational assistance programs.
  • Commuter benefits including bike share memberships.
  • Company-subsidized lunch program.
Full Job Description
Your Impact at LILA

As a Data Engineer, you'll build ETL pipelines and data models for Lila's scientific data platform, working at the intersection of data engineering, computational biology, chemistry, and materials science.

You'll partner with AI researchers and experimentalists to turn raw lab instrument outputs into validated, analysis-ready datasets. The core challenge is data modeling: transforming messy, per-instrument measurements into clean, well-typed data that is efficient to query, reliable to use, and ready for downstream analysis.

You'll also build domain-specific analysis functions and reusable data pipelines that help scientists and AI researchers move faster without re-deriving bespoke solutions.

What You'll Be Building
• Design pipelines that turn raw lab output into analysis-ready scientific data. • Model heterogeneous data from bio, chemistry, and materials instruments. • Build validation checks, schema-evolution gates, and data quality workflows. • Develop reusable analysis functions for scientific and AI research workflows. • Improve automation and observability across instrument-to-result data flows. • Build canonical datasets that scientists and AI researchers can trust. • Use AI coding tools to accelerate pipeline development and team velocity.

What You'll Need to Succeed
• 2-6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. •
Strong Python skills, including typed, tested, production-quality code.
• Strong SQL skills, especially with Postgres or similar relational databases.
• Experience building ETL pipelines, data models, and reusable data transformations.
• Data science foundation, including statistics and pandas, NumPy, or similar tools.
• Experience translating noisy scientific measurements into accurate, validated datasets.
• Workflow orchestration experience, ideally Flyte, Airflow, Prefect, Dagster, or Nextflow.
• Active use of AI coding tools in day-to-day engineering work.

Bonus Points For
• Experience with columnar or lakehouse stacks such as Parquet, Iceberg, DuckDB, Polars, or Ibis.
• Familiarity with event-driven pipelines such as NATS or Kafka. • Exposure to lab instrument data formats, LIMS, or ELN systems.
• Familiarity with life sciences assays, sequencing, imaging, or flow cytometry.
• Familiarity with materials or chemistry methods such as XRD, XRF, SEM, TGA, or DSC.
• Experience with curve fitting, peak detection, or unit and dimensional analysis.

Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range

$144,000-$240,000 USD

Similar Jobs

More Jobs at Lila Sciences

More Pharmaceuticals & Biotech Jobs

Find similar Senior Data Engineer, Bioinformatics, Cheminformatics, Materials jobs: