Data Engineer

Matterworks, Inc

$110K — $130K *
Pharmaceuticals & Biotech
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years of production experience building data pipelines
  • Proficient in Python and SQL
  • Familiarity with cloud data infrastructure, ideally with Argo Workflows, Metaflow, Glue, Athena, DuckDB, and Terraform
  • Proven ownership of a pipeline or dataset from start to finish
  • Ability to handle messy data formats and resolve discrepancies
  • Experience using AI coding tools with a critical approach to their output
  • Strong written communication skills, especially in technical documentation
  • Interest in science - experience in life sciences, biotechnology, or biochemistry is a plus.

Responsibilities

  • Build and operate end-to-end data pipelines, ensuring they are designed, tested, documented, and running on schedule.
  • Transform raw data into usable datasets with consistent schemas and trustworthy metadata.
  • Design and enhance data quality checks to ensure pipeline integrity before publication.
  • Ingest and manage new datasets from public and partner sources, accommodating messy formats.
  • Collaborate with AI teams to ensure datasets meet scientific standards and usability requirements.

Benefits

  • Flexible hybrid work model allowing remote options
  • Health and dental insurance
  • Vision coverage
  • Long- and short-term disability insurance
  • Life insurance
  • 401k plan with company match
  • Unlimited time off policy
  • Company support for continued education and conference participation.
Full Job Description
Position Overview

Matterworks is seeking a Data Engineer to build and run the pipelines behind our models. Our platform acquires mass spectrometry and molecular data at scale and turns it into the datasets our AI team trains on and our product reasons over. You will own real pieces of that path end to end.

Our data serves two different customers. The AI team needs training corpora that are complete, correctly split, and reproducible. The product and agentic layer needs values a scientist can explain and we can safely show a customer. You will build inside the contracts and checks that keep both honest, and you will help extend them.

This is a hands-on role with a clear growth path. You will start by owning well-scoped pipelines and datasets and grow toward owning larger parts of the platform. You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team.

Key Responsibilities
  • Build and Operate Pipelines: Own well-scoped pipelines end to end: designed, tested, instrumented, documented, and running on a schedule.
  • Labels and Enrichment: Turn raw data into datasets people can actually use, with consistent schemas, trustworthy metadata, and documented definitions.
  • Data Quality Checks: Design, build and extend our quality checks so that each build gets compared against the last one before it publishes, and a failure stops the pipeline instead of shipping.
  • Ingest and Acquisition: Bring new public and partner datasets into the platform: fetching, converting, validating, and reconciling them against what we already hold. Expect messy scientific and vendor formats and file that require continuous improvements to our systems to handle at scale.


About You
  • 2+ years of professional experience building data pipelines in production.
  • Proficient in Python and SQL.
  • Working knowledge of cloud data infrastructure. We run Argo Workflows and Metaflow on EKS, Glue and Athena over Apache Iceberg and Parquet, DuckDB, and Terraform. Depth in any comparable stack transfers fine.
  • Demonstrated experience owning a pipeline or dataset end to end, including the tests, the monitoring, and the failures.
  • Comfort with messy data and messy formats, and the patience to track down why two sources disagree.
  • Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness.
  • Clear written communication, particularly when explaining what broke and what you changed.
  • Curiosity about the science. Experience in life sciences, biotechnology, or biochemistry is a plus but not a requirement, and you will work alongside strong in-house chemistry every day.
  • A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos.


Working at Matterworks

Given the cross-disciplinary and innovative nature of our work, effective collaboration and communication are critical to our progress. We operate in a flexible hybrid model that accommodates both fully remote team members and those who work full-time from our Somerville, MA office. While some positions may require regular in-person presence for hands-on work or local collaboration, many roles can be performed remotely with team members distributed across various locations.

Compensation and Benefits

Matterworks offers full-time employees a competitive base salary, stock options, and benefits (health & dental, vision, long- and short-term disability, life insurance, 401k with company match). Employees enjoy a flexible work & unlimited time away policy, commuter benefits and parking, regular team meals and outings, and company support for continued education/coursework and conference participation.

Similar Jobs

More Jobs at Matterworks, Inc

  • Data Engineer
    $110K — $130K *
    Somerville, MA 02145 (Middlesex County)
    Pharmaceuticals & Biotech
    Hybrid

More Pharmaceuticals & Biotech Jobs

Find similar Data Engineer jobs: