Senior Software Engineer, Data Platform

Matterworks, Inc

$130K — $155K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Professional experience in building production data systems and pipelines.
  • Proficient in Python and SQL for large-scale data processing.
  • Knowledge of Kubernetes-native orchestration and modern data lake technologies.
  • Experience designing stable identifiers for large, dynamic datasets.
  • Familiarity with integrating LLMs into production data workflows.
  • Skilled in using AI coding tools with a critical perspective on output accuracy.
  • History of taking projects from inception to a validated, running state.

Responsibilities

  • Own and scale data pipelines to effectively manage multiple petabytes of data.
  • Design systems for efficient data acquisition, storage, and serving.
  • Transform raw data into usable datasets with consistent schemas.
  • Automate quality checks to maintain system integrity and prevent regressions.
  • Manage user interfaces that various teams interact with for data services.
  • Oversee data operations to meet service level agreements and secure processing.

Benefits

  • Competitive base salary and stock options.
  • Health, dental, and vision insurance.
  • Short- and long-term disability and life insurance coverage.
  • 401k plan with company match.
  • Flexible work environment with hybrid options.
  • Unlimited time off policy and commuter benefits.
  • Support for continued education and participation in conferences.
Full Job Description
Position Overview

As a Senior Software Engineer you'll work to build the connective tissue of our data platform, in both what you build and how you build it. Design, build and scale systems to enrich data from raw samples and information into readily usable datasets enriched with biological context. The data produced by you and the team will serve our customers through both our ML research and model development activities as well as our product.

You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team.

Key Responsibilities
  • Build and Scale Data Contracts: Own the pipelines and system other teams consume from. Design and implement systems that scale to multiple petabytes of data in an effective way.
  • Serving and Cost at Scale: You will build and scale systems to acquire, store and serve data in a fast and affordable as it grows: data layout, featurization throughput, Kubernetes-native orchestration, and cost surfaced before it adds up.
  • Labels and Enrichment: Turn raw data into datasets people can use, with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies down to their spectra and underlying features.
  • Quality Gates: Automate quality checks that enable increasing capability without regression and promote only on a pass.
  • Interfaces People Use: Own the surfaces AI, chemistry, product, and agents call, from the SDK used to build datasets to the tools that expose platform capabilities.
  • Operations and Data Rights: Ensure effective operations of our data needs meeting our designed service level agreements, while providing high quality, provenance, and secure data processing in line with our customer needs.


About You
  • Significant professional experience building production data systems and pipelines. We level on scope and judgment rather than years.
  • Proficient in Python and SQL for large-scale data processing.
  • Proficient in Kubernetes-native batch orchestration and modern data lake technologies (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform). Airflow or Dagster experience transfers fine.
  • Demonstrated experience designing stable identifiers for a large, changing corpus, and building validation that gates a publish rather than reporting on it after the fact.
  • Experience putting an LLM or agent component into a production data path, including the eval loop, the gold set, and cost per record.
  • Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness.
  • A track record of owning work through to a running, validated system, including the unglamorous parts: fixing the malformed dataset, writing the backfill, debugging last night's bad publish.
  • Comfort with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar). Engineering depth is the requirement.
  • A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos.


Working at Matterworks

Given the cross-disciplinary and innovative nature of our work, effective collaboration and communication are critical to our progress. We operate in a flexible hybrid model that accommodates both fully remote team members and those who work full-time from our Somerville, MA office. While some positions may require regular in-person presence for hands-on work or local collaboration, many roles can be performed remotely with team members distributed across various locations.

Compensation and Benefits

Matterworks offers full-time employees a competitive base salary, stock options, and benefits (health & dental, vision, long- and short-term disability, life insurance, 401k with company match). Employees enjoy a flexible work & unlimited time away policy, commuter benefits and parking, regular team meals and outings, and company support for continued education/coursework and conference participation.

Similar Jobs

More Jobs at Matterworks, Inc

  • Data Engineer
    $110K — $130K *
    Somerville, MA 02145 (Middlesex County)
    Pharmaceuticals & Biotech
    Hybrid

More Information Technology Jobs

Find similar Senior Software Engineer, Data Platform jobs: