ML Data Platform Engineer

Parisi Labs

$175K — $245K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years experience in data engineering or machine learning data systems.
  • Strong proficiency in Python and SQL.
  • Familiarity with data quality tooling and workflow systems.
  • Experience with object storage and columnar formats.
  • Ability to work autonomously on small teams.

Responsibilities

  • Build and enhance data ingestion and validation processes.
  • Define data contracts and semantics for model training.
  • Create reusable workflows for integrating new data sources.
  • Establish data quality and access controls.
  • Develop efficient datasets for machine-learning usage.
  • Diagnose failures in data and model pipelines.
  • Collaborate with engineers to convert data requirements into reliable software.

Benefits

  • Location flexibility in New York City or Boston/Cambridge.
  • Opportunity for meaningful early-stage equity.
  • Medical and dental benefits provided.
  • Focus on collaboration and in-person team engagement.
Full Job Description
About The Role

We are looking for an ML Data Platform Engineer to make the data behind our models and products dependable, understandable, and easy to use.

This role sits where data engineering meets machine learning. You will turn messy, changing real-world sources into durable datasets and interfaces that researchers, product engineers, and customer-facing technical teams can trust. The goal is not to build a large platform for its own sake. It is to make each new model, product surface, and authorized data source faster to bring online without compromising correctness.

What You Will Own

- Build and improve ingestion, backfill, validation, and observability for high-volume, time-dependent data.

- Define clear data contracts and point-in-time semantics for model training, evaluation, and product use.

- Create reusable workflows for bringing new public and customer-authorized sources into the system.

- Build quality, lineage, freshness, and access controls that make data trustworthy in repeated use.

- Develop efficient datasets and query interfaces for machine-learning and product workloads.

- Diagnose whether failures originate in source data, transformations, model inputs, or serving systems.

- Work closely with research and product engineers so data requirements become reliable software, not recurring manual projects.

- Exercise judgment about which abstractions should become durable infrastructure and which should remain purpose-built.

First 90 Days

- 30 days: Understand the data lifecycle behind Ask The Grid and our ML work; identify the most consequential reliability and usability gaps.

- 60 days: Ship a reusable ingestion, backfill, validation, or dataset primitive used in active product or research work.

- 90 days: Own a dependable end-to-end data workflow, with documented contracts, quality checks, and clear operational visibility.

Requirements

You May Be A Fit If

- You enjoy making difficult real-world data useful, not merely moving it between systems.

- You understand how warehouse or streaming data becomes training, evaluation, and product data.

- You care about temporal correctness, reproducibility, leakage, lineage, and source rights.

- You can design practical systems without reaching immediately for a large-company platform.

- You are comfortable debugging incomplete APIs, changing schemas, and surprising data behavior.

- You communicate clearly with researchers, product engineers, and customer-facing teammates.

- You want substantial ownership on a small team and can make progress without a mature data organization around you.

Helpful Background

- Strong Python and SQL experience.

- Experience with data engineering, ML data systems, dataset engineering, platform engineering, or high-quality analytics engineering.

- Experience with object storage, warehouses, streaming or workflow systems, columnar formats, APIs, and data-quality tooling.

- Experience preparing data for model training, evaluation, or scientific computing.

- Familiarity with time-series, geospatial, weather, market, event, or other temporally sensitive data is useful.

- Startup or small-team experience is helpful, but evidence of unusually strong ownership matters more than a particular company background.

Benefits

Location And Working Style

New York City or Boston/Cambridge. We work together in person regularly and will choose the home office based on the strongest candidate and their closest collaborators.

Compensation And Benefits

Base salary range: $175K-$245K, plus meaningful early-stage equity. Parisi Labs provides medical and dental benefits. Final compensation depends on level, location, experience, and role scope.

Interview Process

- Founder conversation with the CTO.

- Technical working session on a representative data-platform problem.

- Collaboration conversation with a research or product teammate.

- In-person final.

- Offer review.

Similar Jobs

More Jobs at Parisi Labs

More Information Technology Jobs

Find similar ML Data Platform Engineer jobs: