Data Infra - Signals and Evaluations (IC)

Matter Intelligence

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience in ML evaluation and reinforcement-learning environments.
  • Strong foundations in statistics, experimental design, and benchmark validity.
  • Fluency in Python and modern machine-learning frameworks.
  • Experience converting qualitative model failures into reproducible datasets.
  • Ability to reason across the complete data path affecting correctness.

Responsibilities

  • Build versioned dataset workflows for curation and evaluation.
  • Design benchmark suites for various AI tasks.
  • Create simulation infrastructure for model training.
  • Define evaluation methods across multiple performance metrics.
  • Turn field and customer failures into test cases and scenarios.
  • Build automated promotion gates for release communication.

Benefits

  • Competitive compensation based on experience.
  • Early-stage equity package.
  • 100% employer-paid health, dental, and vision coverage.
  • Opportunity to work on cutting-edge AI systems with real-world applications.
Full Job Description
About the Role

Matter is hiring a Signal, Data, and Evaluation Engineer to build the datasets, benchmarks, environments, and evidence that make our models and AI systems scientifically defensible. Reporting to Ignacio Cases Martin, this individual contributor will work across data curation, simulation and learning environments, model and agent evaluation, and release-quality evidence.
Key Responsibilities
  • Build versioned dataset workflows for collection, curation, filtering, labeling, ground truth, synthetic data, hard-negative mining, contamination controls, and train-evaluation isolation.
  • Design benchmark suites across perception, multimodal reasoning, physics-informed prediction, retrieval, planning, tool use, world modeling, and customer tasks.
  • Create simulation and environment infrastructure for reinforcement learning, imitation learning, offline learning, model-based learning, and agent training.
  • Define repeatable evaluation methods across accuracy, calibration, generalization, robustness, safety, latency, cost, and operational outcomes.
  • Turn field, mission, and customer failures into versioned datasets, benchmark cases, adversarial tests, and environment scenarios.
  • Build automated promotion gates and communicate clearly what has been tested, what remains unproven, and what evidence supports release.
Qualifications
Required
  • Experience in ML evaluation, dataset engineering, reinforcement-learning environments, simulation, scientific ML, test infrastructure, or production AI systems.
  • Strong foundations in statistics, experimental design, benchmark validity, distribution shift, calibration, reward design, and evidence-based acceptance.
  • Fluency in Python and modern machine-learning frameworks, with the software engineering discipline to build tested modules, data contracts, or services.
  • Experience converting qualitative model or agent failures into reproducible datasets, tests, or environment scenarios.
  • Ability to reason across the complete data path and identify how acquisition, transformation, indexing, orchestration, and presentation affect correctness.
Preferred
  • Experience with dataset curation, benchmark platforms, deep reinforcement learning, world models, synthetic data, multimodal evaluation, or scientific validation.
  • Experience with PyTorch or JAX, distributed evaluation, experiment tracking, environment frameworks, or model and agent observability.
  • Experience with uncertainty quantification, calibration, out-of-distribution detection, conformal methods, or selective prediction.
  • Experience evaluating models deployed to aircraft, satellites, robotics, industrial systems, or other constrained environments.
What Success Looks Like
  • Model and agent releases are supported by reproducible datasets, benchmarks, and clearly stated evidence.
  • Failures from experiments and deployments become durable test cases that improve future systems.
  • Evaluation results preserve scientific meaning and make capability, uncertainty, and residual risk legible to the team.
Location

This role is based in San Francisco, CA, and requires onsite work.
ITAR Requirements

To comply with U.S. export regulations, applicants must be one of the following:
  • A U.S. citizen or national
  • A lawful permanent resident (green card holder)
  • Eligible to obtain required authorizations from the U.S. Department of State
Employee Offerings and Benefits

At Matter, we believe in rewarding high performance and providing the support you need to thrive. Our compensation and benefits package includes:
  • Competitive compensation based on experience
  • Early-stage equity package
  • 100% employer-paid health, dental, and vision coverage
  • Opportunity to work on novel sensing, data, and AI systems with real-world deployment paths to the largest industries in the world

Similar Jobs

More Jobs at Matter Intelligence

More Information Technology Jobs

Find similar Data Infra - Signals and Evaluations (IC) jobs: