Data Infra - Telemetry and Observability (IC)

Matter Intelligence

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience with ML observability, telemetry, or distributed training systems.
  • Strong engineering skills in observability, automation, and systems debugging.
  • Understanding of metrics, logs, traces, and incident diagnosis in AI.
  • Ability to analyze AI failure modes like bad datasets and training instabilities.
  • Experience designing telemetry schemas across complex systems.

Responsibilities

  • Build a telemetry model connecting mission states and AI components.
  • Instrument various elements of distributed AI training and deployment.
  • Manage workflows of AI agents in complex environments.
  • Connect mission data with hardware and model results for insight.
  • Establish alerts and escalation paths for operational issues.
  • Ensure visibility into capacity and costs across AI processes.

Benefits

  • Competitive compensation based on experience.
  • Early-stage equity package.
  • 100% employer-paid health, dental, and vision coverage.
  • Work on innovative AI systems with real-world impact.
Full Job Description
About the Role

Matter is hiring a Telemetry and Observability Engineer to make datasets, training runs, models, environments, agents, and missions observable as one AI system. Reporting to Ignacio Cases Martin, this individual contributor will connect flight and mission reality with infrastructure, model, and product behavior through shared telemetry, traces, alerts, reliability mechanisms, and operational context.
Key Responsibilities
  • Build a common telemetry model linking mission state, datasets, code, checkpoints, accelerators, model and prompt versions, environment episodes, agent steps, tools, decisions, cost, latency, and feedback.
  • Instrument distributed training, data loading, GPU utilization, checkpointing, experiment health, batch and online inference, and model-serving behavior.
  • Instrument agent workflows across prompts, context, retrieval, memory, planning, tool calls, graph state, human intervention, evidence, outcomes, and safety controls.
  • Connect planned-versus-actual mission state, hardware context, time, geometry, and geolocation to relevant datasets and model results.
  • Establish actionable service objectives, alerts, incident interfaces, rollout and rollback signals, and escalation paths across platform, model, environment, and agent failures.
  • Build capacity and cost visibility across storage, processing, accelerators, training, evaluation, inference, retrieval, and agent execution.
Qualifications
Required
  • Experience with ML observability, telemetry, time-series systems, distributed training or inference, cloud infrastructure, or production AI platforms.
  • Strong hands-on engineering skills in observability, automation, APIs, infrastructure, incident tooling, and systems debugging.
  • Understanding of metrics, logs, traces, events, model and data monitoring, service objectives, capacity, release safety, and incident diagnosis.
  • Ability to reason about AI failure modes including bad datasets, training instability, stale checkpoints, drift, version mismatch, environment bugs, and agent-tool failure.
  • Experience designing telemetry schemas and correlation workflows across multiple systems, time domains, and ownership boundaries.
Preferred
  • Experience with MLOps, GPU workloads, model serving, agent observability, reinforcement-learning environments, mission telemetry, or scientific instrumentation.
  • Experience operating high-throughput inference, streaming systems, geospatial pipelines, or mixed cloud and edge deployments.
  • Experience with reliability platforms or observability products used by multiple engineering teams.
  • Familiarity with aerospace, defense, regulated operations, or other settings requiring formal change control and incident evidence.
What Success Looks Like
  • Teams can correlate mission, data, infrastructure, model, environment, and agent behavior during normal operation and incidents.
  • Alerts and service objectives identify actionable failures without obscuring scientific or operational context.
  • Capacity, cost, version, and reliability signals support safe deployments and faster diagnosis across the AI stack.
Location

This role is based in San Francisco, CA, and requires onsite work.
ITAR Requirements

To comply with U.S. export regulations, applicants must be one of the following:
  • A U.S. citizen or national
  • A lawful permanent resident (green card holder)
  • Eligible to obtain required authorizations from the U.S. Department of State
Employee Offerings and Benefits

At Matter, we believe in rewarding high performance and providing the support you need to thrive. Our compensation and benefits package includes:
  • Competitive compensation based on experience
  • Early-stage equity package
  • 100% employer-paid health, dental, and vision coverage
  • Opportunity to work on novel sensing, data, and AI systems with real-world deployment paths to the largest industries in the world

Similar Jobs

More Jobs at Matter Intelligence

More Information Technology Jobs

Find similar Data Infra - Telemetry and Observability (IC) jobs: