Hadrian

ML Platform Engineer

Hadrian$170K — $300K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience building and operating production ML infrastructure for multiple models or inference workloads.
  • Proficient in production-level Python and SQL, particularly in typing, testing, packaging, and API design.
  • Hands-on experience with Kubernetes, containers, and distributed systems' failure modes.
  • Background with model registries, feature systems, or CI/CD workflows in ML.
  • Strong practical judgment regarding latency, throughput, and infrastructure optimizations.
  • Proven ability to build stable interfaces and collaborate with diverse technical teams.

Responsibilities

  • Construct the production platform for safe model deployment in automated factories.
  • Develop shared batch and online serving for various workloads with clear SLAs.
  • Implement standardized release and evaluation processes including automated tests.
  • Ensure integrity of online feature serving and manage model integrity proactively.
  • Create operational tools for telemetry, incident management, and resource management.
  • Design APIs, SDKs, and templates for team-wide adoption without extensive support.

Benefits

  • Medical, dental, vision, and life insurance
  • 401(k) plan
  • Flexible vacation policy
Full Job Description
The Role

This is an ML infrastructure role at the core of Hadrian's technology stack. While Data Science, Operations Research, Vision, and Document AI teams build models, you will own the platform that ensures these models remain reliable, effective, and secure in production. You'll standardize our deployment patterns built around MLflow, Dagster, ECR, FastAPI, and EKS, making them the backbone for packaging, evaluating, releasing, serving, monitoring, and rolling back models across Hadrian's automated factories.

What You'll Do
  • Build the production platform that enables Hadrian's factories to safely depend on models for drawing extraction, cycle-time prediction, forecasting, scheduling, and more-with measurable performance and fast rollback.
  • Develop shared batch and online serving for tabular, vision, document-AI, scheduling, graph, and embedding workloads, targeting clear SLAs for latency, availability, and isolation.
  • Create repeatable release and evaluation processes featuring automated tests, reproducible artifacts, lineage, shadow deployments, canaries, and A/B tests.
  • Own online feature serving and maintain contract integrity with offline feature tables; proactively detect and address training-serving skew, feature drift, bad data, and model degradation.
  • Build operational tooling for telemetry, incident response, autoscaling, resource and GPU management, cost attribution, and secure model routing.
  • Develop APIs, SDKs, reusable templates, and documentation that teams can adopt without requiring close support from platform engineers.


What We're Looking For
  • Track record building and operating production ML infrastructure across multiple models or inference workloads.
  • Strong production-level Python and SQL skills, including typing, testing, packaging, API design, and building observability features.
  • Hands-on experience with Kubernetes, containers, and handling distributed-system failure modes such as retries, partial failures, idempotence, and resource isolation.
  • Engineering background with model registries, feature systems, batch/real-time inference, experiment tracking, or model CI/CD workflows.
  • Practical judgment around latency, throughput, availability, multi-tenancy, autoscaling, and infrastructure cost optimizations.
  • Ability to build stable interfaces and collaborate closely with engineering and scientific stakeholders.


What Will Set You Apart
  • Experience implementing feature stores (Feast, Tecton, or internal systems).
  • Production work with Ray Serve, KServe, Triton, BentoML, SageMaker, Vertex AI, or custom gRPC inference services.
  • Experience serving and evaluating vision, document-understanding, embedding, or generative pipelines.
  • Expertise in GPU inference optimization, multi-model serving, edge inference, or Go/Rust performance-sensitive AI services.
  • Background in regulated environments or open-source contributions to ML infrastructure projects (MLflow, Feast, KServe, Ray).


Compensation

Salary range: $170,000 - $300,000

This is the lowest to highest salary we reasonably and in good faith believe we would pay for this role at the time of posting. We may ultimately pay more or less than the posted range, and the range may be modified in the future. An employee's pay position within the salary range will depend on several factors, including relevant education, qualifications, certifications, experience, skills, geographic location, performance, and business needs.

Benefits
  • Medical, dental, vision, and life insurance
  • 401(k)
  • Flexible vacation policy


ITAR Requirements

To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR), you must be a U.S. citizen, lawful permanent resident, protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State.

About Hadrian

Hadrianadri?ja?n?s]; 24 January 76 – 10 July 138) was Roman emperor from 117 to 138. He was born in Italica, a Roman municipium founded by Italic settlers in Hispania Baetica and he came from a branch of the gens Aelia that originated in the Picenean town of Hadria, the Aeli Hadriani. His father was of senatorial rank and was a first cousin of Emperor Trajan. Hadrian married Trajan's grand-niece Vibia Sabina early in his career before Trajan became emperor and possibly at the behest of Trajan's wife Pompeia Plotina. Plotina and Trajan's close friend and adviser Lucius Licinius Sura were well disposed towards Hadrian. When Trajan died, his widow claimed that he had nominated Hadrian as emperor immediately before his death. Rome's military and Senate approved Hadrian's succession, but four leading senators were unlawfully put to death soon after. They had opposed Hadrian or seemed to threaten his succession, and the Senate held him responsible for their deaths and never forgave him. He earned further disapproval among the elite by abandoning Trajan's expansionist policies and territorial gains in Mesopotamia, Assyria, Armenia, and parts of Dacia. Hadrian preferred to invest in the development of stable, defensible borders and the unification of the empire's disparate peoples. He is known for building Hadrian's Wall, which marked the northern limit of Britannia.
Learn more about Hadrian

Similar Jobs

More Jobs at Hadrian

  • Hadrian
    Data Scientist
    $170K — $300K *
    Los Angeles, CA 90011 (Los Angeles County)
    Aerospace & Defense
    In-Person
  • Hadrian
    ML Platform Engineer
    $170K — $300K *
    Los Angeles, CA 90011 (Los Angeles County)
    Enterprise Technology
    In-Person
  • Hadrian
    Data Platform Engineer
    $170K — $300K *
    Los Angeles, CA 90011 (Los Angeles County)
    Information Technology
    In-Person
  • Hadrian
    Tech Lead, Offensive Security
    $150K — $180K *
    Los Angeles, CA 90011 (Los Angeles County)
    Aerospace & Defense
    In-Person
  • Hadrian
    Tech Lead, Detection & Response
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person

More Enterprise Technology Jobs

Find similar ML Platform Engineer jobs: