What You'll DoAs a Senior Machine Learning Engineer on the Defensive Agent team, you'll be the person who gets our defensive models into production and keeps them there. Our AI researchers define how NodeZero's agents should reason. You build the pipelines, serve infrastructure, and release machinery that turns that research into something running against thousands of customer tenants every day.
This is the seam where most AI products fail. A model that performs well in a notebook is not a capability. It becomes one when there is a reproducible training pipeline, an evaluation suite that runs in CI, a versioned release path with canaries and rollback, monitoring that catches regressions before customers do, and a cost and latency profile the business can afford. That is your work.
You'll work directly with our AI researchers and closely with backend and infrastructure engineers. You are not being hired to do research, and you are not being hired to do generic platform work. You own the path from model to production.
Responsibilities- Build and own the training and post-training pipelines - data preparation, fine-tuning and preference optimization runs, experiment tracking, artifact management, and reproducibility.
- Build the inference and serving layer: model gateway with provider routing and fallback, regional pinning for data residency, batching, and caching.
- Own the release path for model-layer artifacts: version prompts, model selections, and tool definitions as deployable config; run shadow and canary deployments by tenant; make rollback fast and boring.
- Build monitoring for the model layer - behavioral drift, regression detection, output quality signals, latency, and per-tenant token and cost accounting with budget enforcement.
- Build the data and context pipelines that feed inference, including retrieval and embedding infrastructure over attack path, configuration, and remediation data, with tenant isolation enforced end to end.
- Optimize cost and latency across the inference path, and make the tradeoffs visible so product decisions are made with real numbers.
- Develop core product features in ETL and GraphQL where model outputs, run history, and evaluation results need to reach the product and internal tooling.
- Partner with researchers to move prototypes into production, and feed production constraints and failure data back into research direction.
Required Education / Experience- Bachelor's Degree in Computer Science, Computer Engineering or related field, or equivalent practical experience.
- 5+ yrs professional software engineering experience, with strong production Python.
- Demonstrated experience taking ML or LLM-backed systems from prototype to production and operating them.
- Hands-on experience with ML pipelines and tooling: training or fine-tuning workflows, experiment tracking, artifact and model registries, and reproducible data preparation.
- Experience building and operating inference or model-serving infrastructure in production, including latency and cost optimization.
- Experience building applications on cloud computing platforms such as AWS, Azure, GCP, using container technologies such as Docker and Kubernetes.
- Solid proficiency in SQL and experience with production data pipelines.
What Sets You Apart- Experience operationalizing LLM or agentic systems specifically.
- Hands-on post-training experience: supervised fine-tuning, distillation, preference optimization, or RL, including the infrastructure around the runs.
- Experience with model gateways or multi-provider routing, and with self-hosted or customer-hosted inference (vLLM, TGI, Bedrock, or similar).
- Experience with GPU infrastructure, quantization, or inference optimization.
- Experience with database architectures including relational (PostgreSQL) and graph (Neo4j), and with GraphQL backends.
- Experience with observability tooling (Datadog, Prometheus, Grafana) and distributed tracing, including tracing across model calls.
- Experience shipping ML into regulated, air-gapped, or customer-controlled environments, or under compliance regimes like FedRAMP.
Perks of Horizon3- Inclusive Team: We value diversity and promote an inclusive culture where everyone can thrive.
- Growth Opportunities: Be part of a dynamic and growing team with numerous career development opportunities.
- Innovative Culture: Work in a collaborative environment that encourages creativity and out-of-the-box thinking.
- Hybrid & Remote Work: We embrace a mix of remote and hybrid work models depending on role and location, including our Chicago office, where some roles require regular in-office presence.
- Competitive Compensation: We offer competitive salary, equity and benefits. Our benefits include health, vision & dental insurance for you and your family, a flexible vacation policy, and generous parental leave.
Compensation and ValuesAt Horizon3, we believe that our people are our greatest asset, and our compensation philosophy reflects this core value. We are committed to fostering an environment where all employees feel valued, respected, and rewarded for their contributions. Our compensation structure is designed to be fair, competitive, and transparent, ensuring that every team member is recognized and compensated equitably across roles, levels, and locations.
In accordance with various State's transparency regulations, we provide the following salary range information for this position:
- Base salary range: $211,000 - $249,000 annually. The exact salary will be determined based on the selected candidate's location, qualifications, experience, and relevant skills.
- Additional compensation: All full-time roles are eligible for an equity package in the form of stock options.
Other DutiesPlease note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities, and activities may change at any time with or without notice.