Software Engineer, AI Systems (United States)

AI Fund

• $125K — $150K *
US-AnywhereRemote in United States
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years in AI/ML engineering with an emphasis on LLM systems.
  • Hands-on experience with tool-calling workflows using LangGraph or LangChain.
  • Solid approach to evaluating LLMs with regression testing and human review loops.
  • Proven ability to assemble context for LLM systems.
  • Track record of owning production systems from deployment to incident response.
  • Experience with model evaluation across multiple providers and tradeoffs.
  • Familiarity with Neo4j and Cypher for graph data modeling, strongly preferred.
  • Proficient in Python, with emphasis on FastAPI and maintainable code practices.

Responsibilities

  • Build and operate multi-step LLM pipelines integrating various model and tool calls.
  • Extend and manage the orchestration of agent workflows through evidence to enterprise learning.
  • Design and implement effective context layers for Neo4j graph and hybrid retrieval systems.
  • Develop evaluation frameworks, datasets, and scoring systems for extraction and reasoning.
  • Monitor AI operations, implementing performance metrics and addressing quality assurance.
  • Select models across leading AI providers to inform product direction and AI roadmap.

Benefits

  • Work on meaningful reasoning problems where correctness is paramount.
  • Be part of an evaluation-first culture that prioritizes measurable quality.
  • See direct impact on customer safety teams in regulated industries.
  • Join a small, agile team providing high ownership and direct collaboration with leadership.
Full Job Description
About the Role:

You will work on the AI and LLM engineering layer that connects Haven's evidence base and knowledge graph to the product experiences investigators and safety leaders use. This is a build-and-operate role reporting to the CTO. You will design the reasoning, ship it, instrument it, and improve it using production evidence.

The work is technically demanding and operationally consequential. Haven serves enterprise customers in regulated, safety-critical industries through multi-tenant and dedicated deployments. A plausible answer and a correct answer can look the same until someone acts on it, so precision, traceability, evaluation, and human oversight are core product requirements.

Haven Safety is a VC-backed pre-seed venture. This role will be a important member of the founding team and will require wearing multiple hats. This role is based in the US and relocation will not be considered.

What You Will Own:

  • Production reasoning systems. Build and operate multi-step LLM pipelines that coordinate model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.
  • Agent orchestration. Extend Haven's coordinated agent team and the orchestration layer that carries an incident from evidence through analysis, review, and enterprise learning.
  • Grounding and retrieval. Design the context layer across Neo4j graph traversal, vector search, and hybrid retrieval so every model call receives the right evidence and organizational knowledge.
  • Evaluation. Build datasets, scoring, regression suites, model comparisons, human-label loops, and per-stage quality attribution for extraction and reasoning tasks.
  • Production AI operations. Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards; catch loops, hallucinations, and silent drift before customers do.
  • Technical direction. Select models by task across OpenAI, Anthropic, and Google; partner with product and knowledge engineering; and help shape the AI roadmap.


Requirements For the Role Include:

  • Production LLM systems. AI or ML engineering, including shipping LLM systems that real users depend on.
  • Agentic workflows. Hands-on experience building and debugging multi-step, tool-calling workflows with LangGraph, LangChain, or an equivalent framework.
  • Evaluation discipline. A repeatable approach to LLM evaluation, including representative datasets, regression testing, LLM-as-judge techniques, or human review loops.
  • Retrieval judgment. Experience assembling context for LLMs and a clear point of view on what to retrieve, how much, and why.
  • Production ownership. A track record of owning systems from deployment through monitoring and incident response, including a strong story about a failure or regression you diagnosed and fixed.
  • Model judgment. Comfort working across model providers and explaining tradeoffs in quality, latency, cost, context, and operational risk.
  • Graph reasoning. Comfort with Neo4j and Cypher, or a comparable graph store, and the ability to ramp quickly on graph data modeling. Strongly preferred.
  • Strong Python. Production habits around FastAPI, asynchronous services, testing, observability, and maintainable interfaces are required.


Nice To Haves Include:

  • Deep graph experience. Cypher fluency, schema evolution, MERGE patterns, embeddings, and operating a live knowledge graph.
  • Enterprise AI security. Prompt-injection awareness, context-leak prevention, tenant isolation, role-based access, and policy-layer separation.
  • Azure and hybrid search. Experience running production AI services in Azure and using Azure AI Search, Pinecone, MongoDB Atlas, pgvector, Elasticsearch, or a similar platform.
  • B2B Enterprise SaaS. Prior experience operating in an enterprise-level environment is strongly preferred.


Why Join Haven Safety:

  • Meaningful reasoning problems. Work across evidence, causal pathways, controls, organizational history, and corrective actions in a domain where correctness matters.
  • An evaluation-first culture. Make quality measurable, observable, and improvable instead of relying on demos or intuition.
  • Visible customer impact. Build for safety teams in energy, utilities, infrastructure, construction, and manufacturing, with direct feedback from the people using the output.
  • Small team, high ownership. Work closely with the CTO, product, and knowledge engineering, make consequential technical decisions, and see your work reach customers quickly.


Sponsorship will not be provided for this role. Please no agency inquiries!

Similar Jobs

More Jobs at AI Fund

  • Principal Builder
    $230K — $275K *
    Mountain View, CA 94040 (Santa Clara County)
    Enterprise Technology
    In-Person
  • AI Engineer
    $150K — $180K *
    Mountain View, CA 94040 (Santa Clara County)
    Education, Government & Non-Profit
    In-Person
  • Senior Associate Builder
    $185K — $225K *
    Mountain View, CA 94040 (Santa Clara County)
    Consumer Technology
    In-Person
  • Marketing Engineer
    $150K — $175K *
    Mountain View, CA 94040 (Santa Clara County)
    Consumer Technology
    In-Person
  • Design Engineer
    $110K — $130K *
    Mountain View, CA 94040 (Santa Clara County)
    Consumer Technology
    In-Person

More Enterprise Technology Jobs

Find similar Software Engineer, AI Systems (United States) jobs: