Agentic AI Engineer (Xora Portfolio Company)

Xora Innovation

$130K — $160K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science or related field
  • 5+ years in production software development focused on LLM or agent systems
  • Strong proficiency in Python with good engineering practices
  • Hands-on experience in building agent systems and orchestrations in production
  • Familiarity with multiple model providers and cost/latency trade-offs
  • Experience with end-to-end retrieval systems and LLM evaluation
  • Ability to manage complex systems in a fast-paced, early-stage environment

Responsibilities

  • Build a flexible provider abstraction for model workflows and outputs
  • Develop an agent orchestration for stateful, parallel execution of sub-agents
  • Implement human-in-the-loop checkpoints for critical workflow steps
  • Create a clear interface between agent capabilities and underlying systems
  • Design and develop the retrieval process from end to end
  • Establish a versioned prompt layer with tracked reasoning
  • Instrument agent activity and model interactions for debugging and observability
  • Build an evaluation framework to monitor performance and catch regressions

Benefits

  • On-site or hybrid work model based on location
  • Opportunity to work in a fast-paced, innovative environment
  • Exposure to cutting-edge LLM technologies and frameworks
  • Collaboration with a team dedicated to high-quality software engineering
  • Flexible working arrangements to support work-life balance
Full Job Description
ABOUT THE ROLE

This role builds the layer every LLM-powered feature on our platform is built on: one interface over many model providers, the agent orchestration that runs multi-step workflows, the retrieval and prompt systems that ground them, and the tracing and evaluation that keep them honest. It is foundational and deeply hands-on.

In this role, you'll design agents that plan, call tools, and reason over the results, then make them dependable in production: traceable, evaluated, and safe to run with a person in the loop where it counts. Because the platform runs inside customers' own secure environments, on their compute clusters, in their cloud, or across a hybrid of the two, the reasoning it produces has to stay auditable and its guardrails have to travel with the software.

Every LLM feature the team ships stands on this layer: what it can do sets what they can build, and how faithfully it reasons sets how far they can trust it.

WHAT YOU WILL DO
  • Build the provider abstraction that lets any workflow call, swap, or add a model provider by configuration, across commercial APIs and self-hosted endpoints, with structured-output validation, retries, and cost tracking.
  • Build the agent orchestration where a planning agent dispatches specialized sub-agents in parallel on a stateful framework, with durable checkpoints, conditional branching, and the context and memory management that keeps multi-step workflows coherent across long task horizons.
  • Build human-in-the-loop checkpoints so low-confidence or high-stakes steps route to a person before an agent proceeds.
  • Wrap existing platform capabilities as typed, registered tools the agents call, with a clean boundary between the agent layer and the systems it builds on.
  • Design retrieval end to end, from ingestion, embeddings, and chunking through hybrid search and reranking, and assemble the context that grounds each model call.
  • Build the prompt layer: versioned prompts, few-shot sets, and captured reasoning, so every change is tracked and every call is inspectable.
  • Expose agents and guardrailed model access as tools behind one integration point that backend services, the frontend, and notebooks all consume.
  • Instrument every model call, tool invocation, and agent run as traced spans with prompt, model, and tool lineage, so behavior and cost stay debuggable.
  • Build the evaluation framework, deterministic trace metrics alongside LLM-as-judge scoring for faithfulness, that gates changes and catches regressions before they ship.


WHAT WE ARE LOOKING FOR
  • Bachelor's or Master's degree in Computer Science or a related engineering field, and 5+ years building and shipping production software, with real depth building LLM or agent systems in production.
  • Strong Python and solid engineering practice: async code, typing, testing, modular design, and code review, plus a track record of shipping systems others depend on.
  • Hands-on experience building agentic or LLM systems in production: orchestration loops, tool-calling, structured outputs, and context and memory management for reliable long-running workflows.
  • Experience working across multiple model providers behind a single abstraction, with routing, fallback, and a feel for the cost and latency trade-offs.
  • Experience building retrieval systems end to end: embeddings, chunking, hybrid search, reranking, and vector databases.
  • Experience with LLM evaluation and guardrails: building eval sets and harnesses, LLM-as-judge scoring, regression gating, and output-quality and safety checks.
  • Experience instrumenting LLM systems for observability: tracing model and tool calls, versioning prompts, and using traces to debug and improve real behavior.
  • Comfort owning ambiguous systems end to end in a fast-moving early-stage environment.


NICE TO HAVE
  • Stateful agent-orchestration frameworks such as LangGraph or AutoGen, and durable-execution engines such as Temporal for long-running workflows.
  • Experience building MCP tools or servers, or similar tool-calling integration layers.
  • LLMOps and evaluation tooling such as MLflow or Langfuse for tracing, prompt versioning, and evaluation.
  • Human-in-the-loop and interrupt-driven agent patterns for review and control.
  • Applying LLMs to scientific or technical workflows, grounding reasoning in tool outputs and structured data.
  • Fluency with modern AI coding assistants, or open-source contributions to AI or agent tooling.


LOCATION

Singapore or United States. We're hiring in both to reach the right person. Work model is on-site or hybrid, set per location.

CLOSING NOTE

If you don't tick every box but this is clearly your kind of work, get in touch.

Similar Jobs

More Jobs at Xora Innovation

More Information Technology Jobs

Find similar Agentic AI Engineer (Xora Portfolio Company) jobs: