AI Engineer

Tessera Labs

$130K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in production software development, especially with LLM systems.
  • Experience shipping agentic systems users rely on, including incident ownership.
  • Fluency with current agent toolkit and limitations in tool calling and orchestration.
  • Background in building retrieval systems from messy corpora, with concrete measurement of performance.
  • Solid experience in traditional ML techniques: supervised learning, feature engineering, evaluation.
  • Focus on evaluation as a core engineering practice rather than an afterthought.
  • A strong Python skillset and familiarity with TypeScript.

Responsibilities

  • Design and deploy production agents to facilitate enterprise transformations.
  • Build tools with strict access controls and documentation that empower agents to operate.
  • Enhance agent performance by improving context, tool strategy, and decision-making processes.
  • Develop retrieval strategies for handling enterprise content, including indexing and hybrid search.
  • Manage context efficiently to navigate historical enterprise data and decision-making.
  • Implement classical ML techniques where they provide better solutions than LLMs.
  • Create comprehensive evaluation and monitoring systems for production environments.

Benefits

  • Opportunity to work on mission-critical systems for a Fortune 500 company.
  • Impactful work with clear ownership over project outcomes.
  • A fast-paced environment with a focus on cutting-edge AI technologies.
  • Engaging with large-scale enterprise systems and complex challenges.
Full Job Description
About the role

Turning platform capability into outcomes a Fortune 500 will bet on has two halves - the systems that surround the model, and the model itself. This role owns the first.

You will own agents end to end: the harness they run in, the tools they call, the context they see, the guardrails around them, and the evals that tell us whether any of it is improving. Most of the difficulty is not model access. It's an agent reasoning across a landscape with nineteen years of undocumented decisions embedded in it, where one call silently returns a stale schema forty steps into a plan - and the work of making that failure legible, reproducible, and then impossible.

This is not a prototyping role. Everything you build gets pointed at systems a company's quarter close depends on.

If the model itself interests you more than the systems around it, look at Research Engineer - same team, other half of the problem.

What you'll do
  • Design and ship the production agents that do transformation work: understanding a landscape, planning a change, executing it across process, data, and code, and proving it was correct.
  • Build the tool layer - typed, permissioned, well-documented interfaces that let agents read and act across enterprise systems without ever exceeding what a human approver authorized.
  • Improve agent performance through prompting, context construction, tool-use strategy, and decision logic - and know which of those a given failure calls for.
  • Build the retrieval layer over enterprise artifacts and metadata: chunking and indexing strategies for content that doesn't resemble prose, hybrid search, reranking, grounding, and the evaluation that tells you whether any of it helped.
  • Manage context deliberately. Enterprise artifacts are enormous and a forty-step run accumulates history fast; deciding what the model sees, what gets compressed, and what gets dropped is a first-class engineering problem here, not a prompt detail.
  • Use classical ML where it's the right tool. Plenty of the work inside an agent pipeline - routing, ranking, classification, anomaly detection, confidence estimation - is better served by a small supervised model than by another LLM call, and knowing the difference is part of the job.
  • Build the eval and monitoring layer: task sets drawn from real customer landscapes, regression coverage on every deploy, and alerting that catches a changed schema, a revoked authorization, or a new model version before the customer does.
  • Instrument every run so each model call, tool invocation, decision, and human approval can be reconstructed afterward. This is what "passes audit" means in practice, and it's a product requirement rather than a nice-to-have.
  • Diagnose production failures from execution traces down to root cause, then close the gap in the system rather than in a single prompt.
  • Design the guardrails, approval gates, and rollback paths that make an agent safe to point at a live enterprise landscape.
  • Build capability that generalizes. Anything that only works for one customer is a bug in the product, not a feature of the engagement.
Representative projects
  • Shipping the agent that lets a customer retire a third of their custom estate in a quarter - and the review surface that makes a human comfortable approving each call in minutes rather than days.
  • Designing the approval and write semantics for agents operating in a live landscape, so nothing reaches production without a traceable human decision and no partial failure leaves systems inconsistent.
  • Building the reconciliation agent that makes two merged companies' data agree, plus the eval set that tells us when it's confident and when it's guessing.
  • Standing up the replay-evaluation pipeline that scores a candidate agent version against a month of recorded runs - real system responses, mocked writes - so a regression is caught before anything is touched in a customer environment.
  • Taking a pattern that worked at one customer and turning it into a platform primitive that works at the next four without a human rewriting it.
You may be a good fit if you
  • Have 3+ years building and operating production software, with meaningful recent time on LLM-powered systems.
  • Have shipped agentic systems that real users depend on, and have owned the incident when they broke.
  • Are fluent in the current agent toolkit - tool calling, orchestration, context engineering, RAG, evals, tracing - and can say where each one stops working.
  • Have built retrieval systems against messy real-world corpora, and have measured them rather than assumed them.
  • Have real traditional ML in your background - supervised learning, feature engineering, model evaluation - and reach for it when it beats a prompt.
  • Treat evaluation as engineering rather than a report you generate at the end.
  • Think in systems and customer outcomes, not model metrics.
  • Are comfortable with non-determinism and have opinions about building reliably on top of it.
  • Write strong Python and are comfortable in TypeScript.
  • Ship fast, and verify before you claim it works.
  • Find large, ugly, undocumented enterprise systems interesting rather than beneath you. That instinct is most of the job, and it's rarer than it should be.
Strong candidates may also have
  • Hands-on experience with enterprise platforms, their APIs, and their extension models - SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft, ServiceNow.
  • Experience with knowledge graphs, ontologies, or semantic models over messy real-world systems.
  • Experience with code analysis, program transformation, or automated refactoring at scale.
  • Experience with MCP, sub-agents, or agent skill/plugin architectures.
  • A background in distributed systems or workflow engines, where partial failure is the default case.
  • Experience at an early-stage startup or as a founder, where scope was whatever needed doing that week.
  • Familiarity with enterprise security and compliance realities: SSO, RBAC, segregation of duties, PII handling, SOC 2, data residency.
Join the Future of AI at Tessera Labs

We're looking for someone who enjoys building reliable, scalable systems that help teams move faster. If you take pride in cutting-edge AI, value clear ownership, and want to have a meaningful impact in a fast-moving environment, you'll fit right in.

Similar Jobs

More Jobs at Tessera Labs

  • AI Engineer
    $130K — $160K *
    San Jose, CA 95123 (Santa Clara County)
    Information Technology
    In-Person
  • AI Engineer
    $150K — $180K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Research Scientist
    $120K — $145K *
    San Jose, CA 95123 (Santa Clara County)
    Consumer Technology
    In-Person
  • Cloud Infrastructure/DevOps Engineer
    $120K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Cloud Infrastructure/DevOps Engineer
    $120K — $150K *
    San Jose, CA 95123 (Santa Clara County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar AI Engineer jobs: