Software Engineer, Agent Infrastructure

Sapiom, Inc

$120K — $145K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of relevant engineering experience
  • Proven experience shipping production-ready software
  • Ability to adapt and solve problems with limited structure
  • Comfortable in backend, infrastructure, and tooling roles
  • Strong decision-making skills with incomplete information
  • Proficiency in using AI tools to enhance engineering practices
  • Familiarity with distributed systems and API integrations

Responsibilities

  • Build scalable routing and capacity solutions for model calls
  • Ensure metering and billing accuracy for customer consumption
  • Develop and implement reliability practices at scale
  • Manage permissions and workflow state within the control plane
  • Optimize agent execution across thousands of daily runs

Benefits

  • Flat organizational structure enhancing collaboration
  • Direct access to architecture decision-makers
  • Real opportunities for professional growth and skill enhancement
  • Emphasis on self-directed problem solving and innovation
  • Dynamic work environment with evolving challenges
Full Job Description
About the role

Agents can think now. They still can't act - not economically, not reliably, not with anyone in control. Teams build a great demo in days, then find that running it in production is brutally hard: it breaks when nobody's watching, the bill arrives before the explanation, and nobody can reconstruct what happened. Closing that gap is the whole company.

We're a flat org. Nobody has a lane, and what you work on will change as the problems change - you might be deep in routing one month and in metering the next. That's not disorganization; it's what a company this early looks like when the architecture is still being set.

We hire at two levels: engineers one to three years in, and Staff. There's no layer between you and the people setting the architecture, which is why someone this early gets real surface area here. It also means your design decisions get questioned by engineers who have made these mistakes before, which is the fastest way to get good at this.

What we're working on
  • Routing and capacity. Every model call has to be placed against latency, cost, and quality targets in real time, and the scheduling, reservations, and forecasting underneath have to keep those decisions honest as demand shifts. Getting this right is what makes running agents economical at all.
  • Metering and billing correctness. Two charging layers, hold and capture semantics, and 270M+ consumption events that all have to reconcile - in a ledger and in a usage number a customer is reading right now. Metering is the product here, not a feature, which means the correctness bar is unusually high.
  • Reliability at sustained scale. We've taken a 10-20x increase that kept compounding week over week for months. The open question is what reliability should mean here, and what observability, incident practice, and architecture get us ahead of the growth rather than reacting to it.
  • The control plane. Credential-scoped permissions, workflow state limits, gateway hardening. These are the decisions that determine how much authority a customer can safely hand an agent - largely unsettled, and consequential.
  • Agent execution. Sandboxed runs, state, memory, tools, and the external integrations agents depend on, at tens of thousands of runs a day.


You may be a fit if
  • You have strong engineering fundamentals and write good code quickly.
  • You've shipped something that ran in production and that you were responsible for when it broke - at a job, an internship, or a project of your own.
  • You take problems from "this should exist" to shipped without much structure around you.
  • You move between backend, infrastructure, and tooling as the problem requires, rather than staying where you're most comfortable.
  • You make reasonable calls with incomplete information and revise them when you learn more.
  • You close gaps fast - you'd rather learn the thing you're missing than route around it.
  • You use AI tools in your own engineering workflow - not instead of judgment, but as a multiplier.
  • Nice to have: experience with distributed systems, queues, storage, or observability; systems that talk to a lot of third-party APIs; LLM inference or agent architectures; anything you've built and operated yourself at real scale, including side projects.

If this reads like a stretch, apply anyway. At this stage we care more about how fast you close gaps than which ones you've already closed.

Applying

A recruiter screen, then a technical screen. If those go well, a three-part loop: an architecture deep dive, a hands-on AI project, and a conversation with our founder.

Everyone on our engineering team carries the title Member of Technical Staff internally. We post by level so the scope is clear.

Similar Jobs

More Jobs at Sapiom, Inc

More Information Technology Jobs

Find similar Software Engineer, Agent Infrastructure jobs: