Infrastructure Engineer

Elicit

$185K — $260K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of hands-on infrastructure/SRE/platform engineering experience.
  • Proficient in Terraform as the primary infrastructure-as-code tool.
  • Strong Kubernetes expertise with production cluster operations experience.
  • AWS experience required, GCP familiarity is a plus.
  • Fluency in GitOps and CI/CD practices, particularly with Argo CD and GitHub Actions.
  • Experience building observability stacks and incident response processes.
  • Awareness of security and compliance frameworks like SOC 2.

Responsibilities

  • Own and manage cloud infrastructure on AWS and GCP including Kubernetes clusters and CI/CD processes.
  • Scale single-tenant deployments to meet diverse customer requirements.
  • Establish observability and incident response practices to enhance team efficiency.
  • Drive compliance and security operations to ensure adherence to established policies.
  • Manage infrastructure costs and capacity planning for optimal spend.
  • Enhance developer experience through improvements to CI/CD performance and tooling.
  • Collaborate on backend systems where infrastructure intersects with application logic.

Benefits

  • Flexible work environment with options for remote or office-based work in Oakland, CA.
  • Fully covered health, dental, vision, and life insurance, including family options.
  • Generous flexible vacation policy with a minimum recommendation of 20 days per year.
  • Monthly wellbeing stipend of $200 for health and wellness support.
  • 401K plan with a 6% employer match.
  • Budget for workstation setup including a Mac and ongoing financial support.
  • Quarterly budget of $1,000 for AI experimentation and learning opportunities.
Full Job Description
Why we9re hiring for this role

Elicit is an AI research platform used by scientists, pharma companies, and decision-makers for high-stakes evidence synthesis. A single session could trigger hundreds of thousands of language model invocations across multiple providers, which means our infrastructure decisions directly impact cost, reliability, and the quality of research outcomes for users making decisions worth millions of dollars.

Our infra is well-architected using best practices - Terraform, Kubernetes, Argo CD, GitHub Actions. But we9re at an inflection point: enterprise contracts are getting larger, single-tenant deployments are multiplying, and the surface area that needs dedicated attention has outgrown what our current team can cover part-time. This is the first dedicated infrastructure hire and you9ll define how this function works at Elicit.

James (Head of Engineering, ex-Square) set up the original infrastructure and will be your close partner. This role will own and evolve the infrastructure platform that underpins Elicit9s product. You will ensuring it is reliable, secure, cost-efficient, and ready for the demands of a growing enterprise customer base. Under your ownership, our systems will scale gracefully across single-tenant deployments, our SLAs will be backed by real engineering rigor rather than best intentions, and our compliance posture will be a selling point rather than an afterthought.

What you9ll own
  • Own our cloud infrastructure across AWS and GCP - Kubernetes clusters, networking, databases (Aurora PostgreSQL, Redis, MongoDB Atlas), Cloudflare, and our CI/CD pipeline.
  • Scale single-tenant deployments from a handful to many - each with distinct data retention, geographic, monitoring, and compliance requirements. Make a private cloud deployment a repeatable, low-overhead operation.
  • Build our observability and incident response practice - proactive monitoring, alerting, SLA tracking, and structured post-mortems that make the whole team better at diagnosing and resolving issues.
  • Drive compliance and security operations - ensure we follow through on the policies we9ve written (SOC 2, NIST AI framework, EU Cyber Resilience). Own disaster recovery exercises, database restoration drills, and security event monitoring (SIEM).
  • Manage infrastructure cost and capacity - make smart decisions about where we run workloads (AWS, CoreWeave, Parasail), optimize spend, and plan capacity as usage grows.
  • Improve developer experience - CI/CD pipeline performance, preview environments, local development tooling, and deployment confidence.
  • Contribute to backend systems where infrastructure and application intersect - circuit breakers, inference routing, data connector infrastructure for enterprise customers bringing their own data.


What success will look like (6-12 months)
  • A private cloud deployment is a ~1-day turnkey operation. Playbooks and templated Terraform make standing up Elicit in a customer9s cloud routine, which opens up 8-figure enterprise deals.
  • Our observability signal:noise ratio improves 10-fold. Health monitors cover every endpoint and job, and an alert firing means something needs attention.
  • Disaster recovery is practiced. We run database restoration drills and provider-outage dry runs on a schedule, with post-mortems that make the whole team better at diagnosis.
  • Our SLAs are backed by engineering rigor. We follow through on SOC 2, NIST AI framework, and EU Cyber Resilience commitments, and enterprise security reviews go faster because of it.
  • Inference is faster and cheaper. You9ve found and executed opportunities like shifting load between providers to cut p95 latency and cost at the same time.


What we9re looking for
  • 5+ years of hands-on infrastructure/SRE/platform engineering experience.
  • An AI-native way of working. Agentic coding tools (Claude Code, Cursor, Devin, etc.) are how we build at Elicit, and infrastructure is no exception: agents help us write IaC and investigate incidents. You should be an enthusiastic practitioner who uses AI to multiply your impact, and excited to find new places agents can safely take on infrastructure work. Bonus: you9ve written about, spoken about, or built projects demonstrating this.
  • Solid Terraform experience. this is our primary infrastructure-as-code layer and the most important technical requirement.
  • Strong Kubernetes expertise. You9ve operated production clusters, not just deployed to them. Comfortable with EKS, networking, autoscaling (Karpenter), and debugging cluster-level issues.
  • AWS experience (primary), with GCP familiarity a plus.
  • GitOps and CI/CD fluency. Argo CD, GitHub Actions, or equivalent. You understand deployment automation, rollback strategies, and change management.
  • SRE mindset. You9ve built or significantly improved observability stacks (DataDog or equivalent), incident response processes, and on-call practices.
  • Security and compliance awareness. Experience with SOC 2 or similar frameworks, SIEM tooling, and translating compliance requirements into engineering practice.
  • Ability to write software. You can contribute to our backend codebases where infrastructure meets application logic.


Am I a good fit?

Strong applicants will find it easy to answer these questions:
  1. Can you describe a time you designed and executed a multi-tenant or single-tenant deployment architecture for enterprise customers?
  2. How have you approached disaster recovery planning and testing at a previous company?
  3. Walk me through how you9d evaluate whether to build vs. buy for a new infrastructure component at a ~30-person startup.
  4. Have you owned compliance follow-through (not just policy writing) for a framework like SOC 2?
Who will I work with?
  • James (Head of Engineering): set up the original infrastructure and will be your closest partner. You9ll own execution, with James as sounding board and advocate.
  • Panda: the engineer who has been covering infrastructure part-time, with deep context on our cluster bootstrapping, Cloudflare setup, and inference providers.
  • Product: PMs covering the core product, ML, and evals. They carry the customer side of enterprise deployments, so you9ll work together to turn requirements like data residency, compliance commitments, and SLAs into architecture, and to weigh the cost and latency tradeoffs behind product decisions. Eval infrastructure is a shared surface with Ben, from CI integration to inference capacity.
  • The whole engineering team: we9re ~30 people company-wide, so you9ll work directly with the engineers whose developer experience you9re improving, and pair with them where infrastructure meets application code.
  • Andreas and Jungwon (cofounders): you9ll meet both during the interview process, and infrastructure decisions with strategic weight (enterprise deployments, compliance posture) get their direct attention.
On-call & incident expectations

We don9t have a formal on-call rotation yet. Incidents today are handled by the engineers closest to the affected system. Part of this role is building the incident response practice we should have: sensible alerting, SLA tracking, structured post-mortems, and eventually a rotation designed so it doesn9t burn anyone out. You9d design the on-call setup you9ll then live with.

Why join us?

The work matters. Our mission is to radically improve reasoning for high-stakes decisions. Over 2 million people use Elicit, including pharma teams making R&D decisions worth tens of millions of dollars, and the reliability of our platform is part of what makes those decisions sound.
  • Infrastructure work is close to the business. Single-tenant deployments open enterprise deals, and inference routing choices show up directly in our costs and latency. You9ll see the effect of your work in the company9s trajectory.
  • Built for the long term. We9re a Public Benefit Corporation that spun out of a non-profit AI research lab. Long-term impact and AI safety are part of the corporate charter.
  • High agency, low bureaucracy. ~30 curious, slightly weird people who write things down and trust each other to run with a vague brief.
  • Serious investment in your growth. $1,000 per quarter for every person to explore AI tools, courses, and events, plus quarterly in-person team retreats.


Location and travel

We have a great office in Oakland, CA, and we9d love to see you there if you9re local. That said, we9re just as happy for you to work remotely. We do get the whole team together for a quarterly retreat somewhere fun, because in-person time matters to us.

Benefits and perks

In addition to working on important problems as part of a productive and positive team, we also offer great benefits (with some variation based on location):
  • Flexible work environment: work from our office in Oakland or remotely with time zone overlap (between GMT and GMT-8), as long as you9re comfortable traveling for quarterly in-person offsites.
  • Fully covered health, dental, vision, and life insurance for you, generous coverage for the rest of your family (FSA/HSA, too).
  • Flexible vacation policy, with a minimum recommendation of 20 days / year and plenty of company holidays.
  • Every Elician receives a $200 monthly wellbeing stipend to spend on whatever supports your health and wellbeing.
  • 401K with a 6% employer match.
  • A new Mac + $1,000 budget to set up your workstation or home office in your first year, then $500 every year thereafter.
  • $1,000 quarterly AI Experimentation & Learning budget, so you can freely experiment with new AI tools to incorporate into your workflow, take courses, purchase educational resources, or attend AI-focused conferences and events.
  • A team administrative assistant who can help you with personal and work tasks.
  • You can find more reasons to work with us in this thread!


Compensation

For all roles at Elicit, we use a data-backed compensation framework to keep salaries market-competitive, equitable, and simple to understand. For this role, we target starting ranges of:
  • Senior (L4): $185-260k + equity
  • Expert (L5): $250-280k + equity
  • Principal (L6): >$260 + significant equity

We9re optimizing for a hire who can contribute at a L4/senior-level or above.

We offer above-market equity for all roles at Elicit, as well as employee-friendly equity terms.

Similar Jobs

More Jobs at Elicit

  • Infrastructure Engineer
    $185K — $260K *
    Oakland, CA 94601 (Alameda County)
    Information Technology
    Hybrid
  • Infrastructure Engineer
    $185K — $260K *
    Remote
    Information Technology
    Remote in United States
  • Engineering Manager
    $250K — $300K *
    Remote
    Information Technology
    Remote in United States
  • Engineering Manager
    $250K — $300K *
    Oakland, CA 94601 (Alameda County)
    Information Technology
    Hybrid
  • AI Engineer
    $195K — $260K *
    Oakland, CA 94601 (Alameda County)
    Information Technology
    Hybrid

More Information Technology Jobs

Find similar Infrastructure Engineer jobs: