AI-First SRE/DevOps Engineer

Axiad

$120K — $160K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-8 years of SRE, DevOps, or platform engineering experience.
  • Builder mentality focused on creating automation and tools.
  • Ownership of problems from ambiguity to resolution.
  • Strong operational experience with Kubernetes in production.
  • Demonstrated AI-First mindset and daily use of AI development tools.
  • Proficient in infrastructure-as-code, GitOps, and CI/CD workflows.
  • Experience with containerization and service mesh concepts.

Responsibilities

  • Own reliability and observability for a cloud-native Kubernetes platform.
  • Build CI/CD pipelines and infrastructure-as-code for rapid deployment.
  • Advocate AI-First operations by automating incident responses and remediation processes.
  • Create infrastructure components for AI-native features, including inference gateways.
  • Instrument systems for SLOs, error budgets, and tracing.
  • Harden platforms by managing secrets and ensuring supply-chain security.
  • Troubleshoot production issues with AI-powered tools and collaboration with engineers.

Benefits

  • Equity participation in a startup environment.
  • Fast-paced work culture that empowers ownership.
  • Opportunity to mentor and lead AI-first operational practices.
  • Direct collaboration with product and platform teams for impactful solutions.
Full Job Description
Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5-8 years of hands-on infrastructure and platform engineering experience to help build and run Mesh, our Identity Visibility and Intelligence Platform (IVIP) - a cloud-native microservices platform on Kubernetes spanning human identity, non-human identity (NHI), post-quantum cryptography, and agentic AI identity risk. The ideal candidate has a builder mentality and a strong AI-First mindset: automation and AI are the default, not the afterthought, and infrastructure is something you create, not just maintain.

This is a startup environment. You will own real surface area end-to-end, move fast, and ship. The role requires deep operational expertise in Kubernetes, CI/CD, and infrastructure-as-code, along with practical experience running AI/LLM systems in production. If your instinct when facing a repetitive task is to script it, agent-ify it, or delete it entirely - you'll fit right in.

Role Responsibilities
  • Own reliability, observability, and delivery for a multi-tenant, cloud-native Kubernetes platform - from design through production, yours to run and yours to improve.
  • Build (not just operate) CI/CD pipelines, infrastructure-as-code, and GitOps-driven progressive delivery that let a small team ship many times a day, safely.
  • Embrace and advocate AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the loop to triage, diagnose, and propose fixes where it makes sense. Treat toil as a bug.
  • Build the infrastructure that AI-native features run on: inference gateways, LLM cost/latency observability, prompt/version pipelines, eval harnesses, and guardrails for agentic workloads.
  • Instrument everything - SLOs, error budgets, and distributed tracing across services and data pipelines.
  • Harden the platform: secrets management, supply-chain security, and least-privilege everywhere.
  • Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling.
  • Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure - no throwing tickets over a wall; if you see a problem, it's yours to solve.
  • Mentor engineers in adopting AI-first operational practices and automation-by-default culture.

Skills and requirements
  • 5-8 years of professional experience in SRE, DevOps, or platform engineering roles.
  • Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them.
  • Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act.
  • Strong Kubernetes operational experience - running it in production, not just deploying to it.
  • Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf). Daily use of at least one AI development tool is a must.
  • Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred).
  • Cloud-native depth on at least one major cloud provider.
  • Solid observability expertise and SLO-driven operations experience.
  • Experience with containerization (Docker) and service mesh concepts.
  • Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams.
  • A bias for shipping - startup pace energizes you rather than stresses you.

Preferred Qualifications
  • Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration.
  • Data-pipeline and streaming/CDC experience.
  • Security or identity background; familiarity with post-quantum cryptography or supply-chain security.
  • Prior experience at an early-stage startup.

120,000 - 160,000 OTE + Equity + Benefits

Similar Jobs

More Jobs at Axiad

  • Solutions Architect
    $200K — $250K *
    San Jose, CA 95123 (Santa Clara County)
    Information Technology
    In-Person
  • GTM Engineer
    $140K — $175K *
    San Jose, CA 95123 (Santa Clara County)
    Business Services
    In-Person

More Enterprise Technology Jobs

Find similar AI-First SRE/DevOps Engineer jobs: