Staff Platform & Reliability Engineer

Interface AI

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of hands-on experience with production systems and real uptime commitments
  • Expertise in Kubernetes on AWS and GitOps delivery
  • Strong background in infrastructure-as-code at multi-account scale
  • Experience building SLO and error-budget practices that impact release decisions
  • Proven disaster recovery capabilities with defined RTO/RPO
  • Proficient in security practices related to IAM, secrets management, and compliance
  • Familiarity with advanced AI tools and daily use of AI in operations
  • Programming proficiency in TypeScript and/or Python and Bash, with clear documentation skills

Responsibilities

  • Define and manage reliability metrics (SLIs and SLOs) for the platform
  • Own the disaster recovery strategy and ensure real failover tests
  • Implement a GitOps-based deploy path with automated rollback
  • Manage cloud infrastructure as code and optimize for cost and scale
  • Oversee end-to-end incident management processes
  • Establish observability frameworks, including alerting and logging
  • Automate operations for a small team using AI-driven solutions

Benefits

  • 100% paid health, dental, and vision care
  • 401(k) and financial wellness perks
  • Daily meals provided
  • Commuter benefits offered
  • Monthly wellness stipend
  • Access to cutting-edge AI tools
  • Brand-new office space with impressive views
  • Discretionary PTO and paid parental leave
  • Comprehensive mental health and wellness benefits
  • Opportunity to work closely with executive leadership in a mission-driven environment
Full Job Description
The Role

You own how the platform ships and how it stays up - the reliability and security backbone behind every conversation on interface.ai.

Our customers are banks and credit unions, and the bar for trust is high: real money, real regulators, and the uptime our customers count on. This role holds the platform to the enterprise-grade, 99.99% reliability we're known for and sets the engineering practices that keep it there as we scale. You'll define the SLOs, own the deploy path, run incidents when they happen, and build the platform so a small, senior team working with AI tooling can operate it with confidence - and you'll set the reliability and security bar the rest of the org builds against.

This is a hands-on, senior-most IC seat. What we mean by senior isn't tenure - it's judgment: you've run production systems people depend on, you think in failure modes and containment, and AI tooling is already part of how you work every day.
What You'll Own
  • Reliability & SLOs - define customer-facing SLIs and SLOs across product surfaces, run an error-budget program that governs release decisions, and hold the platform to its reliability bar.
  • Resilience & disaster recovery - own the DR strategy: regional failover, written RTO/RPO per tier, resilience against third-party dependency failure, and recovery you've actually tested rather than just planned.
  • Deploy & delivery - a GitOps deploy path with progressive delivery, automated analysis, and one-click rollback for every service, plus drift detection and a full change audit trail.
  • Cloud & infrastructure-as-code - the AWS foundation, fully managed as code; the Kubernetes platform and service mesh; capacity, cost, and scale.
  • Incident management - end to end: paging and severity policy, an incident-commander rotation, status-page automation, blameless post-mortems with tracked actions, and SLA reporting to customers.
  • Observability - metrics, logging, and distributed tracing across services; dashboards and alerts as code; burn-rate alerting that pages on what actually matters.
  • AI-native operations - build the automation and guardrails that let a small, senior team operate the platform with AI in the loop: runbooks encoded for safe automation, self-healing for routine work, and golden-path templates that ship new services with SLOs, alerts, secrets, and policy built in.
Security & Compliance

You own the infrastructure's security posture - including the parts unique to running an AI platform in regulated financial services.
  • Cloud & cluster security - IAM least-privilege across accounts, secrets management with a single source of truth, admission control and network policy on every cluster, image signing / SBOM / scanning in CI, and cloud-audit and threat-detection feeds as first-class alert sources.
  • Compliance evidence - own the infrastructure evidence for SOC 2 Type II and for customer and regulator reviews (NCUA/FFIEC examinations, GLBA): change management, access reviews, backup/restore, and DR test records, automated wherever possible.
  • Securing AI systems - protect the platform against AI-specific threats such as prompt injection, data exfiltration through tools, tenant isolation for retrieval stores, and abuse detection on public-facing endpoints - and set the guardrails and credential boundaries for AI-assisted engineering workflows.
What We're Looking For
  • A senior-most, hands-on IC who has run production systems with real uptime commitments - and carried the pager for them.
  • Deep production Kubernetes on AWS, including service mesh, with GitOps-based delivery across many services.
  • Infrastructure-as-code at multi-account scale, including taking over and reshaping a large existing estate.
  • You've built an SLO and error-budget practice that actually changed release decisions, with alerting tuned to burn rate rather than noise.
  • You've delivered multi-region or DR capability with defined RTO/RPO and proven it with real failover tests.
  • Security as daily practice, not a checklist - least-privilege IAM, secrets management, admission and network policy, software supply-chain controls - and you've produced evidence that satisfied auditors (SOC 2, ISO 27001, PCI, or FFIEC-style exams).
  • Extreme AI fluency - you use frontier AI tools (Claude Code, Cursor) daily and have clear opinions about the boundaries agents should operate within.
  • Strong programming in TypeScript and/or Python, plus Bash - and writing clear enough that a regulator could follow your post-mortem.
  • BS/BA in Computer Science required; MS or PhD a strong plus. San Francisco-based and committed to working onsite; on-call participation. H1B transfers welcome.
Bonus Points
  • Real-time voice or telephony systems, or other latency-critical streaming workloads.
  • Operating streaming and analytical data platforms and durable workflow engines at scale.
  • Chaos engineering and running resilience exercises against live environments.
  • Regulated-industry background (fintech, banking, healthcare), including vendor-risk management.
  • Threat modeling for LLM and agentic applications (e.g., OWASP LLM Top 10) and defending tool and agent integrations.
  • Cloud cost engineering, including compute-fleet optimization and LLM spend attribution.
What This Role Is - And Isn't

This is a hands-on, senior-most IC seat - not a management track, and not a hands-off "architect" role. You'll be measured by what you ship, what stays up, and the reliability and security bar you set for the engineers around you. If you want the title without owning the pager, this isn't it.

It's also not a role for someone who needs a mature SRE org and a runbook for everything handed to them. You'll set the standards here, not inherit them, in a space moving faster than most companies can track. If that's the leverage point you've been looking for, we want to talk.
Benefits

100% paid health, dental & vision care
401(k) & financial wellness perks
Daily meals on us
🚇 Commuter benefit
Monthly wellness stipend
Claude Enterprise + frontier AI tools for every employee - build with the best
🏙 Brand-new 21st-floor SF office at 44 Montgomery - floor-to-ceiling views, worth showing up to
Discretionary PTO + paid parental leave
Mental health, wellness & family benefits
A mission-driven team shaping the future of banking

You'll work directly with Bruce Kim (CTO / Co-Founder), who sets the technical bar and goes deep on the hardest problems, and Srinivas Njay (CEO), who is hands-on daily across product and engineering. The team is small enough that your judgment shapes everything, and the market is large enough that what you build will matter for a long time.

Similar Jobs

More Jobs at Interface AI

More Information Technology Jobs

Find similar Staff Platform & Reliability Engineer jobs: