Staff Cloud/ML Ops Engineer

Ivo

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Minimum 7 years of experience in infrastructure engineering
  • Deep, hands-on experience with Kubernetes in production environments
  • Strong experience with infrastructure as code tools like Pulumi or Terraform
  • Solid understanding of distributed systems architecture, including scheduling and networking
  • Experience managing multi-cluster or multi-region setups with GitHub CI/CD

Responsibilities

  • Own and evolve Kubernetes platform across AWS, GCP, and Azure
  • Design and operate multi-cluster architectures with robust failover strategies
  • Build internal tooling for cluster provisioning and standardize environments
  • Develop strategies to isolate ML workloads from API workloads
  • Implement security controls at the platform layer ensuring compliance and data isolation
  • Collaborate with SRE and ML teams to establish realistic SLOs

Benefits

  • Competitive compensation based on experience and expertise
  • Equity options for meaningful ownership in a fast-scaling company
  • Relocation assistance and visa support for successful applicants
  • Comprehensive health, dental, and vision insurance
  • Access to HSA and FSA accounts, plus life insurance coverage
  • 401(k) savings plan
  • Commuter benefits for convenience
  • Unlimited PTO for work-life balance
  • Vibrant office perks including catered lunches, snacks, a gym, and a dog-friendly environment
Full Job Description
The Role: Why, What and the Who

Why? Infrastructure Engineers build the foundation for Ivo's entire platform. Customers are cagey about their contracts, so each customer gets their isolated environment with containers, database, VPC, etc. Things break. Regions go down. Cloud and LLM providers have "incidents." Customers still expect us to hit our SLAs.

What? We're looking for a Cloud/MLOps Engineer as part of Infrastructure team to:
  • Own and evolve our Kubernetes platform across AWS/GCP/Azure
  • Design and operate multi-cluster / multi-region architectures with failover and disaster recovery strategies adhering to secure cluster isolation boundaries
  • Build internal tooling for cluster provisioning and lifecycle management, standardizing environments (dev 12 staging 12 prod)
  • Design strategies to isolate ML vs API workloads while optimizing for cost, performance, and reliability
  • Implement security and compliance controls at the platform layer with RBAC, workload identity, secrets management while preserving data isolation aligned with residency requirements and auditability for enterprise customers
  • Partner with SRE + ML teams to ensure SLOs are realistic and enforceable and models are deployed in production environments

Who? We need someone who has:
  • Minimum 7 years of experience
  • Deep, hands-on experience with Kubernetes in production (you've debugged it at 2am, not just deployed to it)
  • Strong experience with infrastructure as code (Pulumi, Terraform, etc.)
  • Strong understanding of cluster architecture, scheduling, networking, storage primitives and failure modes in distributed systems
  • Experience managing multi-cluster or multi-region setups with Github CI/CD

This isn't a "keep the lights on" role. You'll be building the system that keeps the company running. In addition to helping us run a solid, high-performance distributed system, we'd love someone who's as excited about LLMs as we are. You'd be deeply embedded into the engineering team and highly encouraged to push the frontiers.

Ivo might be a good fit for you if you:
  • You love writing code, but you love having impact more: We're a team of engineers at heart, but our #1 goal is building the best possible product. That means making pragmatic choices and looking for 80/20 solutions.
  • Would describe yourself as being relentlessly resourceful.
  • You have a strong internal sense of urgency. You have a bias towards doing things *today*, rather than tomorrow.
  • Experience working in a startup environment is preferred but not required.
  • Are excited about the adventure of building a company!


What We Offer
  • Competitive Compensation: Final offer details are determined based on experience, expertise, and overall fit.
  • Equity: Meaningful ownership in a company that's scaling fast
  • Relocation and Visa Support: We also offer relocation assistance for successful applicants moving to SF, as well as support for visa and green card applications where applicable.
  • Health & Wellness: Comprehensive medical, dental, and vision plans to suit the needs of you and your family.
  • Flexible Spending & Insurance: Access to HSA and FSA accounts, plus life insurance coverage.
  • 401(k) Program: Save for the future with our 401(k) program.
  • Commuter Benefits: We help make getting to and from the office easier and more convenient.
  • Unlimited PTO: So you can take the time you need to recharge, stay healthy, and bring your best self to work.
  • Office Perks: Enjoy a vibrant Downtown San Francisco office with catered lunch five days a week, premium snacks and coffee, an in-building gym, and a dog-friendly environment.


FAQ
  • What stage of growth is Ivo at?: We launched in early access in 2023. Since then, we've had an incredible response from the market and are growing rapidly. We 6x'd in ARR in the last 12 months. Our clients include companies like Uber, Reddit, IBM, Canva, Pinterest, WordPress, and more. We're happy to share more details with candidates who go through our interview process.
  • Is this a chill gig?: Startups are very hard, especially if they're growing fast. You'll have a ton of responsibility, and there's always an enormous amount of stuff to do. It's hard work but the payoff is uncapped.
  • Can I work remotely?: We require candidates to work with us in-person 5 days a week in our San Francisco office.

Similar Jobs

More Jobs at Ivo

  • Staff Cloud/ML Ops Engineer
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Senior AI Researcher
    $160K — $190K *
    San Francisco, CA 94112 (San Francisco County)
    Legal & Accounting
    In-Person
  • Senior DevOps Engineer
    $135K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Staff DevOps Engineer
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Designer
    $142K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Business Services
    In-Person

More Information Technology Jobs

Find similar Staff Cloud/ML Ops Engineer jobs: