Patreon

Senior Machine Learning Engineer, Infrastructure

Patreon$150K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in building and maintaining production-grade ML infrastructure.
  • Expertise in low-latency live inference pipelines and feature store architectures.
  • Strong knowledge of distributed systems and backend engineering principles.
  • Proficiency in Python for writing robust, maintainable code.
  • Experience with debugging high-throughput systems and performance issues.
  • Excellent communication skills for documentation and cross-team collaboration.

Responsibilities

  • Architect and maintain high-throughput live inference infrastructure for relevance systems.
  • Manage the end-to-end feature store lifecycle, ensuring availability and consistency.
  • Design monitoring and validation frameworks for performance detection.
  • Collaborate with product and engineering teams to develop scalable solutions.
  • Automate model deployment and testing for system reliability.
  • Debug and resolve performance bottlenecks in relevance systems.

Benefits

  • Competitive healthcare coverage.
  • Flexible time off and company holidays.
  • Commuter benefits and lifestyle stipends.
  • Stipends for learning and professional development.
  • Parental leave and 401k plan with matching contributions.
Full Job Description
This role is based in San Francisco or New York as an in-office 3 days per week on a hybrid work model.

About the Team

You'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how content surfaces across Patreon. The team is responsible for search, feed ranking, and creator-fan matching. You'll work closely with a small, collaborative group of MLEs on shared infrastructure, code reviews, and roadmap alignment, while partnering cross-functionally with Product, Data Engineering, and Trust & Safety to deliver measurable impact across the platform.

About the Role
  • Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support our relevance systems.
  • Own the end-to-end feature store lifecycle-from ingestion and transformation to production serving, ensuring high availability and consistency between online and offline features.
  • Design and implement observability, monitoring, and validation frameworks to detect performance gaps, latency spikes, and production drift.
  • Collaborate with cross-functional partners, such as product, data engineering, and trust and safety, to translate product requirements into robust, scalable infrastructure solutions.
  • Automate model deployment and reliability testing to improve developer velocity and ensure system stability.
  • Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability issues.

About You
  • You have deep experience building, deploying, and maintaining production-grade ML infrastructure at scale, specifically with low-latency live inference pipelines and feature store architectures.
  • You have a strong background in distributed systems and backend engineering, with the ability to write robust, maintainable code in Python.
  • You have a systematic approach to debugging complex, high-throughput systems and performance bottlenecks.
  • You are energized by building '0 to 1' infrastructure systems that stand the test of time and provide a reliable foundation for the team.
  • You possess strong communication skills and are effective at creating clear documentation for system architectures and infrastructure strategies.
  • You have a growth mindset, a keen eye for detail in code reviews, and a passion for empowering your teammates by improving developer velocity.


Patreon offers a competitive benefits package including and not limited to salary, equity plans, healthcare, flexible time off, company holidays and recharge days, commuter benefits, lifestyle stipends, learning and development stipends, patronage, parental leave, and 401k plan with matching.

Patreon operates under a hybrid work model, where employees based in office locations are expected to come into the office two days per week, excluding sick time and paid leave. The goal of this policy is to be intentional about the in-person time we spend together to strengthen the feeling of community at Patreon. Candidates hired into remote-eligible roles are not expected to meet the same requirements.

At Patreon, we believe in fair and transparent pay. In compliance with New York and California pay transparency laws, we are sharing the expected salary range for this role.

The posted salary range is dependent on the location and the level. This range may encompass multiple levels within the role's job family. The final offer will be based on candidate's experience, skills, competencies, and geographic location, aligning with the appropriate job level within Patreon's leveling framework. For remote employees located outside CA and NY, salary may vary based on location and local market conditions.

Patreon reserves the right to modify or update compensation and benefits at any time

About Patreon

Patreon is a membership platform that provides business tools for creators to run a subscription content service. The company was founded in 2013 by musician Jack Conte and developer Sam Yam. Patreon allows creators to earn a monthly income by providing exclusive content and experiences to their subscribers. The platform has over 200,000 active creators and over 6 million active patrons.
Learn more about Patreon
Size
200 employees
Industry
Founded
2013

Similar Jobs

More Jobs at Patreon

More Enterprise Technology Jobs

Find similar Senior Machine Learning Engineer, Infrastructure jobs: