GoFundMe

Manager, Machine Learning Engineering

GoFundMe$219K — $329K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of experience in building production machine learning systems.
  • 1-3 years of experience managing engineers in an MLOps or ML infrastructure role.
  • Proficiency in Python and ML libraries like PyTorch and TensorFlow.
  • Experience with real-time model serving and scalable inference.
  • Strong data engineering skills with SQL and Spark/Databricks.
  • Experience with ML monitoring for technical and business metrics.
  • Familiarity with generative AI infrastructure is a plus.

Responsibilities

  • Own the scalability and operational health of ML/AI production systems.
  • Lead and develop a team of ML/AI operations engineers.
  • Streamline deployment processes from model development to production.
  • Drive standards for model observability and incident response.
  • Establish on-call processes and operational practices for ML systems.
  • Develop operational strategy for traditional and generative AI systems.
  • Collaborate with cross-functional teams to align business goals with technical priorities.

Benefits

  • Healthcare, dental, vision, and life insurance.
  • 401(k) saving program with company contributions.
  • Equity options available for employees.
Full Job Description
Join GoFundMe as our next Manager, Machine Learning Engineering (ML and AI Operations). In this role, you will lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe's machine learning and AI systems reliable, scalable, and safe in production. This role requires strong technical judgment across the ML lifecycle (data 12 training 12 online inference 12 monitoring), a strong understanding of how to enable AI applications to operate safely at scale, and a proven ability to build and lead a high performance team that operates production ML/AI systems with the same rigor as core infrastructure. Candidates considered for this role will be located in the San Francisco Bay Area. There will be an in-office requirement of 3x a week. The Job 12 Own the reliability, scalability, and operational health of ML/AI production systems across GoFundMe, including training pipelines, feature stores, model serving, and monitoring/observability infrastructure. 12 Lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and shared best practices for production ML and AI systems. 12 Partner with data science and ML engineering teams to streamline the path from model development to production deployment, including CI/CD for ML, model packaging, versioning, and rollback strategies. 12 Establish ML operational excellence org-wide by driving standards for model observability (latency, errors, drift, calibration, business KPI deltas), automated retraining triggers, and incident response playbooks. 12 Build and mature on-call processes, SLOs/SLAs, and postmortem practices for ML/AI systems, treating model incidents with the same discipline as production infrastructure incidents. 12 Drive operational strategy for GoFundMe's generative AI systems alongside traditional ML, balancing innovation velocity with safety, compliance, cost, and reliability. 12 Collaborate cross-functionally with Product, Engineering, Design, and Legal/Privacy stakeholders to translate business goals into team priorities and measurable operational outcomes. 12 Manage vendor and platform relationships (e.g., cloud ML platforms, LLM providers) and make build-vs-buy calls that balance cost, control, and speed. 12 Report on team health, system reliability metrics, and operational risk to senior engineering leadership. 12 Employ a diverse set of tools and platforms, including Python, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, and GitHub, to guide your team in developing, deploying, and maintaining scalable and robust machine learning systems. You 12 7+ years of hands-on experience building and shipping production machine learning systems, with demonstrated ownership of backend services and ML pipelines in a high-availability environment. 12 1-3+ years of experience directly managing engineers, ideally in an MLOps, ML platform, or infrastructure context, with a track record of hiring and developing strong teams. 12 Strong proficiency in Python and ML libraries/frameworks such as PyTorch, TensorFlow, Scikit-learn, plus strong software engineering fundamentals (testing, code review, CI/CD, API design, performance, and reliability) - enough depth to stay hands-on and credible with your team. 12 Experience designing and operating real-time model serving at scale, including containerization, scalable inference, feature retrieval, and safe rollout strategies (canaries, shadowing, backward-compatible schema evolution). 12 Strong data engineering fluency: building reliable datasets and features using SQL, Spark/Databricks, and warehouse technologies (e.g., Snowflake), with an understanding of event semantics, identity resolution, and data quality controls. 12 Proven experience implementing ML monitoring for both technical and business metrics (drift, calibration, segment performance, latency, error budgets) and running models reliably in production. 12 Familiarity with generative AI/LLM infrastructure and operational considerations (latency, cost, safety guardrails) is a strong plus. 12 Ability to break down ambiguous, high-impact problems, define crisp interfaces and success metrics, and deliver iteratively while managing stakeholder expectations across engineering leadership, product, and data science. 12 Strong leadership and mentoring skills and a proven ability to raise the bar on architecture, engineering quality, and operational rigor for production ML/AI systems. 12 Advanced degree (Master's or Ph.D.) in Computer Science, Statistics, Data Science, or a related technical field is preferred. 12 Sense of humor is optional but appreciated. The annual U.S. salary range for this full-time position is $219,000 - $329,000. The company also offers equity and other benefits to employees, including healthcare, dental, vision, life insurance and 401(k) saving program. In addition to this wage, there are geolocation differentials that will increase pay depending on the work location. Additionally pay may vary depending on other factors including skills, experience, education, or training. Your recruiter can share more about the specific total compensation package based on your location during the hiring process.

About GoFundMe

GoFundMe is a crowdfunding platform that allows people to raise money for various causes and projects. The company was founded in 2010 and is headquartered in Redwood City, California. GoFundMe has helped people raise over $9 billion for various causes, including medical expenses, education, and disaster relief. The platform is easy to use and allows users to create and share their campaigns on social media. GoFundMe charges a 2.9% processing fee for each donation, plus a $0.30 transaction fee. The company has been recognized for its impact on charitable giving and has won numerous awards for its platform and services.
Learn more about GoFundMe
Size
200 employees
Industry
Founded
2008

Similar Jobs

More Jobs at GoFundMe

More Enterprise Technology Jobs

Find similar Manager, Machine Learning Engineering jobs: