Wealthsimple

Senior Software Developer, ML Platform & Infrastructure

Wealthsimple • $110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of software engineering experience in ML Infrastructure, MLOps, ML Tooling, or Data Platform engineering.
  • Deep experience in MLOps/ML Platform practices with a proven track record in self-serve ML platforms and production serving infrastructure.
  • Advanced proficiency in Python, Kubernetes, Terraform, and AWS cloud services.
  • Strong desire to specialize in LLM serving and tackle LLM-specific challenges.
  • Experience designing highly available, observable microservices for real-time requests.
  • Proven capability to lead architectural roadmaps and maintain complex platform systems.

Responsibilities

  • Transition traditional ML lifecycle and serving patterns into LLM inference engines and GPU orchestration systems.
  • Design low-latency routing frameworks to direct requests across cloud providers and self-hosted models.
  • Architect and manage high-performance GPU serving environments on Kubernetes.
  • Build automated evaluation and observability frameworks for model quality and performance validation.
  • Partner with product engineering and data science teams to create framework-agnostic platform tooling.
  • Optimize cost and performance across self-hosted and managed inference systems.

Benefits

  • Top-tier health benefits and life insurance.
  • Long-term group savings with employer match.
  • 20 vacation days, 4 wellness days, and unlimited sick and mental health days per year.
  • Work outside Canada for up to 90 days per year.
  • Inclusive employee resource groups for diverse communities.
Full Job Description
ML Platform & Infrastructure Team

The Machine Learning Infrastructure & Platform team builds the foundational architecture powering AI and GenAI initiatives across Wealthsimple. We sit at the intersection of production MLOps and cutting-edge GenAI enablement.

As our AI footprint expands rapidly, our priority is evolving our robust MLOps foundations into a scalable, high-performance LLM serving and routing platform. We build self-serve systems that allow Data Scientists and Engineers to host open-source LLMs reliably, optimize inference latencies, manage GPU infrastructure, and benchmark model performance safely in production.

The Role

We are looking for an experienced MLOps or ML Platform Engineer who is excited to pivot their deep background in model orchestration, serving, and platform tooling toward solving the unique challenges of LLM inference and infrastructure.

In this role, you will bridge the gap between traditional MLOps (model lifecycles, pipeline orchestration, serving infrastructure) and modern GenAI stack requirements (vLLM, GPU cluster management, intelligent model routing, and automated Evals). You will take end-to-end ownership of setting the technical direction for operating enterprise-grade LLM systems company-wide.

In this role, you will have the opportunity to:
  • Turn MLOps expertise to GenAI: Transition traditional ML lifecycle and serving patterns into state-of-the-art LLM inference engines and GPU orchestration systems.
  • Build model-routing architecture: Design low-latency routing frameworks (e.g., LiteLLM integration) to dynamically direct requests across managed cloud providers (AWS Bedrock) and self-hosted open-source models.
  • Provision & scale GPU infrastructure: Architect and manage high-performance GPU serving environments on Kubernetes using engines like vLLM, Ray, and Triton.
  • Develop evaluation & benchmarking tooling: Build automated Evals and observability frameworks to empower engineers and data scientists to validate model quality, latency, and drift against production requirements.
  • Empower self-serve ML across Wealthsimple: Partner with product engineering and data science teams to build framework-agnostic platform tooling that abstracts infrastructure complexity.
  • Drive cost & performance optimization: Improve price-performance across self-hosted and managed inference by optimizing capacity, utilization, batching, routing, and model selection while meeting quality and reliability objectives


We are looking for people who have:
  • 7+ years of software engineering experience in ML Infrastructure, MLOps, ML Tooling, or Data Platform engineering.
  • Deep experience in MLOps/ML Platform practices: Proven track record building and operating self-serve ML platforms, model registry workflows, experiment tracking, or production serving infrastructure (Kubeflow, MLflow, Ray, Triton, SageMaker).
  • Strong platform fundamentals: Advanced proficiency in Python, container orchestration via Kubernetes, infrastructure-as-code (Terraform), and cloud provider ecosystem (AWS).
  • Strong appetite to specialize in LLM serving: A genuine desire to leverage your existing MLOps skillset to tackle LLM-specific challenges (vLLM, model routing, prompt engineering tooling, vector databases, GPU memory optimization, or LLM evaluation frameworks).
  • Backend performance & observability focus: Experience designing highly available, observable microservices (e.g., FastAPI) handling real-time, low-latency requests.
  • End-to-end technical ownership: Proven capability to lead architectural roadmaps, guide multi-functional projects with high autonomy, and maintain complex platform systems for the long run.


Nice-to-haves (or areas you will learn on the job):
  • Direct experience serving open-source Large Language Models in production (vLLM, SGLang, TensorRT-LLM, Dynamo).
  • Hands-on work with CUDA, GPU partitioning, or distributed inference frameworks (Ray Serve).
  • Familiarity with vector search and retrieval engines (Elasticsearch, Qdrant, Pinecone).


Our Stack Includes:
  • Container & Infrastructure: Kubernetes, Terraform, AWS GPU Infrastructure
  • ML Serving & LLM Tooling: vLLM, Ray, Triton, LiteLLM, AWS Bedrock, MLflow, SageMaker
  • Languages & Frameworks: Python, FastAPI, PyTorch
  • Data & Streaming: Kafka, Postgres, Redshift, Snowflake


Why Wealthsimple?

🌸 Top-tier health benefits and life insurance

Long-term group savings with employer match, through Wealthsimple for Business

20 vacation days, 4 wellness days, and unlimited sick and mental health days per year*

90 days away: work outside Canada for up to 90 days per year*

Employee resource groups, including Rainbow (2SLGBTQ), Women of WS, and Black at WS

We are a hybrid team with over 1,500 employees across North America. The people are one of the best parts of working here: you'll collaborate with incredibly talented, curious, and driven teammates who are deeply committed to doing great work.

*Unlimited paid sick days, Wellness Days and the 90 day away program do not apply to certain roles.

About Wealthsimple

Wealthsimple is a financial services company that provides online investment management and trading services. The company's platform allows users to invest in a variety of financial products, including stocks, bonds, and exchange-traded funds (ETFs), and offers a range of tools and resources to help users manage their investments. Wealthsimple also offers a high-interest savings account and a tax preparation service. The company was founded in 2014 and is headquartered in Toronto, Canada.
Learn more about Wealthsimple
Size
500 employees
Industry
Founded
2014

Similar Jobs

More Jobs at Wealthsimple

More Information Technology Jobs

Find similar Senior Software Developer, ML Platform & Infrastructure jobs: