Bank of Montreal

Principal AI Cloud Engineer

Bank of Montreal$105K — $215K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's/Master's/PhD in Computer Science, Engineering, or related field
  • 7+ years experience in large-scale distributed cloud infrastructure
  • 5+ years of hands-on experience with Azure or AWS
  • Proven expertise in AI/ML infrastructure: GPU clusters, Kubernetes, CI/CD
  • Strong knowledge of Infrastructure as Code (IaC) using Terraform or Bicep

Responsibilities

  • Design and build cloud-native AI infrastructure for machine learning workloads
  • Operationalize large language models with strong compliance and observability
  • Develop secure APIs and microservices for AI applications
  • Establish and maintain best practices for AI platform architecture
  • Mentor engineers and disseminate reusable infrastructure components

Benefits

  • Influence the AI technical direction and platform development
  • Work on impactful systems used across various business lines
  • Engage in full stack development from infrastructure to model serving
  • Receive mentorship and support from an invested leadership team
Full Job Description

Application Deadline:

11/30/2026

Address:

100 King Street West

Job Family Group:

Data Analytics & Reporting

The Impact
As a Principal AI & Cloud Engineer, you are a hands-on technical developer who designs, builds, and scales cloud-native AI solutions and products. You help set engineering standards, establish patterns, mentor senior engineers, and partner with multiple teams to deliver resilient, governed, and cost-efficient AI at enterprise scale. You99ll help shape and evolve our AI cloud strategy from model serving and LLMOps to security, observability, and compliance so teams across the bank can innovate safely and rapidly.


You will advance BMO99s Digital First strategy by:

  • Defining reference and production-grade solutions for AI/GenAI on cloud (Azure/AWS preferred; multi-cloud aware).
  • Building reusable, secure, and observable components (APIs, SDKs, microservices, pipelines).
  • Operationalizing LLMs and RAG with strong controls and Responsible AI guardrails.
  • Driving platform roadmaps that enable faster delivery, lower risk, and measurable business outcomes.


What99s In It for You

  • Influence the technical direction of AI and the platform primitives others build on.
  • Ship high-impact systems used across many business lines and products.
  • Work across the full stack: cloud infra, data/feature pipelines, model serving, LLMOps, and DevSecOps.
  • Partner with a leadership team invested in your growth and thought leadership.


ResponsibilitiesInfrastructure & Platform Builder

  • Design, build, and operate cloud-native AI infrastructure for ML/GenAI workloads:
  • Compute: GPU/CPU clusters, autoscaling, spot instance strategies
  • Networking: Azure VNet, Private Link, peering, multi-region HA/DR
  • Storage & Databases: high-performance data lakes (e.g., Azure Data Lake Storage), relational DBs, vector DBs (FAISS, Milvus, Pinecone, pgvector)
  • Security: IAM, Key Vault-backed secrets management, encryption, policy-as-code
  • Implement observability and reliability for AI infra:
  • Metrics (latency, throughput, GPU utilization, cost)
  • Logging/tracing (OpenTelemetry), SLOs/SLIs for infra services
  • Build CI/CD and GitOps pipelines for infrastructure-as-code (Terraform/Bicep) and AI platform components
  • Drive FinOps for AI infra: GPU rightsizing, caching, inference optimization, cost governance

Application & Service Enablement

  • Enable frontend and backend services for AI platforms:
  • Secure APIs, microservices, and event-driven architectures
  • Integration with custom model runtimes (TensorRT-LLM, vLLM, Triton/KServe)
  • Provide infrastructure support for RAG systems: embeddings, chunking, retrieval pipelines
  • Ensure scalable serving infrastructure for LLMs and ML models with caching and token optimization

Strategy & Architecture

  • Define and evolve AI infrastructure reference architecture for cloud (Azure preferred):
  • Container orchestration (Kubernetes), service mesh, ingress
  • Serverless/event-driven patterns for AI pipelines
  • Multi-region, HA/DR, compliance-ready designs
  • Establish standards and best practices for containerization, IaC, and secure networking for AI systems

Security, Risk & Governance

  • Implement defense-in-depth for AI infra:
  • IAM least privilege, private networking, KMS/Key Vault, SBOM, image signing
  • Ensure compliance and Responsible AI controls at infra level:
  • Data residency, encryption, lineage, audit readiness

Delivery & Operations

  • Lead infrastructure discovery and solution design with stakeholders
  • Operate platforms with SRE principles: error budgets, incident response, chaos testing
  • Mentor engineers; create reusable IaC modules, templates, and golden paths

Must-Have Qualifications

  • Bachelor99s/Master99s/PhD in CS, Engineering, or related field
  • 7+ years building large-scale distributed cloud infrastructure
  • 5+ years hands-on with Azure/AWS
  • Proven experience with AI/ML infra: GPU clusters, Kubernetes, CI/CD, observability
  • Strong in IaC (Terraform/Bicep), Kubernetes, networking, security
  • Expertise in cloud-native patterns: containers, service mesh, serverless
  • Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs
  • Programming in Python (infra automation) and one of Go/TypeScript for tooling
  • Understanding of frontend/backend integration for AI services
  • Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs
  • Programming in Python (infra automation) and one of Go/TypeScript for tooling
  • Understanding of frontend/backend integration for AI services

Nice-to-Have

  • GPU optimization (CUDA/NCCL, TensorRT-LLM)
  • Observability tools (Prometheus, Grafana, OpenTelemetry)
  • Event streaming (Kafka/Azure Event Hubs), real-time systems
  • Experience with AI platform products (Azure ML, MLflow, KServe, Hugging Face)

Success Metrics

  • Reliability & Performance: SLOs met for infra services, GPU utilization optimized
  • Security & Compliance: Zero critical findings, auditable infra
  • Cost Efficiency: Reduced GPU/infra spend via FinOps strategies
  • Developer Velocity: Faster provisioning and deployment of AI infra
  • Technical Leadership: Influence on infra standards, mentorship, reusable patterns

Salary:

$105,000.00 - $215,000.00

Pay Type:

Salaried

About Bank of Montreal

The Bank of Montreal is a Canadian multinational investment bank and financial services company. It provides a wide range of personal and commercial banking, wealth management, and investment banking products and services. The bank had revenues of CAD 23.6 billion in 2020.
Learn more about Bank of Montreal
Size
45,454 employees
Market Cap
$60.9 billion
Industry
Founded
1817
5 Year Trend
+9.1%
NASDAQ

Similar Jobs

More Jobs at Bank of Montreal

  • Bank of Montreal
    Wealth Advisor
    $90K — $180K *
    Wausau, WI 54401 (Marathon County)
    Finance & Insurance
    In-Person
  • Bank of Montreal
    Lead Developer
    $70K — $150K *
    Toronto, ON M3C 0E3
    Information Technology
    In-Person
  • Bank of Montreal
    Options Product Manager
    $120K — $250K *
    New York, NY 10025 (New York County)
    Finance & Insurance
    In-Person
  • Bank of Montreal
    Strategic Sourcing Manager
    $70K — $150K *
    Toronto, ON M3C 0E3
    Business Services
    In-Person
  • Bank of Montreal
    Financial Advisor
    $65K — $140K *
    Chanhassen, MN 55317 (Carver County)
    Finance & Insurance
    In-Person

More Information Technology Jobs

Find similar Principal AI Cloud Engineer jobs: