AstraZeneca

Senior Cloud Platform Engineer (AWS), AI Infrastructure - Evinova

AstraZeneca • $134K — $176K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years in platform engineering, infrastructure, SRE, or DevOps with a focus on infrastructure as code.
  • Deep AWS cloud engineering experience, particularly in production environments.
  • Strong skills in AWS services like Bedrock, ECS, IAM, and CloudWatch.
  • Proficient in AWS CDK (TypeScript/Python) or Terraform for infrastructure as code.
  • Experience with Docker and Amazon ECS for containerized workloads.
  • Familiarity with CI/CD and GitOps pipelines for cloud workloads.
  • Strong software engineering skills in Python and/or TypeScript.

Responsibilities

  • Design and operate scalable AWS cloud platform capabilities for ML/AI workloads.
  • Create reusable infrastructure and deployment patterns for AI teams.
  • Write and maintain AWS IaC primarily using AWS CDK in TypeScript/Python.
  • Build and manage containerized workloads with Amazon ECS.
  • Develop platform capabilities across compute, networking, and storage.
  • Establish CI/CD workflows for automated deployments.
  • Collaborate with engineers to transition prototypes into production services.

Benefits

  • Competitive Flex Benefits & Retirement Savings Program.
  • 4 weeks of paid vacation and annual Personal Days.
  • Variable Pay Bonus/Short Term Incentive opportunity.
  • Eligibility for equity-based long-term incentive program.
Full Job Description
This role is located in the Greater Toronto Area and follows a hybrid work model. Candidates must reside within commuting distance of the GTA or be willing to relocate for this opportunity.

Introduction to Role

The Machine Learning and Artificial Intelligence Operations team (ML/AI Ops) is the cloud platform engineering team responsible for building and operating the infrastructure that enables our AI Engineers and Data Scientists to deploy Generative AI applications reliably, securely, efficiently and at scale.

As a Senior Cloud Platform Engineer on the ML/AI Ops team, you will design, build and operate the AWS platform that runs our production Generative AI, agentic AI and conversational AI workloads. Most of the code you write will be AWS CDK (TypeScript and/or Python) that creates the infrastructure for other teams to build on.

This is a cloud and platform engineering role rather than an AI application development role. You will partner closely with the engineers and Data Scientists who build agents and models, and provide the infrastructure, deployment patterns, model access, observability and operational capabilities they need to move solutions from experimentation into reliable production environments.

You will work across AWS infrastructure, infrastructure as code (IaC), Amazon Bedrock AgentCore, Amazon ECS, CI/CD, SageMaker Unified Studio, AI gateways, observability, scalability, reliability, security, governance and cost optimization. Your work will establish reusable platform capabilities that allow teams across Evinova to deploy and operate solutions faster and more reliably while meeting the requirements of a highly regulated pharmaceutical environment.

Accountabilities

Cloud Platform Engineering
  • Design, build and operate scalable AWS cloud platform capabilities for production ML/AI and Generative AI workloads.
  • Create reusable infrastructure, tooling and deployment patterns that enable AI Engineers and Data Scientists to independently deploy and operate their applications.
  • Write and maintain AWS IaC primarily AWS CDK in TypeScript and/or Python, including reusable CDK constructs that other teams consume.
  • Build and operate containerized workloads using Amazon ECS and AgentCore.
  • Develop reusable platform capabilities across compute, networking, IAM, secrets management, storage, model access and workload isolation.
  • Build and maintain CI/CD and GitOps workflows that enable safe, automated and repeatable deployments across environments.
  • Partner with engineers and Data Scientists to transition prototypes and research workloads into resilient, production-grade services.
  • Build self-service capabilities and automation that improve developer experience and reduce operational toil.


Reliability, Scalability & Operational Excellence
  • Engineer platform capabilities that improve the availability, scalability, resiliency and performance of production GenAI workloads.
  • Design and implement autoscaling, load balancing, retries, timeouts, fallback, rate limiting and failure-recovery strategies.
  • Establish monitoring, alerting, SLOs, runbooks and production-readiness standards for ML/AI workloads.
  • Troubleshoot complex production issues across AWS infrastructure, container, application and model-provider layers.
  • Automate operational processes and proactively identify opportunities to improve platform reliability and performance.
  • Drive cloud and model cost optimization through capacity management, workload optimization and data-driven analysis.


AI Gateway & Model Access
  • Build and operate a centralized AI gateway that routes requests across model providers.
  • Provide secure, reliable and governed access to foundation models through Amazon Bedrock, OpenAI, Anthropic, Microsoft Foundry, Gemini Enterprise Agent Platform (formerly Vertex AI) and other model platforms.
  • Implement model/provider routing, fallback, authentication, rate limiting, quotas and cost controls.
  • Enable AI teams to evaluate and change model providers without tightly coupling their applications to individual model endpoints.
  • Provide the compute, networking, storage and runtime infrastructure to reliably operate agentic AI, RAG and conversational AI workloads in production.


Observability, Governance & Cost Optimization
  • Build platform-level observability for GenAI workloads, including token consumption, latency, throughput, errors, model/provider performance and cost attribution by team and application.
  • Implement standardized tracing, logging, metrics and alerting capabilities that can be adopted across AI applications.
  • Integrate observability technologies such as Amazon CloudWatch, OpenTelemetry, Datadog and Splunk.
  • Build the telemetry and data pipelines that AI teams use to evaluate and monitor production LLM behavior.
  • Build appropriate security, auditability and governance controls for AI workloads operating within a regulated environment.
  • Support compliance with applicable industry standards and practices, including Good Clinical Practice and Good Machine Learning Practice and other GxP related standards and practices.


Representative Projects
  • Build a library of AWS CDK constructs that gives an AI team a production-ready Amazon ECS service, with IAM, secrets, networking and observability, from a single import.
  • Deploy and operate a centralized AI gateway on Amazon ECS, with multi-provider routing, fallback, quotas and rate limiting.
  • Build multi-account CI/CD with CDK Pipelines or GitOps, with safe, staged deployments across environments.
  • Implement token cost attribution by team and application with SLO and burn-rate alerts.
  • Migrate container workloads from Amazon EKS to Amazon ECS without loss of reliability.


Essential Skills/Experience
  • Minimum of 4 years of hands-on experience in a platform engineering, infrastructure, site reliability engineering (SRE) or DevOps role where infrastructure code was your main output.
  • Deep hands-on AWS cloud engineering experience, including designing, deploying and operating production cloud-native infrastructure.
  • Strong experience with AWS services such as Bedrock, ECS, IAM, VPC, load balancing, S3, CloudWatch, Secrets Manager and related services.
  • Strong hands-on experience with IaC. AWS CDK using Python and/or TypeScript is preferred. Strong Terraform engineers who are willing to work in CDK are welcome.
  • Strong experience with Docker and Amazon ECS, including workload deployment, autoscaling, resource management, health checks, networking, security and production troubleshooting.
  • Experience designing and maintaining CI/CD and/or GitOps pipelines for production cloud workloads.
  • Strong understanding of cloud and platform engineering principles, including high availability, scalability, resiliency, observability, performance, security and cost optimization.
  • Experience with on-call rotations, incident response and root-cause analysis for production systems.
  • Strong software engineering skills in Python and/or TypeScript, with experience building production-quality platform tooling and automation.
  • Experience implementing production observability using technologies such as CloudWatch, OpenTelemetry, Prometheus, Splunk, Datadog, Pydantic Logfire, Langfuse or comparable tools.
  • Demonstrated experience partnering with software, data or ML/AI teams to transition workloads from experimentation into reliable production environments.
  • Proven ability to collaborate effectively across ML/AI engineering, software engineering, security, infrastructure and product teams.
  • Strong written and verbal communication skills, including the ability to clearly document infrastructure architecture, operational processes and platform standards.


Preferred Skills/Experience
  • Experience building internal developer platforms or self-service cloud capabilities used by multiple engineering teams.
  • Experience supporting Generative AI/LLM workloads in production, with an understanding of their unique operational challenges including model availability, latency, token consumption, evaluation and cost.
  • Hands-on experience with Amazon Bedrock or another enterprise foundation-model platform.
  • Experience implementing and/or configuring AI/model gateways, proxies or routing layers, including model routing, fallback, rate limiting and provider abstraction.
  • Experience operating infrastructure supporting agentic AI, multi-agent systems, RAG pipelines or conversational AI.
  • Experience with AgentCore or comparable managed agent-runtime capabilities.
  • Experience supporting Model Context Protocol (MCP) tools, servers or services.
  • Experience with LLM evaluation and observability platforms such as Arize Phoenix, Langfuse, Braintrust, Freeplay or Logfire.
  • Experience supporting production ML/AI infrastructure within pharmaceutical, healthcare, financial services or another highly regulated industry.
  • Understanding of GxP and/or other controls applicable to regulated ML/AI systems.
  • Experience with Kubernetes (Amazon EKS), including migrating container workloads to Amazon ECS.


This Role Is Not a Good Fit If
  • Your main experience is building agents, RAG pipelines, prompts or LLM applications, and you have not owned production infrastructure.
  • You have not written or maintained infrastructure as code for production systems.
  • You prefer research or model development to operating production platforms.


Personal Attributes
  • Platform-minded: You think beyond individual applications and build reusable capabilities that enable multiple engineering teams.
  • Operationally focused: You care about reliability, scalability, observability, security, performance and what happens after a workload reaches production.
  • Hands-on: You are comfortable writing code, building infrastructure, operating container platforms and troubleshooting production systems.
  • Customer-obsessed: You view the AI Engineers and Data Scientists using the platform as your customers and continually look for ways to improve their developer experience.
  • Organized and attentive to detail: You effectively manage multiple initiatives and competing priorities.
  • Collaborative and inclusive: You foster a positive engineering culture where teams share knowledge and solve problems together.
  • Comfortable working autonomously: You like working in an evolving environment where you can help establish new systems, patterns, and best practices.
  • Curious and committed to staying current: You love learning and applying that knowledge to the systems you build. You stay current with the latest AI trends and technologies and apply that knowledge to your work pragmatically.

Annual base salary for this position ranges from 134,855.20 to 176,997.45.
AstraZeneca is committed to providing fair and equitable compensation opportunities to all colleagues. Our compensation policies and practices have been designed to allow colleagues to progress through the salary range over time as they progress in their role. The range provided in this posting represents an offer pay range used in a majority of situations. The base pay offered will vary depending on multiple individualized factors, including the candidate's skills and experience, job-related knowledge, and other specific business and organizational needs. In some cases, offers outside the range may also be considered to address unique circumstances.

In addition, our permanent positions offer an annual Variable Pay Bonus/Short Term Incentive opportunity as well as eligibility to participate in our equity-based long-term incentive program (if applicable to role). Benefits offered for permanent roles include a competitive Flex Benefits & Retirement Savings Program, 4 weeks' paid vacation, and annual Personal Days. Fixed Term Contract/Temporary positions (excluding stud

About AstraZeneca

AstraZeneca is a British-Swedish multinational pharmaceutical company that specializes in the research, development, and manufacturing of prescription drugs. The company was formed in 1999 through the merger of Astra AB and Zeneca Group plc. AstraZeneca's products are used to treat a wide range of medical conditions, including cancer, cardiovascular disease, respiratory disease, and diabetes. The company has operations in over 100 countries and employs more than 76,000 people worldwide. AstraZeneca is committed to developing innovative medicines that improve the health and well-being of people around the world.
Learn more about AstraZeneca
Size
83,100 employees
Market Cap
$211.5 billion
Industry
Net Income
$3.1 billion
Founded
1999
5 Year Trend
+10.2%
Revenue
$26.6 billion
NASDAQ

Similar Jobs

More Jobs at AstraZeneca

More Information Technology Jobs

Find similar Senior Cloud Platform Engineer (AWS), AI Infrastructure - Evinova jobs: