AI Platform Operations Manager

STACK Infrastructure

$128K — $146K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or equivalent experience.
  • 5+ years in DevOps, site reliability, or platform engineering with strong delivery track record.
  • Proficiency in Infrastructure as Code using Terraform, Bicep, or ARM.
  • Experience with CI/CD pipelines in Azure DevOps or GitHub Actions.
  • Hands-on knowledge of Docker, Azure Kubernetes Service (AKS), and Azure Container Apps.
  • Solid understanding of Azure services - compute, networking, storage, identity, Key Vault.
  • Scripting skills in Python, PowerShell, or Bash.

Responsibilities

  • Build and maintain infrastructure-as-code modules for Azure using Terraform, Bicep, or ARM.
  • Automate AI platform component provisioning in Azure as reusable patterns.
  • Ensure environment consistency across development, test, and production environments.
  • Automate routine platform operations like patching, backup validation, and disaster recovery.
  • Design and operate CI/CD pipelines for application code and infrastructure.
  • Implement deployment automation for AI services and integration workloads on Azure.
  • Monitor platform performance and provide incident response and triage.

Benefits

  • Comprehensive Healthcare plan.
  • Dental and Vision Insurance.
  • Life Insurance coverage.
  • Paid Time Off for vacation and personal days.
  • Paid Leave Programs for various life events.
Full Job Description
THE POSITION:

The DevOps Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK's enterprise AI platform on Azure. This is a hands-on engineering role focused on build and run - not oversight.

Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure-as-code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably. The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps - turning platform architecture into automated, governed, observable, and cost-efficient environments that teams across the organization build on.

KEY RESPONSIBILITIES

Infrastructure Automation & Infrastructure as Code
  • Build, maintain, and version infrastructure-as-code modules for Azure environments using Terraform, Bicep, or ARM, including compute, networking, storage, identity, and AI platform resources.
  • Automate provisioning of AI platform components - Azure AI Foundry, Azure OpenAI Service, Azure AI Search, Cosmos DB, ADLS Gen2, and Databricks - as reusable, parameterized deployment patterns.
  • Maintain environment parity across development, test, and production, including configuration management, drift detection, and remediation.
  • Implement and enforce tagging, naming, and resource organization standards that support governance, chargeback, and lifecycle management.
  • Automate routine platform operations - patching, certificate rotation, key and secret rotation, backup validation, and disaster recovery testing.

CI/CD & Release Engineering
  • Design, build, and operate CI/CD pipelines in Azure DevOps or GitHub Actions for application code, infrastructure code, container images, and AI/agent deployments.
  • Implement automated build, test, security scanning, artifact management, and promotion gates across environments.
  • Establish branching strategies, code review standards, and release management practices in partnership with AI engineering and enterprise application teams.
  • Build deployment automation for agentic AI services, MCP (Model Context Protocol) servers, and integration workloads running on Azure Container Apps and Azure Kubernetes Service (AKS).
  • Support model and prompt release workflows - versioning, staged rollout, evaluation gates, and rollback procedures for LLM-based applications.

Container Platform & AI Workload Operations
  • Operate and tune AKS and Azure Container Apps, including cluster upgrades, node pool sizing, autoscaling, ingress, networking, and workload isolation.
  • Build and maintain container images, base image standards, and registry governance in Azure Container Registry.
  • Manage compute scheduling and scaling for AI workloads, including GPU-backed and inference-heavy workloads where required.
  • Implement resiliency patterns - health probes, retries, throttling, quota management, and failover - for AI endpoints and integration services.

Observability, Reliability & Incident Response
  • Instrument platform and AI services with logging, metrics, tracing, and alerting using Azure Monitor, Log Analytics, Application Insights, and equivalent open-source tooling.
  • Build dashboards and service-level indicators covering platform availability, latency, throughput, error rates, token consumption, and model endpoint performance.
  • Participate in on-call rotation, lead incident triage and resolution for platform issues, and drive root cause analysis and corrective actions.
  • Develop and maintain runbooks, operational documentation, and automated remediation for recurring issues.

Security, Governance & Cost Optimization
  • Implement DevSecOps practices - secrets management in Azure Key Vault, managed identity usage, least-privilege access, dependency and container vulnerability scanning, and policy-as-code.
  • Partner with Information Security to ensure pipelines and environments meet enterprise security, data residency, and compliance requirements.
  • Support Azure FinOps practices through cost visibility, rightsizing, reserved capacity, and automated controls on non-production and idle resources.
  • Maintain audit trails and change records for infrastructure and release activity.

Delivery & Cross-Functional Collaboration
  • Work directly with AI engineers, data engineers, and enterprise application teams to remove deployment friction and improve time-to-production for AI solutions.
  • Translate platform architecture and standards into automated, self-service capabilities that teams can consume without deep infrastructure knowledge.
  • Contribute to platform engineering standards, reference implementations, and internal documentation.
  • Provide technical escalation support for build, deployment, and environment issues.


THE DETAILS:
  • Location: Denver, CO
  • Travel:
  • Benefits: Healthcare, Dental Care, Vision Insurance, Life Insurance, Paid Time Off, and Paid Leave Programs
  • Must be eligible to work in the United States
  • Must pass comprehensive background and drug screening


MUST-HAVE QUALIFICATIONS:
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience.
  • 5+ years of hands-on experience in DevOps, site reliability engineering, or platform engineering roles with a strong delivery track record.
  • Strong proficiency in Infrastructure as Code - Terraform, Bicep, or ARM - including module design, state management, and reusable patterns.
  • Proven experience building and operating CI/CD pipelines in Azure DevOps or GitHub Actions.
  • Hands-on experience with containerization and orchestration - Docker, Azure Kubernetes Service (AKS), and Azure Container Apps or equivalent.
  • Solid working knowledge of Azure core services - compute, networking (VNet, NSG, Private Endpoints), storage, identity (Entra ID), and Key Vault.
  • Strong scripting and automation skills in Python, PowerShell, or Bash.
  • Experience with monitoring and observability tooling - Azure Monitor, Log Analytics, Application Insights, Prometheus, or Grafana.
  • Working knowledge of Git-based workflows, code review practices, and artifact/registry management.
  • Demonstrated ability to troubleshoot production issues across infrastructure, network, and application layers.


PREFERRED QUALIFICATIONS:
  • Microsoft Certified: DevOps Engineer Expert (AZ-400), Azure Administrator (AZ-104), or Certified Kubernetes Administrator (CKA).
  • Experience deploying and operating AI/ML workloads - model endpoints, RAG pipelines, vector databases, or agentic services in production.
  • Familiarity with MLOps tooling and practices - Azure Machine Learning, MLflow, Databricks, or equivalent model lifecycle platforms.
  • Experience deploying MCP (Model Context Protocol) servers or similar integration services connecting AI agents to enterprise systems.
  • Exposure to agentic AI frameworks such as Semantic Kernel, LangGraph, or AutoGen from a deployment and operations perspective.
  • Experience with GPU compute provisioning, quota management, and inference cost optimization.
  • Knowledge of FinOps frameworks and Azure cost optimization practices.
  • Experience integrating with enterprise systems such as Microsoft 365, Freshworks ITSM, Workday, NetSuite, or Procore.
  • Experience in data center, hyperscale, or infrastructure-intensive industry environments.


Compensation Range:
$128,260.00 - $146,017.59

THIS MIGHT BE RIGHT FOR YOU IF:

  • You are a strong communicator, you are persuasive and clear, blending analytics with experience in decision-making.
  • You do not get flustered easily. You can juggle multiple priorities while balancing urgent requests with shifting timelines and deliverables.
  • You are a team builder. You take the time to understand and develop the strengths of your resources while formulating long-term plans for the growth and success of the team.
  • You are naturally curious and driven toward continual improvement. While you celebrate your successes, you take time to review and analyze campaigns for future learning.

Similar Jobs

More Jobs at STACK Infrastructure

More Information Technology Jobs

Find similar AI Platform Operations Manager jobs: