Western Governors University

Principal AI Platform Operations Engineer

Western Governors University$202K — $313K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years experience in deploying and operating ML/AI systems at scale.
  • 5 years working with Data Scientists or ML Engineers.
  • 3 years building large-scale ML or deep learning models on cloud platforms.
  • Master's degree in a related field (Computer Science, Data Science, etc.).
  • Expertise in AI safety, responsible AI, and enterprise governance.

Responsibilities

  • Define the multi-year technical vision and strategy for AI operations.
  • Establish foundational architectures and engineering principles for AI systems.
  • Lead adoption of advanced AI operations tools and practices.
  • Drive complex AI operations programs and large-scale initiatives.
  • Advise executives on AI strategy, risk, and investment priorities.

Benefits

  • Medical, dental, vision, and telehealth services.
  • Flexible paid time off and sick time.
  • Retirement savings plan with matching contributions.
  • Discounted tuition for further education at WGU.
  • Comprehensive life and disability insurance options.
Full Job Description
The salary range for this position takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organizational needs.

At WGU, it is not typical for an individual to be hired at or near the top of the range for their position, and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is:

Grade: Technical 413

Pay Range: $202,000.00 - $313,000.00

Job Description

The Principal AI Platform Operations Engineer is the highest technical individual contributor and authority within the AI Platform & Operations team. The Principal Engineer is responsible for setting the long-term AI operations and platform vision and strategy for WGU. This role operates at the intersection of frontier AI operations practices and enterprise-scale production systems, translating emerging developments into strategic technical investments that advance the organization's mission. The Principal AI Platform Operations Engineer shapes the AI operations discipline itself, defining the practices, standards, and culture that govern how the organization deploys and operates AI, while advising executive leadership, influencing the product roadmap, and representing WGU's AI platform capabilities externally.

Primary Responsibilities
  • Defines the multi-year AI operations and platform technical vision and strategy for WGU, aligned to organizational mission and emerging technological opportunity.
  • Establishes the foundational reference architectures and engineering principles that govern how AI/ML systems are deployed, operated, and governed across the organization.
  • Leads the evaluation and adoption of frontier AI operations paradigms, tooling, and infrastructure; authors technical strategy documents and architectural decision records that guide org-wide direction.
  • Drives the most complex, highest-stakes AI operations programs in the organization, including large-scale platform, migration, and reliability initiatives.
  • Advises executives such as the Chief Technology Officer, VP of Engineering or Security, and senior product leadership on AI operations strategy, risk, and investment priorities.
  • Defines AI operations governance including responsible AI standards, model risk management, security, cost accountability, and compliance frameworks for the organization.
  • Identifies and incubates emerging operational capabilities; sponsors proof-of-concept initiatives that create future organizational leverage.
  • Serves as the primary external technical spokesperson for WGU's AI operations work; represents the organization at industry conferences, in partnerships, and with strategic vendors.
  • Designs and evolves the team's operating model, hiring criteria, competency framework, and engineering culture.
  • Develops Staff, Senior, and II-level engineers through technical mentorship, sponsorship, and organizational knowledge-building programs.
  • Authors internal and external thought leadership: technical blogs, whitepapers, architectural guides, and research contributions.
  • Partners with legal, compliance, and privacy teams to ensure AI systems meet regulatory requirements and institutional risk standards.
  • Performs other job-related duties as assigned.


This job description includes a general representation of job requirements rather than a comprehensive inventory of all required responsibilities or work activities. The contents of this document or related job requirements may change at any time with or without notice.

Qualifications
Knowledge, Skills, and Abilities
  • Recognized expertise across the full spectrum of AI operations: deployment, platform engineering, observability, reliability engineering, governance, and LLMOps/MLOps at enterprise scale.
  • Ability to define and communicate multi-year technical strategy; translate organizational goals into architectural vision and operational roadmaps.
  • Deep expertise in frontier AI operations practices and the ability to assess, adapt, and productionize emerging techniques at organizational scale.
  • Expert-level knowledge of AI safety, responsible AI, model risk, security, and enterprise AI governance including regulatory and compliance considerations.
  • Proven ability to influence organizational direction at the executive level; able to build consensus across competing priorities and stakeholders.
  • Strong platform and systems thinking: ability to design AI infrastructure that enables teams to move faster and build more reliably.
  • Track record of establishing engineering culture, standards, and practices that durably improve organizational capability.
  • Exceptional written and verbal communication; able to author technical strategy documents, present to boards and executives, and represent the organization externally.
  • Experience in technical hiring, team-building, and developing talent across multiple seniority levels.
  • Deep familiarity with the AI operations vendor landscape, open-source ecosystem, and technology frontier; able to make informed, timely, well-reasoned bets.
  • Strong cross-functional leadership; comfortable driving alignment across engineering, product, data science, legal, and executive stakeholders.

Education
  • Master's Degree in Computer Science, Software Engineering, Data Science, Machine Learning, Math, Physics, or a related field

Experience
  • 7+ years of hands-on experience deploying and operating ML/AI systems in production at scale.
  • 5 years of experience working in an AI/ML context alongside Data Scientists or ML Engineers
  • 3 years of experience in building large-scale machine learning or deep learning models on a cloud platform
  • Demonstrated experience setting operational strategy or architectural vision at an organizational level.
  • Proven track record of delivering transformative AI platform or operations programs that produced measurable business impact.
  • Experience advising or influencing senior/executive leadership on AI operations strategy and investment.
  • Experience leading or significantly contributing to AI governance, responsible AI, or model risk frameworks.
  • Experience building and scaling AI engineering or operations teams, including hiring, competency-building, and culture-setting.
  • Experience representing an organization externally through conference talks, publications, or strategic partnerships.
  • Deep hands-on expertise in production AI reliability engineering, observability, and deployment at scale.

Preferred Qualifications
  • 10+ years of experience in software engineering, data science, or machine learning.
  • Experience with the Databricks platform; Databricks certifications (e.g., Databricks Certified Machine Learning Professional, Databricks Certified Data Engineer, or Databricks Certified Associate Developer for Apache Spark) are a plus.
  • AWS cloud platform experience; AWS certifications (e.g., AWS Certified Machine Learning, Specialty, AWS Certified DevOps Engineer) a plus.
  • Experience with infrastructure-as-code (Terraform, CloudFormation) and container orchestration (Kubernetes) for AI/ML workloads.
  • Experience in EdTech, personalized learning, or student-facing AI/ML platforms.
  • Experience with enterprise AI governance and compliance frameworks (FERPA, GDPR, etc.).
  • Contributions to open-source MLOps/LLMOps tooling, or published work and conference presentations in the AI/ML operations domain.
  • PhD in Computer Science, AI/ML, or a related field.
  • Experience operating in highly regulated environments (e.g., EdTech, healthcare, finance) with associated compliance requirements.
  • Recognized thought leadership in the AI/MLOps community through writing, speaking, or advisory roles.


This position requires occasional travel of up to 20%, including required attendance at designated company summits (typically one to two per year). Additional travel may include conferences, visits to company locations, and other business-related events as needed. Additional travel may be assigned as needed to support business requirements.

#LI-AW2

Position & Application Details

Full-Time Regular Positions (classified as regular and working 40 standard weekly hours): This is a full-time, regular position (classified for 40 standard weekly hours) that is eligible for bonuses; medical, dental, vision, telehealth and mental healthcare; health savings account and flexible spending account; basic and voluntary life insurance; disability coverage; accident, critical illness and hospital indemnity supplemental coverages; legal and identity theft coverage; retirement savings plan; wellbeing program; discounted WGU tuition; and flexible paid time off for rest and relaxation with no need for accrual, flexible paid sick time with no need for accrual, 11 paid holidays, and other paid leaves, including up to 12 weeks of parental leave.

How to Apply: If interested, an application will need to be submitted online. Internal WGU employees will need to apply through the internal job board in Workday.

Additional Information

Disclaimer: The job posting highlights the most critical responsibilities and requirements of the job. It's not all-inclusive.

About Western Governors University

Western Governors University (WGU) is a private, nonprofit online university based in Salt Lake City, Utah. The university was founded by 19 U.S. governors in 1997 with a mission to expand access to higher education. WGU offers undergraduate and graduate degree programs in business, information technology, education, and healthcare. The university is accredited by the Northwest Commission on Colleges and Universities and has been recognized by the White House as an example of excellence in education innovation. WGU has a competency-based learning model, which allows students to progress through their coursework at their own pace based on their mastery of the material.
Learn more about Western Governors University
Size
5,000 employees
Industry

Similar Jobs

More Jobs at Western Governors University

More Technical Services Jobs

Find similar Principal AI Platform Operations Engineer jobs: