Lead AI Infrastructure Operations Engineer

The Mutual Group

$130K — $150K *
US-AnywhereRemote in Iowa, LA
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science, engineering, information technology, or a related field, or equivalent practical experience.
  • 8+ years of IT experience, with 5+ years in infrastructure operations or related fields.
  • Experience with cloud-hosted, data-intensive, or AI-enabled production applications.
  • Experience implementing observability, monitoring, and operational reporting.
  • Knowledge of cloud infrastructure, networking, APIs, and security.
  • Experience coordinating infrastructure services and changes across teams and partners.
  • Familiarity with AWS services related to infrastructure operations.
  • Experience with automation, CI/CD pipelines, and infrastructure-as-code practices.

Responsibilities

  • Partner with cross-functional teams to understand the infrastructure needs for AI use cases.
  • Translate AI solution designs into technical requirements for cloud environments and services.
  • Coordinate infrastructure provisioning and changes with cloud service providers.
  • Ensure operational readiness for AI solutions transitioning to production.
  • Establish reusable infrastructure patterns and operational standards across initiatives.
  • Monitor and improve observability and operational performance for AI applications.
  • Support AI-related incident management and operational reviews.

Benefits

  • Competitive base salary plus incentive plans for eligible team members.
  • 401(K) retirement plan with a company match of up to 6%.
  • Free basic life and disability insurance.
  • Healthcare plans including medical, dental, and vision coverage.
  • Generous time off program including personal and volunteer paid time off.
  • Flexible work schedules and hybrid/remote options for eligible positions.
  • Educational assistance opportunities.
Full Job Description
Department:
Information Technology

Job Description:

The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions.

As TMG expands its AI capabilities, AI CoE will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their infrastructure, environment, connectivity, observability, security, capacity, and production-readiness needs.

The engineer in collaboration with other Infrastructure resources will focus on the AI application, platform, and operational layers. TMG's managed services provider operates the underlying AWS cloud infrastructure. This role will translate AI solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose issues, and validate that environments meet reliability, security, performance, and operational expectations.

This is a hands-on role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed-services partners.

Work Arrangement:
  • Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in-office schedule. The initial training period may require additional in-office days.


Accountabilities:

Enable AI Infrastructure and Environments
  • Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases.
  • Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity, and supporting cloud services.
  • In Collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access, and changes with TMG's AWS managed-services provider and other technology partners.
  • Support operational readiness for AI solutions transitioning into production.
  • Identify infrastructure dependencies, constraints, risks, costs, and lead times early in the delivery lifecycle.
  • Establish reusable infrastructure patterns, operational standards, dashboards, runbooks, and production-readiness requirements across AI use cases.
  • Validate that environments and supporting services are appropriately configured, monitored, secured, scalable, and ready for production use.
  • Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails.
  • Follow and enforce established security, compliance, and operational controls across AI platforms and infrastructure.

Operate and Improve Production AI Solutions
  • Work with managed services to Implement and maintain observability, logging, tracing, monitoring, dashboards, and alerting for production AI applications.
  • Define and track operational metrics covering AI quality, reliability, latency, cost, utilization, adoption, and business outcomes.
  • Work with AI engineering team to diagnose production issues across AI applications, models, prompts, retrieval systems, data, integrations, and supporting infrastructure.
  • Support root-cause analysis and coordinate resolution with AI engineers, application teams, platform teams, and managed-services providers.
  • Support AI-related incident management, problem management, operational reviews, release validation, and production-readiness activities.
  • Create and maintain service-health dashboards, runbooks, troubleshooting guidance, support procedures, and escalation paths.
  • Define and track service health indicators, SLIs, SLOs, and operational KPIs for production AI solutions.
  • Use telemetry, evaluations, and production data to validate fixes, releases, configuration changes, and system improvements.

Support AI Risk and Operational Governance
  • Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements.
  • Operationalize evaluation processes and support AI Governance team to monitor AI performance, regressions, drift, grounding, retrieval quality, and overall effectiveness.
  • Support the collection and retention of operational evidence, including model and prompt versions, evaluation results, incidents, exceptions, and corrective actions.
  • Identify material changes in AI behavior and help ensure they are evaluated, documented, and appropriately addressed.
  • Analyze trends and proactively identify degradation, drift, capacity constraints, reliability risks, and quality issues before they become production incidents


Key Outcomes
  • AI teams receive timely and consistent infrastructure and operational support.
  • AI use cases move efficiently from experimentation to reliable production operation.
  • Reusable infrastructure and operational patterns are applied across AI initiatives.
  • End-to-end visibility exists across AI applications and supporting services.
  • AI quality, reliability, cost, usage, and business impact are consistently measured.
  • Production issues are detected, diagnosed, and resolved more quickly.
  • Releases result in fewer regressions and operational disruptions.
  • Secure, compliant, and audit-ready AI platforms meeting enterprise governance and risk requirements


Qualifications
  • Bachelor's degree in computer science, engineering, information technology, or a related field, or equivalent practical experience.
  • 8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering.
  • Experience supporting cloud-hosted, distributed, data-intensive, or AI-enabled production applications.
  • Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting.
  • Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines.
  • Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed-services providers.
  • Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations.
  • Experience with automation, CI/CD pipelines, infrastructure-as-code, configuration management, and release-management practices.
  • Experience operating or supporting LLM, generative AI, machine-learning, RAG, or agent-based applications in production is strongly preferred.
  • Familiarity with AI observability, evaluation, model monitoring, hallucination detection, grounding, retrieval quality, model or data drift, and prompt-related risks.
  • Familiarity with vector databases, model APIs, AI gateways, prompt-management platforms, or agent orchestration frameworks.
  • Strong analytical, documentation, communication, and cross-functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives.
  • Experience in insurance, financial services, or another regulated industry, along with relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications, is preferred.


Pay Range:

Anticipated Hiring Range:
  • $130,000 - $150,000 annual base salary depending on experience, qualifications, and geographic location


Benefits:

We are proud to offer our full-time regular employees a robust benefits suite that includes:
  • Competitive base salary plus incentive plans for eligible team members
  • 401(K) retirement plan that includes a company match of up to 6% of your eligible salary
  • Free basic life and AD&D, long-term disability and short-term disability insurance
  • Medical, dental and vision plans to meet your unique healthcare needs
  • Wellness incentives
  • Generous time off program that includes personal, holiday and volunteer paid time off
  • Flexible work schedules and hybrid/remote options for eligible positions
  • Educational assistance


Similar Jobs

More Jobs at The Mutual Group

More Information Technology Jobs

Find similar Lead AI Infrastructure Operations Engineer jobs: