AI Operations and Monitoring Engineer

Choctaw Nation of Oklahoma

$90K — $110K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science or related field, or 4 years relevant experience.
  • 3 years' experience in DevOps, Site Reliability Engineering (SRE), or Machine Learning Operations (MLOps).
  • Proficiency with tools like Kubernetes, Docker, and various cloud services.
  • Experience with observability tools to monitor system performance.
  • Familiarity with machine learning deployment workflows.

Responsibilities

  • Monitor AI/ML system health, drift, and anomalies to ensure optimal functioning.
  • Build and maintain dashboards and alerts for real-time system performance insights.
  • Implement and manage CI/CD pipelines for efficient model deployments.
  • Ensure stability and optimization of training and deployment environments.
  • Develop automation for scaling and retraining AI models based on demand.
  • Respond to incidents and outages, conduct thorough root cause analysis.
  • Collaborate on security and compliance to meet organizational standards.
  • Document procedures through runbooks and maintain audit logs for operational consistency.

Benefits

  • Hybrid work environment with flexible scheduling.
  • Weekly earned wage access available.
Full Job Description
Job Description

Monday-Friday 8:00AM-4:30PM| Hybrid Position| Weekly Earned Wage Access is an option for this position.

Job Purpose or Goals: The AI Operations and Monitoring Engineer is responsible for ensuring the reliability, performance, documentation, and compliance of production AI/ML systems, keeping them stable and functioning as intended. They play a key role in maintaining system uptime and driving effective incident response to minimize outages and protect overall service quality.

Tasks:

1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.

2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.

3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.

4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.

5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.

6. Respond to incidents and outages to restore services quickly, conduct root cause analysis, and prevent future disruptions to AI/ML workloads.

7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.

8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.

9. Performs other duties as may be assigned.

Job Requirements:

Bachelor's degree in computer science or related field, or 4 years relevant professional experience

3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services

Experience with observability tools

Familiarity with ML deployment workflows

Responsibilities

1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.

2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.

3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.

4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.

5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.

6. Respond to incidents and outages to restore services quickly, conduct root cause analysis, and prevent future disruptions to AI/ML workloads.

7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.

8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.

9. Performs other duties as may be assigned.

Qualifications

Bachelor's degree in computer science or related field, or 4 years relevant professional experience

3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services

Experience with observability tools

Familiarity with ML deployment workflows

Similar Jobs

More Jobs at Choctaw Nation of Oklahoma

More Information Technology Jobs

Find similar AI Operations and Monitoring Engineer jobs: