Job DescriptionThis role is an individual contributor position responsible for designing, developing, fine-tuning, and operationalizing AI/ML capabilities for the AIOps platform within the Observability product. The candidate will work closely with the Lead Engineer, SRE teams, Platform team, Data Ingestion team, Platform DevOps team, Visualization team, and other portfolio teams.
As part of the AIOps Platform team, you will design and build intelligent systems that improve observability, incident response, and operational efficiency. This includes developing machine learning models for forecasting, anomaly prediction, alert classification, event intelligence, and causal analysis, as well as building AI agents and multi-agent workflows for RCA summarization, investigative assistance, and SRE productivity use cases.
This position will be based out of Phoenix, Arizona or Pleasanton CA.
Main responsibilities:- Design, develop, and productionize AI/ML capabilities for the AIOps platform to support intelligent observability and operational decision-making.
- Build machine learning systems for time-series forecasting, anomaly prediction, incident prediction, alert classification, noise reduction, and event correlation.
- Develop causal ML and statistical inference solutions to identify likely root causes, dependency impacts, and relationships across systems and services.
- Create and fine-tune models for incident intelligence use cases such as forecasting service degradation, capacity risk prediction, alert prioritization, and anomaly explanation.
- Design feature pipelines and model training workflows using telemetry, log, metric, trace, topology, and incident data.
- Build intelligent RCA summarization capabilities using LLMs and agentic frameworks such as LangChain and LangGraph.
- Develop AI agents and multi-agent systems for use cases such as Multi-Agent RCA, SRE Assistant, remediation guidance, incident triage, and operational knowledge retrieval.
- Design prompt orchestration, reasoning workflows, retrieval pipelines, tool usage patterns, and memory/context handling for AI agents.
- Integrate AI/ML services with observability platforms, event systems, knowledge bases, CMDB, incident management tools, and automation platforms.
- Collaborate with platform and engineering teams to build scalable model-serving and agent-serving architectures.
- Define and implement evaluation frameworks for model quality, agent effectiveness, hallucination reduction, relevance, and operational usefulness.
- Ensure AI/ML systems are scalable, reliable, explainable, and aligned with enterprise security, governance, and responsible AI practices.
- Build and maintain APIs and microservices for model inference, online scoring, batch predictions, and agent orchestration.
- Partner with SREs, observability engineers, and product stakeholders to translate operational pain points into ML and AI-driven solutions.
- Continuously improve model performance, feature quality, inference latency, agent reliability, and business impact through experimentation and monitoring.
- Establish engineering best practices for ML development, prompt engineering, evaluation, model deployment, testing, versioning, and documentation.
- Support production incident analysis for AI/ML services and drive root cause identification and remediation for model or agent failures.
- Create technical documentation covering model design, feature logic, training pipelines, evaluation metrics, deployment architecture, and agent workflows.
- Drive innovation in AI-enabled observability, causal intelligence, and agentic SRE workflows to enhance the value of the AIOps platform.
We are searching for someone with the following skills:- Strong experience designing and building AI/ML systems for real-world production use cases.
- Solid hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, and NumPy.
- Experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, or recommendation problems.
- Strong understanding of time-series modeling techniques for forecasting and operational prediction use cases.Experience with alert classification, incident prediction, event deduplication, prioritization, or
- signal correlation in observability or IT operations contexts.
- Knowledge of causal ML, causal inference, graph-based reasoning, and dependency-aware analysis techniques for RCA and impact analysis.
- Hands-on experience building AI applications using LLM frameworks such as LangChain and LangGraph.
- Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, or assistant workflows.
- Strong understanding of prompt engineering, RAG architecture, embeddings, vector stores, tool calling, memory handling, and agent evaluation techniques.
- Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, and operational systems.
- Experience building backend services and APIs for AI/ML model inference and agent orchestration.
- Good understanding of observability data such as logs, metrics, traces, topology, incidents, and alerts.
- Experience with data engineering concepts including feature engineering, data preprocessing, model pipelines, and batch or streaming inference.
- Familiarity with graph databases such as Neo4j and their use in dependency mapping, causal analysis, and knowledge-driven AI systems.
- Experience with REST APIs, microservices architecture, Docker, Kubernetes, and cloud-native deployment patterns.
- Familiarity with CI/CD, MLOps, model lifecycle management, experiment tracking, and model versioning practices.
- Knowledge of OpenTelemetry, monitoring systems, and observability platforms is highly desirable.
- Strong understanding of software engineering fundamentals, system design, and scalable architecture patterns.
- Strong analytical and problem-solving skills, with the ability to convert ambiguous operational problems into measurable AI/ML solutions.
- Excellent communication and collaboration skills to work with SREs, platform engineers, product owners, and business stakeholders.
- Self-driven mindset with strong curiosity, innovation, and the ability to learn and apply emerging AI techniques effectively.
We believe the successful candidate has these qualifications and experience:- Bachelor's degree in computer science, Information Systems, Engineering, Data Science, Artificial Intelligence, or a related field, or equivalent practical experience.
- 6 to 10 plus years of overall experience in software engineering, machine learning, or AI system development.
- 3 plus years of hands-on experience building and deploying machine learning systems in production.
- Strong experience in Python-based AI/ML development is required.
- Experience working on observability, monitoring, or AIOps-related platforms is strongly preferred.
- Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred.
- Experience in AIOps, Observability, SRE, IT operations, or incident management domains.
- Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, or remediation recommendations.
- Familiarity with knowledge graphs and graph-based ML techniques for dependency-aware intelligence.
- Experience using vector databases and retrieval frameworks for enterprise search and agentic applications.
- Experience integrating AI services with tools such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, or similar platforms.
- Familiarity with MCP-based client or agent integrations is a plus.
We also provide a variety of benefits including:- Competitive wages paid weekly
- Access to up to 50% of your earned wages before payday, via our partnership with Stream
- Associate discounts
- Health and financial well-being benefits for eligible associates (Medical, Dental, 401k and more!)
- Time off (vacation, holidays, sick pay). For eligibility requirements please visit myACI Benefits
- Leaders invested in your training, career growth and development
- An inclusive work environment with talented colleagues who reflect the communities we serve
Our Values - Click below to view video: ACI Values
A copy of the full job description can be made available to you.
#LI-MF1
About the TeamPay Transparency:Starting rates will be no less than the local minimum wage and may vary based on criteria such as location, experience, and qualifications.
Candidates with unique qualifications may be considered for compensation above this range. Benefits may include medical, dental, vision, disability and life insurance, sick pay, PTO/Vacation Pay or Flexible Time Off, paid holidays, bereavement pay, and retirement benefits (pension and/or 401k eligibility). [If applicable:] Associates in this position may be eligible for a quarterly bonus.