Senior AIOps ML Engineer

Prophecy Technologies

$130K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in data engineering or machine learning roles
  • Proficiency in Kafka and streaming technologies (Flink/Spark)
  • Experience with Lakehouse architectures such as Delta Lake or Apache Iceberg
  • Strong background in machine learning models, particularly for anomaly detection and time-series analysis
  • Familiarity with observability tools like OpenTelemetry (OTel) and Application Performance Monitoring (APM)
  • Expertise in MLOps practices including feature store management, model drift detection, and automated retraining
  • Advanced skills in SQL and Python programming
  • Knowledge of Dynatrace for monitoring performance.

Responsibilities

  • Design and evolve the Lakehouse schema for multi-domain observability data
  • Build and maintain data ingestion pipelines ensuring schema enforcement
  • Implement transformation models to prepare data for analysis
  • Define data quality contracts and establish SLAs for various metrics
  • Optimize query performance using advanced strategies tailored for time-series data
  • Deploy machine learning models for streaming anomaly detection and incident forecasting
  • Build low-latency streaming inference pipelines for real-time applications
  • Develop log intelligence models and conduct analysis on log data
  • Implement methods for user experience monitoring and correlation analysis
  • Manage the ML feature store and oversee feature engineering processes
  • Track model performance and establish monitoring protocols
  • Design and operate end-to-end AIOps workflows with auto-remediation capabilities
  • Build infrastructure for high-performance model serving with strict latency requirements
  • Integrate AIOps insights into incident management systems
  • Define metrics to measure incident impacts on business KPIs
  • Collaborate with security teams to enhance data handling and analysis procedures
  • Mentor junior engineers and perform design reviews for new data schemas.

Benefits

  • Flexible work arrangements including remote work options
  • Professional development opportunities and mentorship programs
  • Access to cutting-edge technology and tools
  • Collaborative and inclusive work environment
  • Comprehensive health and wellness programs
Full Job Description
Role Overview:

The role is for a Senior AIOps ML Engineer responsible for designing, building, and optimizing a Lakehouse architecture for multi-domain observability data at petabyte scale. This includes developing robust data ingestion and transformation pipelines, implementing advanced machine learning models for AIOps (anomaly detection, root-cause analysis, incident forecasting), and managing the full MLOps lifecycle. The engineer will also focus on platform productionization, security, compliance, and fostering engineering standards.

Key Responsibilities:
  • Design and evolve the Lakehouse schema (Delta Lake / Apache Iceberg) for multi-domain observability data at petabyte scale.
  • Build and maintain robust ingestion pipelines from the OTel Collector through Kafka to the Lakehouse, ensuring exactly-once semantics and strict schema enforcement.
  • Implement dbt transformation models to generate mart-ready, denormalized fact and dimension tables for each of the six domains.
  • Define and enforce data quality contracts, establishing SLAs for data freshness, completeness, and cardinality budgets per mart.
  • Optimize query performance utilizing partitioning strategies, Z-ordering, bloom filters, and materialized views tailored for time-series patterns.
  • Design, train, and deploy machine learning models for streaming multivariate anomaly detection, root-cause analysis, and incident forecasting across all six mart domains.
  • Build low-latency streaming inference pipelines (Flink / Spark Streaming) for real-time anomaly scoring on APM, infrastructure, and security signals.
  • Develop sophisticated log intelligence models-including clustering (DRAIN3 / LogBERT), NLP classification, and error deduplication-over the Log mart.
  • Implement unsupervised and semi-supervised methods for User Experience frustration detection and KPI correlation analysis.
  • Own the ML feature store, managing feature engineering, versioning, backfill pipelines, and point-in-time correct joins for training datasets.
  • Instrument model performance tracking, including drift detection, accuracy monitoring, and automated retraining triggers.
  • Design and operate the end-to-end AIOps workflow, spanning signal ingestion, feature computation, model inference, alert routing, and auto-remediation hooks.
  • Build high-performance model serving infrastructure-supporting real-time REST/gRPC endpoints and async batch scoring-with strict p99 latency SLOs.
  • Integrate AIOps insights with incident management platforms (PagerDuty, Opsgenie) and internal runbooks to deliver enriched, noise-reduced alerting.
  • Define and publish metrics from the Business KPI mart to quantify the blast radius, revenue loss, and affected user counts for each incident.
  • Partner with the Security team to build the Security mart schema, including threat feed ingestion, UEBA baselines, and CVE correlation pipelines.
  • Train anomalous-access and lateral-movement detection models, tuning precision/recall thresholds in collaboration with the SOC team.
  • Ensure all data handling across the marts adheres strictly to data residency requirements, PII masking standards, and audit-log protocols.
  • Define telemetry schema contracts with the OTel Instrumentation team to guarantee high upstream signal quality for downstream ML models.
  • Author ML platform RFCs and contribute actively to observability data model standards across the broader engineering organization.
  • Mentor junior ML and data engineers, and conduct rigorous design reviews for new mart schemas and model architectures.

Required Skills:
  • Kafka + Streaming (Flink/Spark)
  • Lakehouse (Delta / Iceberg)
  • ML (Anomaly detection + time-series)
  • Observability (OTel, APM, Logs)
  • MLOps (feature store, drift, retraining)
  • SQL + Python (strong)
  • Dynatrace

Qualifications:
  • 10+ years of experience

Similar Jobs

More Jobs at Prophecy Technologies

More Information Technology Jobs

Find similar Senior AIOps ML Engineer jobs: