Senior AIOps ML Engineer

Prophecy Technologies

$130K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in data engineering and machine learning
  • Strong proficiency in Kafka and streaming technologies (Flink/Spark)
  • Experience with Lakehouse architectures (Delta Lake/Iceberg)
  • Advanced skills in machine learning for anomaly detection and time-series analysis
  • Proficient in MLOps practices including feature store management and model retraining
  • Strong SQL and Python programming skills
  • Familiarity with observability tools like Dynatrace, OTel, and APM.

Responsibilities

  • Design and evolve the Lakehouse schema for observability data at petabyte scale.
  • Build and maintain robust ingestion pipelines ensuring exactly-once semantics.
  • Implement dbt transformation models for generating mart-ready tables.
  • Define and enforce data quality contracts with established SLAs.
  • Optimize query performance using advanced partitioning and materialization techniques.
  • Deploy machine learning models for streaming anomaly detection and incident forecasting.
  • Develop low-latency streaming inference pipelines for real-time anomaly scoring.

Benefits

  • Opportunity to work with cutting-edge machine learning and data technologies.
  • Collaborative and innovative team environment.
  • Focus on professional growth and mentoring.
  • Involvement in high-impact projects across multiple domains.
  • Access to resources and training for further skills development.
Full Job Description
Role Overview:

The role is for a Senior AIOps ML Engineer responsible for designing, building, and optimizing a Lakehouse architecture for multi-domain observability data at petabyte scale. This includes developing robust data ingestion and transformation pipelines, implementing advanced machine learning models for AIOps (anomaly detection, root-cause analysis, incident forecasting), and managing the full MLOps lifecycle. The engineer will also focus on platform productionization, security, compliance, and fostering engineering standards.

Key Responsibilities:
  • Design and evolve the Lakehouse schema (Delta Lake / Apache Iceberg) for multi-domain observability data at petabyte scale.
  • Build and maintain robust ingestion pipelines from the OTel Collector through Kafka to the Lakehouse, ensuring exactly-once semantics and strict schema enforcement.
  • Implement dbt transformation models to generate mart-ready, denormalized fact and dimension tables for each of the six domains.
  • Define and enforce data quality contracts, establishing SLAs for data freshness, completeness, and cardinality budgets per mart.
  • Optimize query performance utilizing partitioning strategies, Z-ordering, bloom filters, and materialized views tailored for time-series patterns.
  • Design, train, and deploy machine learning models for streaming multivariate anomaly detection, root-cause analysis, and incident forecasting across all six mart domains.
  • Build low-latency streaming inference pipelines (Flink / Spark Streaming) for real-time anomaly scoring on APM, infrastructure, and security signals.
  • Develop sophisticated log intelligence models-including clustering (DRAIN3 / LogBERT), NLP classification, and error deduplication-over the Log mart.
  • Implement unsupervised and semi-supervised methods for User Experience frustration detection and KPI correlation analysis.
  • Own the ML feature store, managing feature engineering, versioning, backfill pipelines, and point-in-time correct joins for training datasets.
  • Instrument model performance tracking, including drift detection, accuracy monitoring, and automated retraining triggers.
  • Design and operate the end-to-end AIOps workflow, spanning signal ingestion, feature computation, model inference, alert routing, and auto-remediation hooks.
  • Build high-performance model serving infrastructure-supporting real-time REST/gRPC endpoints and async batch scoring-with strict p99 latency SLOs.
  • Integrate AIOps insights with incident management platforms (PagerDuty, Opsgenie) and internal runbooks to deliver enriched, noise-reduced alerting.
  • Define and publish metrics from the Business KPI mart to quantify the blast radius, revenue loss, and affected user counts for each incident.
  • Partner with the Security team to build the Security mart schema, including threat feed ingestion, UEBA baselines, and CVE correlation pipelines.
  • Train anomalous-access and lateral-movement detection models, tuning precision/recall thresholds in collaboration with the SOC team.
  • Ensure all data handling across the marts adheres strictly to data residency requirements, PII masking standards, and audit-log protocols.
  • Define telemetry schema contracts with the OTel Instrumentation team to guarantee high upstream signal quality for downstream ML models.
  • Author ML platform RFCs and contribute actively to observability data model standards across the broader engineering organization.
  • Mentor junior ML and data engineers, and conduct rigorous design reviews for new mart schemas and model architectures.

Required Skills:
  • Kafka + Streaming (Flink/Spark)
  • Lakehouse (Delta / Iceberg)
  • ML (Anomaly detection + time-series)
  • Observability (OTel, APM, Logs)
  • MLOps (feature store, drift, retraining)
  • SQL + Python (strong)
  • Dynatrace

Qualifications:
  • 10+ years of experience

Similar Jobs

More Jobs at Prophecy Technologies

More Information Technology Jobs

Find similar Senior AIOps ML Engineer jobs: