International Data Group

Principal, AI Operations & Infrastructure Automation

International Data Group$139K — $240K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in infrastructure, platform engineering, or IT operations, including 3+ years leading automation or AIOps initiatives.
  • Proven success in redesigning operations around automation, shifting team workflows rather than implementing isolated tools.
  • Hands-on cloud infrastructure automation experience (AWS preferred) with infrastructure-as-code tools like Terraform or CloudFormation.
  • Familiarity with AIOps and automation platforms, with a discerning approach to AI's role in operations.
  • Practical experience in Python or similar language, building integrations with REST APIs.
  • Proficient with observability and monitoring platforms such as Datadog or Splunk, for effective service instrumentation.
  • Strong leadership skills, capable of establishing and managing a transformation roadmap and meeting ambitious deadlines.

Responsibilities

  • Lead the overhaul of the infrastructure operating model, automating key workflows to create self-healing processes.
  • Develop and implement a phased transformation roadmap with clear milestones and automation targets.
  • Drive change management, establishing new roles and operating processes as automation is integrated.
  • Establish AIOps capabilities through platform selection and deployment, enhancing incident management with anomaly detection and predictive alerting.
  • Implement large language model-assisted tools for triaging incidents and automating ticket management processes.
  • Create predictive models for capacity and service failures, enabling proactive actions to prevent breaches.
  • Champion infrastructure-as-code standards to streamline automation and convert manual processes into efficient workflows.

Benefits

  • 15 vacation days (prorated based on start date)
  • 12 company-paid holidays
  • 6 paid sick days (prorated based on start date; may vary by state)
  • Medical, dental, and vision coverage
  • 2 floating holidays (prorated based on start date)
  • 1 volunteer day
  • 401(k) company match (up to 3% on first 6% contributions)
  • Company-paid short-term disability
  • Company-paid life insurance
  • Company-paid parental leave
Full Job Description
Overview

About the Role & Team

IDC is modernizing how infrastructure is run - shifting from manual, ticket-driven operations to an automated, AI-informed model that keeps pace with the speed of the business. We're looking for a Principal of AI Operations & Infrastructure Automation to lead this transformation: designing and standing up the automation, tooling, and operating model that will define how IDC's infrastructure team works going forward. This is a hands-on, high-velocity leadership role for someone who has actually built AIOps and infrastructure-as-code capability before - not just overseen it.

You'll personally write and review automation, get into the tooling directly, and lead by example for the team you manage. You'll partner closely with the CIO, Cyber & Infrastructure leadership, and AWS as a strategic cloud and co-development partner to re-architect how infrastructure operations get done. This role builds and leads the team responsible for automation - setting technical direction, reviewing their work, and growing their skills as the operating model evolves. Additional responsibilities include:
  • Own the end-to-end redesign of the infrastructure operating model, moving core workflows (provisioning, monitoring, incident response, patching, capacity management) from manual execution to automated, self-healing pipelines.
  • Define a phased transformation roadmap with clear milestones, automation-coverage targets, and go-live gates - and drive execution against it at pace.
  • Lead change management for the infrastructure team through the transition - defining new roles, skills, and ways of working as automation is adopted.
  • Build the AIOps layer; evaluate, select, and deploy AIOps platforms and agentic tooling - anomaly detection, predictive alerting, automated remediation - suited to IDC's AWS-centric environment. Implement event correlation and noise reduction so a hundred alerts resolve into one actionable incident with a probable root cause attached.
  • Stand up LLM (large language model)-assisted triage that classifies, enriches, and routes incidents, pulling relevant runbooks, past resolutions, and change history into the ticket automatically.
  • Develop predictive capacity and failure models for critical services, and act on them before thresholds are breached.
  • Automate end to end, build and champion an infrastructure-as-code (IaC) standard across the organization, converting legacy manual processes into version-controlled, repeatable automation.
  • Deliver self-healing automation and closed-loop remediation for the highest-volume recurring incidents - restart, rollback, scale, reroute, and verify without a human in the loop. Automate service request fulfillment: access provisioning, environment setup, onboarding and offboarding, and routine change execution.
  • Own CI/CD (continuous integration/continuous delivery-deployment) pipelines for operational tooling, including testing and safe rollback for automations that touch production.
  • Lead the team and prove the results Directly manage and develop a team of infrastructure and automation engineers - hiring, coaching, and setting technical direction - while staying hands-on in the tooling and code yourself.
  • Partner with AWS and other strategic vendors to pilot and scale automation and AI-driven operations capabilities, including Bedrock/AgentCore-based tooling where relevant.
  • Establish new operational KPIs - automation coverage, mean time to detect and resolve, deployment frequency, toil eliminated - and report progress in concise, executive-ready formats.
  • Set guardrails for AI in operations: human-in-the-loop thresholds, blast-radius limits, audit logging, evaluation of model output quality, and clear rollback paths.
  • Ensure automated operations meet IDC's security, compliance, and resiliency standards, working closely with Cyber & Compliance.

What You Bring

Required
  • 8+ years in infrastructure, platform engineering, or IT operations, including 3+ years leading automation or AIOps initiatives.
  • Demonstrated track record of redesigning an operations function around automation - not just introducing point tools but changing how a team works.
  • Deep hands-on experience with cloud infrastructure automation (AWS strongly preferred), infrastructure-as-code (Terraform, CloudFormation, or similar), and CI/CD pipelines.
  • Working knowledge of AIOps and agentic automation platforms, and a clear point of view on where AI genuinely improves operational outcomes versus adds noise.
  • Practical automation skill in Python or an equivalent language, including building integrations against REST APIs and operational tooling.
  • Hands-on experience with observability and monitoring platforms - Datadog, Splunk, Dynatrace, New Relic, Prometheus/Grafana, Elastic, or similar - including instrumenting services and building useful alerting.
  • Strong program leadership skills: able to define a roadmap, sequence dependencies, and hit aggressive milestones.
  • Excellent executive communication skills - comfortable distilling complex technical transformation into clear, concise updates for senior leadership.

Preferred
  • Experience operating in a regulated or compliance-sensitive environment (SOC 2, ISO 27001, or similar).
  • Prior experience partnering directly with a hyperscaler (AWS, Azure, or GCP) on joint automation or AI initiatives.
  • Background in site reliability engineering (SRE) practices and observability tooling.
  • Experience with agentic AI frameworks, RAG pipelines, or building internal AI copilots for technical teams.
  • Familiarity with ITSM platforms and process - ServiceNow, Jira Service Management, or similar - and comfort automating against them.
  • Familiarity with FinOps practices and cloud cost automation.
  • Relevant certifications (cloud architect or engineer, Kubernetes, ITIL v4, Terraform).

What success looks like

First 90 days

You know our estate - tooling, telemetry, ticket volume, and the top recurring incident categories. The transformation roadmap is published with milestones and automation-coverage targets, and at least one automation is live and measurably reducing manual work.

First two quarters

A documented automation-first operating model is adopted across the infrastructure team. Intelligent alert correlation is live on our highest-noise services, and baseline AIOps KPIs are instrumented and reported to the CIO.

One year

Measurable improvement in mean time to detect and resolve incidents. A growing share of routine operational workflows runs without manual intervention, and infrastructure changes are consistently deployed through version-controlled, automated pipelines.

Why This Role Stands Out

At IDC, your work helps shape how the world understands technology and where it goes next. You collaborate with curious, high-caliber colleagues who value rigor, integrity, and shared success. As the premier global provider of trusted technology intelligence, IDC equips business and technology leaders with the evidence they need to make confident decisions. Our insights inform strategy, investment, and innovation across industries and regions.

Recognized by IIAR as Analyst Firm of the Year for five consecutive years, IDC sets the standard for credibility and impact. With more than 1,000 analysts worldwide and a truly global perspective, we combine deep expertise with practical relevance. Here, your ideas matter, your voice is heard, and your contributions provide the insights leaders rely on every day. It is meaningful work, backed by a culture that supports growth, collaboration, and long-term career development with a globally respected brand.

What We Offer

15 vacation days (prorated based on start date)

12 company-paid holidays

6 paid sick days (prorated based on start date; may vary by state)

Medical, dental, and vision coverage

2 floating holidays (prorated based on start date)

1 volunteer day

401(k) company match (IDC matches 3% on the first 6% of employee contributions)

Company-paid short-term disability

Company-paid life insurance

Company-paid parental leave

Compensation Transparency

At IDC, we are committed to fair and equitable pay practices. Employees are compensated equitably for their work, aligned with their skills and experience. Salary and incentive structures are determined through a rigorous process that considers experience, education, certifications, role-specific requirements, internal equity, and verified U.S. market data from an independent third-party partner.

The expected total annual compensation, depending on location and experience, is between $139,000-$240,000, inclusive of base salary and variable compensation.

If this role relocates to a different country, the salary range will be updated to reflect that country's range, rather than a currency conversion of the original range.

About International Data Group

International Data Group (IDG) is a technology media, data and marketing services company. IDG publishes more than 300 magazines and newspapers, including CIO, Computerworld, GamePro, InfoWorld, Macworld, Network World, PC World, and TechHive. IDG also produces conferences and events, webcasts, and research services. IDG is headquartered in Framingham, Massachusetts, and has offices worldwide. The company was founded in 1964 by Patrick McGovern, who passed away in 2014. IDG is now owned by China Oceanwide Holdings Group, a Chinese conglomerate with interests in financial services, real estate, and technology.
Learn more about International Data Group
Size
3,000 employees
Industry
Founded
1964

Similar Jobs

More Jobs at International Data Group

More Information Technology Jobs

Find similar Principal, AI Operations & Infrastructure Automation jobs: