Software Engineering Director, AI for Incident Response

Meta

• $180K — $220K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of software engineering experience in technical leadership roles
  • 7+ years of engineering management experience leading multiple teams
  • Proven track record in large-scale distributed systems and reliability platforms
  • Experience in managing critical, always-on production workloads
  • History of building widely adopted platforms for cross-organizational outcomes
  • Expertise in evaluating automated systems and implementing production safeguards
  • Strong ability to hire, onboard, and develop engineering talent across levels

Responsibilities

  • Define and drive engineering strategy for AI-powered incident response
  • Lead architecture for an autonomous incident response agent
  • Establish a human-on-the-loop model with clear operating policies
  • Build Opsmate as an extensible platform for incident response
  • Create evaluation systems to connect agent performance with reliability metrics
  • Collaborate with product groups and infrastructure teams for shared success
  • Cultivate a high-performing engineering organization focused on measurable results

Benefits

  • Opportunity to shape the technology and operating model for incident response
  • Work within one of the largest and most complex production environments
  • Lead cutting-edge AI initiatives intersecting with operational excellence
  • Collaborate across diverse teams to drive tangible outcomes
  • Engage in a culture of rapid learning and accountability for customer impact
Full Job Description
Meta is seeking a Software Engineering Director to lead the organization building the future of incident response. As agents increasingly build and operate infrastructure, Meta needs reliability systems that can reason and act at machine speed. This leader will drive Meta's AI-powered incident response agent, Meta's AI agent for incident response, toward a human-on-the-loop model: the agent autonomously gathers context, investigates incidents, identifies root causes, creates a safe mitigation plan, and executes the appropriate action, while humans set boundaries, supervise consequential actions, and intervene when judgment is required. This role combines frontier agent capabilities with one of the world's largest and most complex production environments. You will define the technical, product, and organizational strategy, lead an engineering organization with real production systems, and partner across infrastructure and product groups to make incident response faster, safer, and dramatically less labor-intensive. Few roles offer the opportunity to define both the technology and operating model for autonomous incident mitigation at this scale.

Responsibilities

Define and drive the engineering strategy and roadmap for AI-powered incident response, progressing from investigation and diagnosis to safe, end-to-end mitigation and continuous learning
• Lead the architecture and delivery of an agent that reasons across telemetry, code, configurations, deployments, service dependencies, and incident history to identify root causes and take appropriate action
• Establish a human-on-the-loop operating model with progressive autonomy, explicit policy and permission boundaries, independent validation, auditability, rollback, and clear escalation paths
• Build Opsmate as an extensible platform that combines shared, out-of-the-box capabilities with domain-specific skills and agents contributed by product groups and individual on-call engineers
• Establish trusted evaluation and measurement systems that connect agent quality to customer and reliability outcomes, including investigation accuracy, adoption, time to mitigation, and reduced operational effort
• Partner deeply with product groups, infrastructure teams, product management, and data science to understand what success means for their services and build toward shared outcomes
• Ensure the platform itself meets a high bar for reliability, latency, scale, security, privacy, and cost in Meta's production environment
• Build and grow an engineering organization that consistently delivers measurable results by attracting, developing, and retaining engineers and engineering managers across distributed systems, reliability, and applied AI
• Create a culture of technical depth, rapid learning, operational excellence, and accountability for realized customer impact
• Contribute to company-wide engineering and AI initiatives, including recruiting, technical standards, and responsible AI practices

Minimum Qualifications
• 10+ years of software engineering experience, including experience in technical leadership roles
• 7+ years of engineering management experience, including leading multiple teams or managers and delivering products or systems with measurable impact
• Experience defining and executing engineering strategy for large-scale distributed systems, reliability platforms, developer infrastructure, or production AI systems
• Experience building and operating systems that support critical, always-on production workloads
• Track record of building broadly adopted platforms and driving outcomes across organizational boundaries
• Experience evaluating complex automated systems and establishing safeguards for trustworthy production operation
• Experience building engineering teams through hiring, onboarding, and developing engineers and managers at multiple levels
• Experience partnering cross-functionally with product, data science, and infrastructure organizations

Preferred Qualifications
• Experience with incident management, SRE, observability, oncall operations, or automated remediation at significant scale
• Experience designing safe execution systems, including identity, authorization, sandboxing, validation, rollback, and auditability
• Experience building extensible platforms that allow domain teams to contribute specialized capabilities while preserving a coherent user experience and control plane
• Track record of taking a technically ambitious product from early adoption to trusted, broad production use
• Experience managing other engineering managers and building a deep technical and management bench across multiple levels
• Experience operating LLM-based or tool-using agents in production, including evaluation, orchestration, and progressive autonomy

Similar Jobs

More Jobs at Meta

More Information Technology Jobs

Find similar Software Engineering Director, AI for Incident Response jobs: