The Production Network Engineering team is responsible for the design, deployment, operation, and continuous improvement of Meta's large-scale production data center networks. In this role, you will drive strategy and execution across complex network systems, lead cross-functional initiatives to improve reliability and performance, and apply deep subject matter expertise to solve problems at a scale few networks in the world match.
Responsibilities
Lead the design and delivery of large-scale production network projects spanning data center fabrics, backbone infrastructure, and edge network services
• Define and drive technical strategy and roadmap contributions for network reliability, capacity, and operational efficiency across multiple teams
• Develop and implement automation frameworks to reduce manual operational overhead and accelerate network design synthesis, build processes, and configuration management
• Own cross-functional incident response and post-incident review processes, driving root cause analysis and systemic improvements to reduce recurrence
• Identify and resolve complex network performance, routing, and reliability issues across multi-vendor, multi-protocol production environments
• Collaborate with network architecture, capacity planning, and software engineering teams to align infrastructure investments with evolving AI and product demand forecasts
• Establish and refine operational standards, runbooks, and monitoring frameworks to improve network observability and reduce mean time to detection and resolution
• Contribute to team-level goals by defining scalable approaches to network operations and influencing tooling and platform decisions across engineering teams
• Advise other engineers on network engineering best practices, operational discipline, and systems thinking across the production environment
• Leverage AI-integrated workflows to accelerate network analysis, anomaly detection, and documentation, sharing learnings to scale adoption across the team
Minimum Qualifications
• Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
• 6+ years of experience in production network engineering, including design, deployment, and operations of large-scale data center or backbone network infrastructure
• Experience with routing protocols and network technologies including BGP, OSPF, IS-IS, MPLS, and large-scale Ethernet fabrics in a production environment
• Experience developing and deploying network automation using scripting or programming languages such as Python to manage configuration, provisioning, or monitoring at scale
• Experience leading cross-functional network reliability or infrastructure projects, including driving incident response, root cause analysis, and systemic remediation
• Experience influencing technical decisions and network strategy across engineering teams through written proposals, design reviews, and stakeholder alignment
Preferred Qualifications
• Experience operating and troubleshooting networks at hyperscale, including multi-vendor spine-leaf data center fabrics and global backbone interconnects
• Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
• Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
• Familiarity with software-defined networking principles, network operating system internals, or co-development of network management platforms with software engineering teams
• Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
• Experience with network observability platforms, traffic engineering, and capacity modeling in high-throughput production environments
• Experience integrating AI tools to redesign network operations workflows and deliver measurable improvements in efficiency or reliability outcomes