Production Network Engineer
Responsibilities
Lead the design and delivery of large-scale production network projects spanning data center fabrics, backbone infrastructure, and edge network services
• Define and drive the technical strategy and roadmap for network reliability, capacity, and operational efficiency across multiple teams
• Develop and implement automation frameworks to reduce manual operational overhead and accelerate network design synthesis, network build and configuration management
• Own cross-functional incident response and post-incident review processes, driving root cause analysis and systemic improvements to reduce recurrence
• Identify and resolve complex network performance, routing, and reliability issues across multi-vendor, multi-protocol production environments
• Collaborate with network architecture, capacity planning, and software engineering teams to align infrastructure investments with evolving AI and product demand forecasts
• Establish and refine operational standards, runbooks, and monitoring frameworks to improve network observability and reduce mean time to detection and resolution
• Contribute to organizational strategy by defining scalable approaches to network operations and influencing tooling and platform decisions across engineering teams
• Mentor other engineers on network engineering best practices, operational discipline, and systems thinking across the production environment
• Leverage AI-integrated workflows to accelerate network analysis, anomaly detection, and documentation, sharing learnings to scale adoption across the team
Minimum Qualifications
• Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
• 8+ years of experience in production network engineering, including design, deployment, and operations of large-scale data center or backbone network infrastructure
• Experience with routing protocols and network technologies including BGP, OSPF, IS-IS, MPLS, and large-scale Ethernet fabrics in a production environment
• Experience developing and deploying network automation using scripting or programming languages such as Python to manage configuration, provisioning, or monitoring at scale
• Experience leading cross-functional network reliability or infrastructure projects, including driving incident response, root cause analysis, and systemic remediation
• Experience influencing technical decisions and network strategy across multiple engineering teams through written proposals, design reviews, and stakeholder alignment
Preferred Qualifications
• Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
• Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
• Experience operating and troubleshooting networks at hyperscale, including multi-vendor spine-leaf data center fabrics and global backbone interconnects
• Experience with network observability platforms, traffic engineering, and capacity modeling in high-throughput production environments
• Experience integrating AI tools to redesign network operations workflows and deliver measurable improvements in efficiency or reliability outcomes
• Familiarity with software-defined networking principles, network operating system internals, or co-development of network management platforms with software engineering teams
• Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies