OverviewEdgewater Federal Solutions is seeking a Subject Matter Data Center Network Engineer to support a major national laboratory.
Responsibilities
- Provide technical knowledge and analysis for specialized applications and operational environments.
- Perform functional systems analysis, design, integration, and documentation.
- Apply advanced principles, methods, and knowledge to solve complex technical problems and develop automated solutions for engineering and scientific applications.
- Assist with system improvements, optimization, development, and maintenance efforts in areas such as information systems architecture, networking, telecommunications, automation, communications protocols, risk management, software lifecycle management, software development methodologies, and modeling and simulation.
- Analyze user needs, define functional requirements, and develop plans to address moderately complex to extremely complex systems.
- Develop recommendations for system improvements and provide expertise recognized within the professional community.
Qualifications
- Requires BS in relevant discipline plus a minimum of 3 years, or more, of directly related experience that demonstrates the knowledge, skills, and ability to perform the duties.
- In lieuof degree, 9 years of Related experience may be substituted for relevant education and vice versa.
- Ability to obtain & maintain a U.S. Dept. of Energy Clearance
- U.S. Citizenship is required.
Required Skills:
- Experience working in missioncritical production environments with change control, incident response, and structured troubleshooting.
- Ability to work safely and effectively in live data center spaces (DCFIT coordination, raised floor environments, cabling standards).
- Strong analytical and troubleshooting abilities, with a demonstrated ability to drive issues to resolution during outages.
- Clear verbal and written communication skills.
- Ability to participate in oncall rotations and support afterhours change windows if needed.
Core Network Engineering Skills
- Strong understanding of L2/L3 networking fundamentals: VLANs, STP, LACP, static routing, OSPF, BGP, ACLs, QoS.
- Experience configuring and supporting major network vendors (any of): Cisco, Arista, Juniper, Mellanox/NVIDIA Networking.
- Familiarity with multipath topologies, spineleaf architectures, and highbandwidth fabrics.
- Understanding of basic firewall and segmentation concepts (paths, zones, NAT, routing symmetry).
- Familiarity with enterprise cabling best practices: fiber types (LR/SR/ER), transceiver selection, rack elevation awareness.
- Working knowledge of monitoring and telemetry: CloudVision, Netropy/Apposite tools, Nagios/Prometheus/Grafana, or similar.
- Ability to follow SOPs/MOPs, support planned change windows, and adhere to structured rollback/validation procedures.
Desired Skills:
Advanced Network Expertise
- Handson experience with Arista EOS, Cisco NXOS, Juniper JunOS, or Mellanox/NVIDIA platforms supporting 40/100/400Gbps networks.
- Experience with BGP tuning, trafficengineering policies, ECMP management, asymmetricrouting detection, and complex routemap design.
- Familiarity with EVPN/VXLAN, modern DC underlays, and highavailability routing gateway designs (HA router pairs, blue/green cutovers)
- Experience troubleshooting firewall pathing issues, zone interactions, NAT64/NAT policy behavior, and segmentation used for HPC.
HPCSpecific Skills
- Exposure to HPC networks and interconnects:
- RDMA, RoCE, basic InfiniBand concepts (subnet managers, fabric behavior).
- Understanding of congestion behaviors common in HPC (parallel I/O bursts, GPUnode communication patterns).
- Familiarity moving or supporting HPC data flows across multisite infrastructures (trilab WAN, DisCom, IHPC routing).
- Understanding of distributed HPC storage systems (Lustre, BeeGFS, Isilon, Spectrum Scale) and the routing or MTU constraints around them.
- Ability to work with HPC teams during large cluster deployments (rack placement coordination, switch firmware updates, cabling checks).
Tooling & Automation
- Handson experience with automation or configuration management tools (Ansible, Python, Terraform).
- Ability to build or maintain dashboards and telemetry for HPC network utilization (CloudVision pipelines, Netropy test data).
- Familiarity with gNMI/gRPC or streaming telemetry for performance analysis and anomaly detection.
Reliability, Capacity & Performance
- Understanding of highavailability designs, reduced failure domains, and deterministic failover patterns in HPC networks.
- Experience monitoring and optimizing latency, packet loss, MTU mismatches, and congestion across crypto tunnels or 100Gbps transport.
- Ability to assist with capacity planning for rapidly growing HPC systems (rack density, port utilization, fiber paths, WAN circuit expansions).