The Manager, NOC Operations & Service Reliability is responsible for establishing, operating, and continuously improving the Valeris Network Operations Center (NOC). The role provides leadership for 24x7 monitoring, incident operations, priority incident and service restoration, operational governance, partner management, reporting, service onboarding, escalation management, and continuous service improvement.
This position serves as the operational owner for service visibility, monitoring coverage, runbook governance, ServiceNow integration, operational reporting, and NOC service delivery across infrastructure, cloud, SaaS, applications, Kubernetes, APIs, databases, storage, and backup services.
The role ensures operational readiness, accountability, and execution of the Valeris NOC operating model while maintaining alignment with business priorities, service commitments, and operational objectives. The NOC operating model includes 24x7 monitoring, event triage, incident coordination, runbook-based remediation, structured escalation, ServiceNow lifecycle management, dashboards, reporting, and continuous improvement.
Key Responsibilities:NOC Operations & Service Monitoring- Lead 24x7 monitoring operations for network, infrastructure, cloud, and SaaS systems
- Ensure visibility across:
- Network (LAN/WAN, firewalls, load balancers)
- Servers, compute, storage, Kubernetes environments, API services, databases
- Azure IaaS/PaaS and SaaS platforms (M365, etc.)
- Establish service-level monitoring aligned to business services and dependencies
- Oversee alerting strategy (thresholds, anomaly detection, service impact)
- Drive end-user experience monitoring and SaaS observability
- Define service criticality, ownership, escalation paths, monitoring standards, and alert requirements.
- Drive monitoring quality, service visibility, alert tuning, and noise reduction initiatives.
- Ensure all monitored services have documented owners, escalation contacts, runbooks, and support procedures.
NOC Deployment & Service Onboarding - Establish and maintain service inventory and ownership records.
- Define service onboarding standards and operational acceptance requirements.
- Lead monitoring deployment activities.
- Validate operational readiness before production support acceptance.
- Ensure onboarded services include:
- Monitoring
- Alerting
- Escalation paths
- Runbooks
- ServiceNow integration
- Reporting requirements
- Govern service transition into 24x7 operational support.
Incident, Problem & Change Management (ITIL) - Own escalation model and execution of:
- Incident management
- Problem management (root cause analysis)
- Change coordination (as applicable)
- Ensure all incidents include:
- Business impact
- Root cause
- Resolution actions
- Lead P1/P2 incident response, including bridge calls and leadership communications
- Drive MTTR reduction and SLA adherence
Vendor & Service Governance - Own the operational relationship with NOC service providers and staff augmentation resources.
- Govern support boundaries and escalation responsibilities.
- Manage service performance review cadence.
- Conduct:
- Daily Service Review (Stand up)
- Weekly Operational Reviews
- Monthly Service Reviews
- Quarterly Business Reviews
- Track SLA attainment and service delivery performance.
- Review partner reports and operating metrics.
- Drive accountability for service commitments.
- Lead continuous service improvement initiatives.
Runbooks, Automation & Knowledge Management - Establish and maintain operational runbooks.
- Govern runbook ownership and lifecycle management.
- Ensure integration with ServiceNow knowledge management.
- Conduct periodic runbook reviews.
- Drive automation and approved recovery actions.
- Identify opportunities for automation and operational efficiency.
- Maintain operational documentation standards and knowledge transfer processes.
Operational Reporting & Visibility - Define operational reporting standards.
- Deliver operational dashboards and management reporting.
- Establish service performance metrics and scorecards.
- Provide operational visibility across:
- Infrastructure
- Cloud
- Network
- SaaS
- Applications
- Kubernetes
- Backup systems
- Lead development of operational and executive dashboards.
Process Improvement & Operational Maturity - Standardize operational processes across the NOC.
- Improve:
- Monitoring coverage
- Alert quality
- Ticket quality
- Runbook coverage
- Escalation effectiveness
- Service visibility
- Drive operational maturity initiatives.
- Identify recurring issues and trend analysis opportunities.
- Lead service improvement plans and preventative action initiatives.
Team Leadership & Development - Lead the NOC organization.
- Define operational roles, responsibilities, and escalation paths.
- Develop a culture of:
- Accountability
- Operational excellence
- Continuous improvement
- Service ownership
- Support development of cloud operations, monitoring, automation, and service management capabilities.
- Utilize Valeris' values as the driving force behind the team's success
- On time adherence to training deadlines for all corporate policies and procedures
- Ensure all SOPs are followed with consistency
- Perform additional tasks or projects as assigned
Qualifications:Technical- 8+ years in IT Infrastructure / Operations (with leadership experience)
- Azure (IaaS/PaaS, governance)
- Networking (firewalls, routing, load balancing)
- Enterprise systems (AD/Entra, M365, VMware)
- Application observability
- Synthetic transaction monitoring
- Monitoring platforms (e.g., Datadog, Azure Monitoring)
Process & Leadership - Strong ITIL experience (Incident, Problem, Change)
- Experience managing NOC or enterprise operations teams
- Ability to drive operational transformation and standardization
- Effective communication skills for management and technical audiences
Success Measures - Incident Detection Performance
- Service Restoration Performance
- SLA adherence across incidents and requests
- MTTR and incident volume reduction
- Monitoring coverage and alert accuracy
- Automation adoption (reduction in manual effort)
- System availability and performance
- Process standardization and audit readiness
- Certificate lifecycle automation ownership
- SaaS monitoring and vendor integration maturity
- Alert Quality Improvement
- Reduction in Recurring Incidents
Physical Demands & Work Environment- While performing the duties of this job, the employee is regularly required to talk or hear. The employee is frequently required to sit for long periods of time, use hands to type, handle or feel; and reach with hands and arms. Prefer candidates who can type at least 35 words per minute with 97% accuracy.
- Although very minimal, flexibility to travel as needed is preferred.
- This job operates in a professional office environment. This role routinely uses standard office equipment such as computers, phones, photocopiers, etc.
Any offer of employment is contingent upon the successful completion of a background check and, depending on the position, a drug screen in accordance with company standards. Please note that this job description is not intended to be an exhaustive list of all duties, responsibilities, or activities associated with the position. Responsibilities and tasks may be modified at any time, with or without notice.