Bachelor's degree in Information Technology, Computer Science, or related field (or equivalent practical experience).
Experience supporting large, distributed operations (100 to 300 locations) with high availability requirements.
Strong knowledge of core infrastructure areas: networking, systems, identity, and endpoint management.
Experience managing business-critical applications in production with engineering teams.
Familiarity with cloud platforms (e.g., Azure, AWS) and their governance.
Proven leadership during major incidents, ensuring clear communication and coordination.
Responsibilities
Build and lead high-performing teams for infrastructure and service desk.
Own strategy and operations for multi-site connectivity and compute/storage.
Define standard technology patterns for automated deployments.
Treat internal infrastructure capabilities as product offerings for engineering.
Partner with clients to ensure stable systems and operational coordination.
Lead ITIL-aligned processes and support models for service delivery.
Support new client launches and ensure reliable site transitions.
Benefits
Comprehensive health, dental, and vision insurance.
Flexible work schedule and remote working opportunities.
Professional development and training support.
Retirement savings plan with company match.
Generous paid time off and holiday schedule.
Full Job Description
What you'll do:
Team leadership: Build and develop high-performing teams across infrastructure, systems and service desk; define on-call models and escalation paths.
Large-scale, multi-site infrastructure leadership: Own strategy, lifecycle, and operations for connectivity (LAN/WAN, Wi Fi), compute/storage, endpoint and mobile device management, and collaboration platforms.
Standardization & repeatability: Define standard technology patterns (network designs, device builds, security baselines, and spares) and enable repeatable, automated deployments through documented templates, "golden paths," and self-service provisioning (where appropriate) so development teams can move faster with fewer handoffs and less variability.
Infrastructure as a Product (IaaP) / Internal Developer Platform: Treat core internal infrastructure capabilities (identity, networking, compute/runtime platforms, endpoint standards, monitoring/observability, secrets/certificates, and deployment standards) as product offerings for internal engineering teams. Deliver developer self-service through a clear service catalog and portal experience (e.g., request/provision workflows, standardized templates, paved-road "golden paths," and reusable modules), integrated with engineering toolchains where appropriate (e.g., CI/CD pipelines, ticketing/ITSM, and approvals). Establish published roadmaps; define SLAs/SLOs and platform KPIs (time-to-provision, change lead time enablement, developer satisfaction); and continuously improve usability, documentation, and support. Expand automation and infrastructure-as-code practices by standardizing version-controlled blueprints (e.g., Terraform modules), GitOps based change workflows, continuous delivery practices, automated environment creation, policy-as-code guardrails, automated testing/validation of infrastructure changes, and repeatable rollback patterns-reducing manual work while maintaining security and compliance requirements.
Client venue enablement: Partner with clients to ensure stable connectivity, device readiness, integrations, and support processes; manage escalation paths and operational coordination to protect service during peak meal periods.
Custom back-office systems (operational partnership): Partner with internal software teams to operationalize custom back-office applications (including Inventory and Menu Builder); ensure environments are monitored, deployments are repeatable, and incidents are triaged effectively.
Service management, field support & SLAs: Lead ITIL-aligned processes and support operating models (service desk, escalations) with a focus on fast restoration and meeting SLAs/OLAs.
Location onboarding & rollouts: Support new client launches, site transitions, and rollouts; ensure repeatability and scalability, readiness checklists, connectivity validation, device provisioning, and cutover support are executed reliably.
Availability, resiliency & DR: Define and test disaster recovery and continuity plans for critical services (network, identity, collaboration, custom field systems, and key integrations); establish recovery objectives and improve resilience.
Operational excellence & observability: Implement monitoring/alerting for networks, endpoints, and key platforms; standardize runbooks; and expand automation for patching, provisioning, and routine maintenance at scale. Operationalize IaC by implementing drift detection, controlled promotion of changes through environments, automated compliance checks, and (where appropriate) auto-remediation for common issues. Track platform telemetry that matters to engineering stakeholders (e.g., provisioning latency, change failure rate, and incident impact) to drive continuous improvement.
Security, privacy & compliance: Partner with Information Security to meet PCI DSS (as applicable), FERPA-aligned expectations at education sites, and corporate security requirements; drive vulnerability remediation, endpoint hardening, identity controls, and logging while minimizing operational disruption.
Release & change governance: Lead change/release governance for infrastructure and field systems; ensure change windows, communication, rollback plans, and site readiness are in place.
Vendor & contract management: Manage MSPs, carriers/ISPs, and hardware vendors; oversee SLAs/SOWs, dispatch processes, spares strategy, and performance reviews; coordinate with vendor partners on support outcomes.
Financial management: Own cost optimization; manage lifecycle refresh and standard configurations to reduce support cost and downtime.
What you'll need:
Bachelor's degree in Information Technology, Computer Science, or related field (or equivalent practical experience).
Demonstrated experience supporting large, distributed, customer-facing operations (100 to 300 locations preferred) with time-sensitive service windows and high availability expectations.
Experience operating and supporting business-critical applications in production in partnership with software engineering teams.
Experience with cloud platforms and operating models (e.g., Azure/AWS), including governance and cost management.
Experience implementing and managing IT service management practices and tooling (e.g., ServiceNow or similar), including incident/problem/change disciplines and SLA performance.
Proven ability to lead through major incidents with clear communications, effective vendor coordination, and structured root-cause analysis.
Experience in contract dining or supporting technology (preferred).
Experience supporting site launches, transitions, acquisitions, or large-scale rollouts (preferred).
Experience with modern operations practices for software (e.g., SRE/DevOps concepts, CI/CD change controls, observability) in partnership with engineering teams (preferred).
Experience with PCI DSS programs, payment ecosystems, and supporting audits/assessments (preferred).
Experience building internal developer platforms and self-service tooling (e.g., developer portals, service catalogs, golden paths/templates) and operating IaC/GitOps practices with policy-as-code guardrails (preferred).
Experience with automation/infrastructure-as-code (e.g., Terraform, Ansible, scripting) (preferred).