About the Role
In this role you will owns our Linux KVM (Kernel-based Virtual Machine) hypervisor platform and the automation used to provision, patch, monitor, and secure it. The successful candidate will define and maintain golden configurations, build the bare-metal provisioning pipeline, and operate the monitoring and log collection stack used across the team.
What You'll Do
- Develop and maintain hypervisor golden configurations: host build standards, kernel and performance tuning, and VM (Virtual Machine) templates, including connectivity testing.
- Manage the hypervisor lifecycle: image builds, patching and update schedules, staged rollouts, and rollback procedures.
- Build and maintain automated bare-metal provisioning: network boot, unattended installation, per-host configuration, and local package repositories.
- Build and operate the monitoring platform: Prometheus metrics collection covering the operating system, hypervisor, and hardware layers, with dashboards and alerting.
- Build and operate centralized log collection for the hypervisor fleet.
- Implement configuration management and drift detection across the hypervisor fleet.
- Maintain the security baseline: CIS (Center for Internet Security) hardening, automated compliance scanning, vulnerability triage and patching, and centralized identity and access management.
- Maintain CI (Continuous Integration) and testing for infrastructure code.
What You'll Need
- Bachelor's or Master's degree in Electrical Engineering, Telecommunications, Computer Science, or a related field.
- 10+ years expert-level Linux administration experience, with production experience on Red Hat Enterprise Linux or its derivatives (Rocky Linux, AlmaLinux, CentOS Stream).
- 5+ years of significant production experience operating Linux KVM/libvirt at fleet scale, including performance-sensitive workloads.
- 10+ years of experience building an automated provisioning pipeline from bare metal to running workload.
- 5+ years of experience with demonstrated host performance tuning: CPU pinning and isolation, NUMA (Non-Uniform Memory Access) memory placement, and interrupt affinity. (Replace Linux with Kubia)
- 5+ years of experience deploying and operating Prometheus and centralized logging.
- 5+ years of experience patching production fleets across multiple sites with minimal service impact.
Preferred Qualifications
- RHCE (Red Hat Certified Engineer) or RHCA (Red Hat Certified Architect) certification is highly desirable.
- Familiarity with SR-IOV (Single Root I/O Virtualization) or PCI passthrough for high-throughput virtual network functions.
- Familiarity with Userspace or kernel-bypass packet processing.
- Experience with immutable or API-managed host operating systems.
- Automation driven from an inventory/IPAM (IP Address Management) system such as NetBox.
- Security hardening in performance-sensitive environments.
What We'll Offer
- Connectivity and content engineering and business knowledge
- Opportunity to work in cross-functional teams
- Competitive Salary
- 401(k) with 100% match up to 4.0% after one year of service
- Competitive health benefits, including optical and dental
- Life Insurance
- Performance Bonus
- Flexible paid time off policy
- Monthly Phone Allowance