Full Job Description
Position Description:
Virtualization Hosting service enablement • Responsible for engineering, deployment, operational administration, and protection of Global enterprise solutions to meet the Virtualization Server Hosting requirements for a variety of infrastructure systems, line of business and third-party application needs • Architect/design and support the installation, and administration of virtualization (VMware and OpenShift Virtualization (OSV)) suite of products • Responsible for the entire lifecycle of technologies globally including infrastructure security vulnerability patching, planning, designing, implementation, maintenance, upgrades and decommissioning of hardware and software • Engineer, test and document procedures, monitoring, logging, disaster recovery process and security policies and guidelines • Demonstrates knowledge of hardware and software products • Research industry best practices and trends • Global large-scale deployments of virtualization technologies • Support HPE Synergy/ProLiant ILO and firmware field testing Capacity Management • Conduct capacity planning and forecasting for the platforms, including Compute/Virtual Machine (VM), memory, storage, and network resources, to ensure scalability and prevent resource exhaustion • Analyze resource utilization trends and make recommendations for infrastructure scaling, consolidation, or optimization • Collaborate with application teams and stakeholders to understand future demand and project capacity needs • Develop and maintain capacity models and reports to support strategic planning Automation & Efficiency • Develop automation solutions (scripts, playbooks) for repetitive VMware/OSV tasks, including configuration changes, VM management (like snapshot removal), auditing, remediation and integration with ticketing systems • Leverage automation to enable delivering operator updates and changes efficiently at scale • Implement Site Reliability Engineering (SRE) principles and practices to improve overall platform stability, performance, and operational efficiency • Role Based Access Control deployment and auditing • Namespace and Resource Quota management (CPU, Disk and Storage) Observability, Monitoring, logging and Troubleshooting • Implement and maintain comprehensive end to end observability solutions (monitoring, logging, tracing) for the VMware/OSV environment, including integration with tools like Dynatrace, RHACM and Prometheus/Grafana • Explore and implement Event Driven Architecture (EDA) for enhanced real time monitoring and response • Develop capabilities to flag and report abnormalities and identify "blind spots " in observability • Perform deep dive Root Cause Analysis (RCA), potentially utilizing available tooling, to quickly identify and resolve issues across the global compute environment • Find the needle in a haystack/unhealthy bits in the compute universe (Globally) for faster time to resolution • Monitor VM health, resource usage, and performance metrics proactively • Monitor for unusual activity that might indicate a compromise or misconfiguration Solution Design & Consulting • Provide technical consulting and expertise to application teams requiring VMware/OSV solutions • Design, implement, and validate custom or dedicated OSV clusters and VM solutions for critical applications with unique or complex requirements (e.g., specialized appliances) Knowledge Management • Create, maintain, and update comprehensive internal documentation and customer facing content to facilitate self service and clearly articulate platform capabilities Support • Participate in L1 - L3 level support to Operations teams environmental related issues. Monthly after hours and weekend work will be required
Skills Required:
Scripting, Automation, Kubernetes, Root Cause Analysis, Troubleshooting (Problem Solving), Cloud Architecture, IT Solutions, GitHub, Cloud Infrastructure, Change Management, Technical Analysis, Developer, Tekton, Utilization Management, VMware, VMware ESX Servers, Platform Support, Infrastructure Architectures
Skills Preferred:
Ansible, GCP, Dynatrace, Powershell, Access Controls, Python, Information Security, Automation, Artificial Intelligence & Expert Systems
Experience Required:
Senior Engineer Exp: Prac. In 2 coding lang. or adv. Prac. in 1 lang.; guides. 10+ years in IT; 8+ years in development Understanding of VMware and Kubernetes concepts Experience with Linux administration and networking fundamentals Proficiency in scripting languages for automation Experience with monitoring tools and logging solutions Understanding of virtualization concepts and technologies (e.g., KVM, VMware) Excellent problem-solving skills and the ability to troubleshoot complex issues across multiple layers of the stack Knowledge of CI/CD pipelines and DevOps methodologies Strong communication and collaboration skills Self-starter. Be on a mission to go where the work is. Look for opportunities to evolve services
Education Required:
Associate Degree, College Senior
Education Preferred:
Certification Program, Bachelor's Degree
Additional Information:
4 days on site