JOB SUMMARY
Client is seeking a Senior DevOps Engineer to provide a 12-month contingent engagement supporting our High Performance Computing (HPC) and Electronic Design Automation (EDA) cloud infrastructure team. This role will work directly within the IT Datacenter (ITDC) organization and is expected to operate independently at a senior level with minimal ramp-up time. The ideal candidate brings strong hands-on experience with Linux HPC environments, infrastructure automation, SLURM workload management, and the Azure cloud environment.
Key Responsibilities
HPC / EDA Platform Operations
Support and administer SLURM-based HPC compute environments, including partition configuration and migration planning using Terraform and Ansible
Author formal Method of Procedure (MOP) documents and runbooks for infrastructure changes and service cutovers
Coordinate cross-functionally with EDA/TD NAND teams, storage teams, and IDAM to deliver coordinated platform changes
Administer Azure EDA user environment utilizing Thinlinc (VNC)
Automation & Infrastructure as Code
Develop, maintain, and extend Ansible playbooks and roles for Linux system setup, authentication, and platform configuration
Ensure multi-version Ansible playbook compatibility across SLES 15
Contribute GitHub pull requests, conduct code reviews, and manage inner-source infrastructure repositories
Drive production environment changes through change management workflows using ServiceNow
Identity & Access Management
Integrate and configure enterprise identity systems including Okta, Active Directory, LDAP, and SSSD for Linux/HPC environments
Audit and reconcile Linux user and group identity data (UID/GID) across multiple directory and authentication domains
Validate authentication methods and access behavior across HPC compute and storage environments
Extend SSSD-based corporate authentication to new compute environments and author corresponding Ansible automation
Monitoring, Logging & Operational Readiness
Assess and implement log management strategies, including evaluation of Splunk integration for HPC system logs
Investigate and remediate operational issues in production Linux services (VNC, AutoFS, Datadog, etc.)
Produce technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence
Required Qualifications
Core Technical Skills
Category
Skills / Tools
HPC / EDA Platforms
SLURM, HPC compute/storage administration, EDA infrastructure, datacenter migrations
Linux / OS
SLES 15, Ubuntu
Provisioning / Automation
Ansible (playbooks, roles, multi-version)
Identity / Auth
SSSD, Okta, Active Directory, LDAP, cross-domain identity management
Storage / Filesystems
NetApp SVM, NFS, AutoFS, RootSquash, storage tier design, IOPS/capacity planning
DevOps / Source Control
Git, GitHub, Artifactory, inner-source repository management
Monitoring / Logging
Splunk integration, Datadog, operational script hardening, log management
Scripting / Languages
Ansible (YAML), Python, Perl (debugging), Bash
ITSM / Documentation
ServiceNow (change requests), MOP authoring, Confluence, Jira, technical diagramming
Experience Requirements
5+ years of experience in a DevOps, Platform Engineering, or Linux Systems Engineering role
Hands-on HPC cluster administration experience, including SLURM or equivalent workload managers
Demonstrated experience supporting EDA or scientific computing environments
Strong Ansible & Terraform automation skills with production-grade playbook and role development
Demonstrated usage and understanding of the Azure cloud compute environment.
Familiarity with enterprise Linux identity and authentication stacks (SSSD, LDAP, AD, Okta)
Experience with NetApp or comparable enterprise storage platforms in HPC contexts
Ability to author formal technical documentation (MOPs, runbooks, architecture diagrams)
Strong written and verbal communication skills; capable of coordinating across multiple teams
Terraform and Ansible
Preferred Qualifications
Experience with SUSE Linux Enterprise Server (SLES) 12 and/or 15 in an enterprise environment
Experience migrating configuration artifacts and binaries to Artifactory
Background in semiconductor, storage, or high-tech manufacturing IT environments