What you will do:- Design, deploy, automate, and operate the compute, storage, licensing, and CI/CD platforms that support High-Performance Computing (HPC) and Electronic Design Automation (EDA) workloads across on-premises, AWS, and Azure environments.
- Build, maintain, and improve Jenkins or equivalent CI/CD pipelines for EDA and HPC engineering workflows, including automated software deployment, environment configuration, validation, and regression testing.
- Standardize EDA tool packaging, deployment, configuration, and version management, and establish repeatable processes for promoting tool and infrastructure changes across development and production environments.
- Manage Red Hat Enterprise Linux and Ubuntu systems and administer virtualized infrastructure based on Xen, VMware, KVM, or similar technologies.
- Architect and operate secure hybrid-cloud infrastructure, including VPCs, AWS Direct Connect, VPN connectivity, auto-scaling, and cloud bursting.
- Administer workload schedulers such as Slurm, Platform LSF, and SGE; automate job submission, resource allocation, and environment configuration; and improve workload throughput, availability, and utilization.
- Manage NAS platforms, including TrueNAS, NetApp ONTAP, and AWS FSx for NetApp ONTAP, and optimize storage performance and reliability for high-throughput HPC and EDA workloads.
- Deploy and maintain FlexLM and other vendor license servers, ensuring license availability, monitoring, compliance, and operational resilience.
- Administer LDAP, Active Directory, SSO, and related identity systems, including domain permissions, role-based access, authentication, and audit logging.
- Implement monitoring and reporting for compute, storage, schedulers, CI/CD pipelines, and license utilization; diagnose performance and capacity issues; and drive cost, reliability, and performance improvements.
What you bring to this role:- Demonstrated experience designing, automating, and operating reliable infrastructure for semiconductor EDA, HPC, or similarly compute-intensive engineering environments.
- Strong experience building and operating CI/CD pipelines using Jenkins or an equivalent platform, including automated deployment, configuration, validation, and operational workflows.
- Strong administration skills with Red Hat Enterprise Linux and Ubuntu, including automation, troubleshooting, performance analysis, and production operations.
- Practical experience with HPC job schedulers such as Slurm, Platform LSF, or SGE, including workload management, resource allocation, and scheduler integration.
- Experience deploying and operating hybrid environments across on-premises infrastructure and AWS or Azure, including secure networking, auto-scaling, or cloud bursting.
- Experience managing NAS and high-performance storage platforms, including TrueNAS, NetApp ONTAP, or AWS FSx for NetApp ONTAP.
- Experience deploying and administering FlexLM or other software-license management systems, with responsibility for availability, monitoring, and compliance.
- Experience with virtualization technologies such as Xen, VMware, or KVM.
- Experience administering LDAP, Active Directory, SSO, access controls, and security or audit logging in an engineering environment.
- Ability to work across software, IT, security, and engineering teams; document technical decisions and operating procedures; and lead infrastructure improvements from design through production operation.
- Typical background: 8+ years of experience in DevOps, infrastructure engineering, systems engineering, or a related discipline, including 5+ years supporting semiconductor EDA or HPC environments. Equivalent demonstrated expertise is also welcome; these are guidelines rather than mandatory minimums.
Bonus points for the following:- Experience with semiconductor design flows, EDA applications, and the compute, storage, and licensing requirements of chip-development teams.
- Experience automating EDA tool deployment, configuration, version management, validation, and regression workflows.
- Experience implementing hybrid-cloud bursting and elastic HPC orchestration.
- Experience deploying or operating AWS ParallelCluster.
- Experience deploying or administering MATLAB Parallel Server, including integration with HPC schedulers and compute clusters.
- Experience integrating Linux-based engineering infrastructure with Windows workstation environments.
- Experience improving platform reliability through monitoring, capacity planning, and continuous operational improvement.
Additional RequirementsThis is a fast-paced, high-impact environment - flexibility to occasionally work extended hours or weekends during critical periods is expected.
$140,000 - $200,000 a year
This is a full time, exempt position, based out of our Santa Clara office. The target base pay for this position is $140,000 - $200,000 annually. The total compensation packaged will be determined by various factors such as your relevant job-related knowledge, skills, and experience.
We are redefining how satellites are designed, manufactured and used-so we're looking for candidates with passion, deep knowledge and direct experience on LEO satellite component development, design and in-orbit activities. If that's your experience - then we'll be immediately wow-ed.
E-Space is not currently able to provide employment sponsorship for candidates who do not hold work authorization for the location of this role.