Job DescriptionWe're looking for a Research Computing Systems Engineer to help operate and grow the high-performance computing (HPC) environment that powers research across every discipline at BU. In this role, you will keep our production HPC cluster running reliably, help deploy and operate a new NIST 800-171-compliant secure research computing environment built on tiCrypt, and modernize the way we monitor, automate, and deliver services. This is a full-stack, operations- and automation-focused position for an engineer who enjoys working on systems end to end - from writing infrastructure as code, to troubleshooting hardware, to planning physical rack deployments. You'll join a small, deeply technical team where everyone stays hands-on and has a real voice in shaping how we operate our infrastructure. You'll work closely with senior engineers, collaborate with other HPC professionals, support researchers, and also work independently. This position is partially remote.
You Will:
- Operate, maintain, and optimize the University's production HPC cluster, including a 1,000+ node bare-metal Linux deployment, the Grid Engine (SGE) scheduler, and GPFS (IBM Storage Scale) parallel file system, to keep research computing services reliable and performant.
- Contribute to the build, deployment, and operations of a new NIST 800-171-compliant secure research computing environment built on tiCrypt, including a Slurm workload manager.
- Help improve monitoring, observability, and configuration management using CI/CD pipelines, Ansible, Git, and related automation tools.
- Automate provisioning, deployment, and routine operations across the full stack, from bare metal through services.
- Identify, diagnose, and resolve complex system, storage, network, and performance issues.
- Contribute to technical documentation and share your expertise with colleagues and the research community.
Required Skills- 3+ years of relevant work experience with a bachelor's degree, relevant post-secondary education, or a combination of these two.
- Linux systems administration experience, including OS patching, upgrades, and security maintenance.
- Experience programming in a Linux environment, preferably Bash and Python.
- Experience in multiple languages such as Perl or C is a plus.
- Ability to work with configuration management systems (e.g., Ansible, Puppet, Chef) and revision control systems (e.g., Git).
- Hands-on experience and understanding of Linux virtualization and containerization technologies (e.g., KVM, OpenStack, OpenNebula, Docker, Kubernetes).
- Strong communication, problem-solving, and teamwork skills.
Boston University offers an excellent benefits package including:
- Time Off: In addition to PTO and leave policy, BU employees have a paid intersession break and 13 paid holidays.
- Retirement: University-funded retirement plan with full vesting after 2 years of eligible service.
- Tuition Assistance Program: Competitive tuition assistance program for yourself and family members.
- Check out https://www.bu.edu/wellness/ and https://www.bu.edu/hr/part-time-employee-perks/ for more information!
Boston University IS&T invests in our staff and their personal and professional growth. We promote staff learning including lunch and learn sessions, an extensive library of online courses, Fun Advisory Board (FAB) arranges a number of events throughout the year and opportunities to engage with peers at NERCOMP and EDUCAUSE events.