Tower Research Capital, LLC

HPC Operations Engineer

Tower Research Capital, LLC$175K — $225K *
Technical Services
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science, engineering, or related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid fundamentals in Linux administration (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting approach to problem-solving.
  • Experience in direct technical support or operations roles.
  • Strong written communication skills for clear documentation and handoffs.
  • Detail-oriented with a disciplined approach to following established processes.

Responsibilities

  • Provide first-line support for HPC users across various issues.
  • Troubleshoot job failures and scheduling errors to resolution.
  • Triage infrastructure incidents, gathering diagnostics and escalating as needed.
  • Monitor fleet health and proactively address potential issues.
  • Execute operational procedures for maintenance and configuration updates.
  • Provision new machines into the Research fleet and handle lifecycle management.
  • Maintain detailed runbooks and knowledge base articles for efficiency.
  • Identify recurring issues and suggest workflow improvements or automation.

Benefits

  • Generous paid time off policies.
  • Savings plans and financial wellness tools tailored to each region.
  • Hybrid working opportunities for work-life balance.
  • Complimentary meals and snacks provided daily.
  • Wellness experiences and reimbursement for select wellness expenses.
  • Company-sponsored sports teams and fitness events.
  • Volunteer opportunities and charitable giving initiatives.
  • Social events and team celebrations throughout the year.
  • Workshops and continuous learning opportunities for career growth.
Full Job Description
Summary:

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day.

The fleet has grown fast, with close to a thousand machines added recently, and this role exists so that infrastructure gets dedicated, high-standard operational care. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation.

You will sit inside the HPC team, next to the engineers who build and run the platform. That proximity matters: your diagnostics feed their root-cause work, your runbooks capture what the team learns, and the recurring issues you surface become candidates for automation and permanent fixes. For someone who wants to grow into HPC engineering, this is a strong place to start.

Responsibilities:
  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and escalate to subject-matter experts when a problem extends beyond defined ownership.
  • Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.
  • Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.
  • Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming back.

Qualifications:
  • A bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux administration fundamentals (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting: you work a problem step by step, know what you have ruled out, and recognize when it is time to escalate.
  • Experience working directly with users in a technical support or operations role.
  • Strong written communication: clear tickets, clear runbooks, clear handoffs.
  • The discipline to follow established processes with genuine attention to detail.


Nice to Have:

  • Enough Bash or Python to script away routine operational tasks.
  • Working knowledge of batch schedulers such as Slurm, HTCondor, or LSF.
  • A good grasp of the plumbing behind networked computing: NFS, automounter, LDAP.
  • Hands-on experience provisioning Linux machines: network boot (PXE), unattended installs (kickstart), or configuration management such as Ansible.
  • Prior exposure to HPC or other large-scale compute environments.
  • Familiarity with monitoring and observability stacks such as Prometheus and Grafana.

Anticipated annual base salary range $175,000-$225,000, plus eligible for discretionary bonus.

Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.

Our benefits include:
  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

About Tower Research Capital, LLC

Tower Research Capital, LLC is a quantitative trading firm that was founded in 1998. The company uses advanced technology and algorithms to trade in multiple asset classes across global markets. Tower Research Capital, LLC is headquartered in New York City and has offices in North America, Europe, and Asia. The company is known for its innovative approach to trading and its use of cutting-edge technology to analyze market data and make trading decisions. Tower Research Capital, LLC is a privately held company and does not disclose its financial information to the public.
Learn more about Tower Research Capital, LLC
Size
1,000 employees
Industry
Founded
1998

Similar Jobs

More Jobs at Tower Research Capital, LLC

More Technical Services Jobs

Find similar HPC Operations Engineer jobs: