Summary:
Trading and research at the firm run around the clock and across the globe, and they run on infrastructure this team designs, builds, and operates.
As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes.
The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.
Responsibilities:
- Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
- Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
- Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
- Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
- Own infrastructure projects end to end, from scope and design through implementation and long-term support.
- Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.
Qualifications:
- 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
- Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.
- Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.
- Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.
- Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.
- Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
- Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.
- Clear communication. You will work daily with researchers, engineers, and vendors.
Nice to Have:
- Experience with the rest of the NVIDIA stack, such as NCCL and NVLink.
Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.
Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.
Our benefits include:
- Generous paid time off policies
- Savings plans and other financial wellness tools available in each region
- Hybrid working opportunities
- Free breakfast, lunch, and snacks daily
- In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
- Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
- Volunteer opportunities and charitable giving
- Social events, happy hours, treats, and celebrations throughout the year
- Workshops and continuous learning opportunities