Summary:
Our GPU fleet is one of the fastest-growing and most critical parts of the firm's infrastructure, and this role is about the software that keeps it healthy, observable, and efficient.
In this role you will build and own the tooling behind day-to-day GPU operations: management, monitoring, metrics collection, maintenance, and network configuration. When something misbehaves anywhere between an application and the kernel, you will help track it down.
The work varies day to day by design. One day you might be shipping a new feature for our fleet automation, the next digging into a stubborn driver issue, the next mining job statistics to find where the fleet could be doing more with what it has.
Responsibilities:
- Build and maintain the tooling that automates GPU fleet operations: management, monitoring, metrics collection, maintenance, and network configuration.
- Troubleshoot software and hardware issues across the fleet, from application and network problems down to the operating system and kernel.
- Work with engineering teams across the firm to tune workloads and processes so they use GPUs more efficiently.
- Analyze GPU job statistics to surface trends, inefficiencies, and opportunities for improvement.
Qualifications:
- BS or MS in computer science or a related field.
- 2+ years of relevant experience, including Python development and hands-on GPU management.
- An automation-first mindset. When a workflow is manual, slow, or error-prone, you reach for code.
- Experience deploying, troubleshooting, and tuning a range of GPU hardware.
- Strong computer science fundamentals and sound software design instincts.
- Solid working knowledge of Linux/UNIX and comfort with open-source software.
- Sharp debugging. You get to the bottom of problems quickly and methodically.
- Familiarity with configuration management and monitoring technologies.
- Organized, adaptable, and collaborative. You can juggle several tasks with careful attention to detail, work independently or with the team, and pick up new skills fast.
Nice to Have:
- Familiarity with continuous integration and deployment (CI/CD) tools and processes.
Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.
Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.
Our benefits include:
- Generous paid time off policies
- Savings plans and other financial wellness tools available in each region
- Hybrid working opportunities
- Free breakfast, lunch, and snacks daily
- In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
- Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
- Volunteer opportunities and charitable giving
- Social events, happy hours, treats, and celebrations throughout the year
- Workshops and continuous learning opportunities