SummaryAs part of Tower Research's Core Engineering team, you will develop a multi-tenant research compute platform capable of dynamically orchestrating large scale ML workloads across a hybrid compute infrastructure of GPUs and CPUs. Your primary mission is to work closely with Quant Researchers, Portfolio Managers, and Infrastructure teams, to build a unified, elastic, and multi-tenant research substrate.
Key Responsibilities- Elevating Researcher Experience: Design intuitive platform abstractions, APIs, scheduler wrappers, and a durable job control plane so researchers can seamlessly launch simulations, distributed training runs, and complex research pipelines without incurring infrastructure overhead
- Architecting Multi-Tenant Scheduling & Isolation: Build a multi-tenant compute substrate combining HPC-grade batch scheduling (gang scheduling, topology awareness, fair share) with cloud-native flexibility, enforcing strict tenant isolation across trading desks
- Enabling Multi-Datacenter & Multi-Cloud Compute Portability: Establish infrastructure agnostic execution abstractions that enable compute workloads to run transparently across multiple on-premise datacenters or burst into external cloud providers
- Standardizing Workflow Orchestration: Evaluate, select, and integrate production-grade job graph orchestration engines to automate multi-stage feature, training, and backtesting pipelines
- Building Fault-Tolerant Research Pipelines: Implement automated failure detection, retry-on-fault mechanisms, and high-performance checkpointing to ensure long-running distributed jobs survive hardware degradation and faults without manual intervention
- Optimizing Data-Paths: Profile and eliminate I/O bottlenecks, ensuring distributed ML workloads align with high-speed network fabrics and high-performance storage
- Delivering Observability & Cost Transparency: Implement comprehensive telemetry to track compute utilization, queue pressure, and GPU/CPU cost attribution, giving Portfolio Managers and Senior Management clear visibility into resource consumption and ROI
Technical Requirements- HPC Administration: Deep experience with HPC job schedulers (e.g., Slurm), including gang scheduling, fair-share priority trees, topology-aware node allocation, and containerized HPC execution
- Kubernetes-native Engineering: Advanced understanding of Kubernetes architecture, CRDs, HPC focused operators, admission controllers and GPU-native schedulers
- Distributed ML Computing Frameworks: Familiarity with distributed computing frameworks (e.g., Ray) and deep learning frameworks (e.g., PyTorch, JAX) with multi-node scaling primitives
- Workflow Orchestration Expertise: Hands-on experience evaluating, architecting, and operating job graph orchestration frameworks
- Multi-Platform Architecture: Experience designing vendor-agnostic infrastructure layers, cloud-bursting strategies, and compute execution spanning multiple on-premise datacenters and public cloud environments
- Storage & Network Performance: Familiarity with high-performance storage solutions and high-speed network fabrics for large scale research workloads
- Compute Observability: Proven track record building cluster-wide telemetry and cost-attribution platforms for multi-tenant environments
- Architectural Leadership Mindset: A strong platform engineering mindset focused on reducing friction for researchers while maintaining rigorous operational efficiency, cost transparency, and system scalability
Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus
Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.
Our benefits include:- Generous paid time off policies
- Savings plans and other financial wellness tools available in each region
- Hybrid working opportunities
- Free breakfast, lunch, and snacks daily
- In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
- Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
- Volunteer opportunities and charitable giving
- Social events, happy hours, treats, and celebrations throughout the year
- Workshops and continuous learning opportunities