Tower Research Capital, LLC

Research Platform Engineer

Tower Research Capital, LLC • $200K — $300K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in HPC job scheduling and administration
  • Advanced knowledge of Kubernetes architecture and GPU-native schedulers
  • Familiarity with distributed ML frameworks like Ray and deep learning tools such as PyTorch
  • Hands-on experience with job graph orchestration frameworks
  • Experience in designing vendor-agnostic infrastructure and cloud-bursting strategies
  • Knowledge of high-performance storage solutions and network fabrics
  • Proven ability to build telemetry and cost-attribution platforms for multi-tenant environments

Responsibilities

  • Design intuitive platform abstractions and APIs for seamless researcher experience
  • Build a multi-tenant compute substrate with HPC-grade scheduling and strict isolation
  • Establish infrastructure-agnostic execution abstractions for multi-cloud portability
  • Evaluate and integrate job graph orchestration engines for automated pipelines
  • Implement automated failure detection and high-performance checkpointing
  • Profile and eliminate I/O bottlenecks for distributed ML workloads
  • Implement telemetry for tracking compute utilization and cost transparency

Benefits

  • Generous paid time off policies
  • Savings plans and financial wellness tools
  • Hybrid working opportunities
  • Free daily meals and snacks
  • In-office wellness experiences and reimbursement for wellness expenses
  • Company-sponsored sports teams and fitness events
  • Volunteer opportunities and charitable giving
  • Social events and celebrations throughout the year
  • Workshops and continuous learning opportunities
Full Job Description
Summary

As part of Tower Research's Core Engineering team, you will develop a multi-tenant research compute platform capable of dynamically orchestrating large scale ML workloads across a hybrid compute infrastructure of GPUs and CPUs. Your primary mission is to work closely with Quant Researchers, Portfolio Managers, and Infrastructure teams, to build a unified, elastic, and multi-tenant research substrate.
Key Responsibilities
  • Elevating Researcher Experience: Design intuitive platform abstractions, APIs, scheduler wrappers, and a durable job control plane so researchers can seamlessly launch simulations, distributed training runs, and complex research pipelines without incurring infrastructure overhead
  • Architecting Multi-Tenant Scheduling & Isolation: Build a multi-tenant compute substrate combining HPC-grade batch scheduling (gang scheduling, topology awareness, fair share) with cloud-native flexibility, enforcing strict tenant isolation across trading desks
  • Enabling Multi-Datacenter & Multi-Cloud Compute Portability: Establish infrastructure agnostic execution abstractions that enable compute workloads to run transparently across multiple on-premise datacenters or burst into external cloud providers
  • Standardizing Workflow Orchestration: Evaluate, select, and integrate production-grade job graph orchestration engines to automate multi-stage feature, training, and backtesting pipelines
  • Building Fault-Tolerant Research Pipelines: Implement automated failure detection, retry-on-fault mechanisms, and high-performance checkpointing to ensure long-running distributed jobs survive hardware degradation and faults without manual intervention
  • Optimizing Data-Paths: Profile and eliminate I/O bottlenecks, ensuring distributed ML workloads align with high-speed network fabrics and high-performance storage
  • Delivering Observability & Cost Transparency: Implement comprehensive telemetry to track compute utilization, queue pressure, and GPU/CPU cost attribution, giving Portfolio Managers and Senior Management clear visibility into resource consumption and ROI
Technical Requirements
  • HPC Administration: Deep experience with HPC job schedulers (e.g., Slurm), including gang scheduling, fair-share priority trees, topology-aware node allocation, and containerized HPC execution
  • Kubernetes-native Engineering: Advanced understanding of Kubernetes architecture, CRDs, HPC focused operators, admission controllers and GPU-native schedulers
  • Distributed ML Computing Frameworks: Familiarity with distributed computing frameworks (e.g., Ray) and deep learning frameworks (e.g., PyTorch, JAX) with multi-node scaling primitives
  • Workflow Orchestration Expertise: Hands-on experience evaluating, architecting, and operating job graph orchestration frameworks
  • Multi-Platform Architecture: Experience designing vendor-agnostic infrastructure layers, cloud-bursting strategies, and compute execution spanning multiple on-premise datacenters and public cloud environments
  • Storage & Network Performance: Familiarity with high-performance storage solutions and high-speed network fabrics for large scale research workloads
  • Compute Observability: Proven track record building cluster-wide telemetry and cost-attribution platforms for multi-tenant environments
  • Architectural Leadership Mindset: A strong platform engineering mindset focused on reducing friction for researchers while maintaining rigorous operational efficiency, cost transparency, and system scalability

Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus

Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.

Our benefits include:
  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

About Tower Research Capital, LLC

Tower Research Capital, LLC is a quantitative trading firm that was founded in 1998. The company uses advanced technology and algorithms to trade in multiple asset classes across global markets. Tower Research Capital, LLC is headquartered in New York City and has offices in North America, Europe, and Asia. The company is known for its innovative approach to trading and its use of cutting-edge technology to analyze market data and make trading decisions. Tower Research Capital, LLC is a privately held company and does not disclose its financial information to the public.
Learn more about Tower Research Capital, LLC
Size
1,000 employees
Industry
Founded
1998

Similar Jobs

More Information Technology Jobs

Find similar Research Platform Engineer jobs: