Tower Research Capital, LLC

Machine Leaning Performance Engineer (Inference)

Tower Research Capital, LLC$200K — $300K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years optimizing deep learning inference in high-throughput environments.
  • Expertise in lower-level ML frameworks (PyTorch/JAX) with strong Python/C++ skills.
  • Experience in custom GPU kernel development and optimization tooling (e.g., Triton, TensorRT).
  • Deep knowledge of GPU architecture and memory optimization.
  • Proven data-driven evaluation of inference performance across architectures.
  • Preferred experience targeting optimized workloads on FPGAs and ASICs.

Responsibilities

  • Lead technical evaluation of various inference platforms (CPUs, GPUs, FPGAs).
  • Analyze and optimize execution across memory hierarchies for resource maximization.
  • Collaborate with infrastructure teams on thermal and power constraints for latency-critical designs.
  • Develop and integrate optimized GPU kernels for maximum throughput.
  • Implement model reduction techniques for compact, low-latency inference.
  • Work closely with cross-functional teams to realize deployment goals.

Benefits

  • Generous paid time off policies.
  • Savings plans and financial wellness tools available.
  • Hybrid working opportunities.
  • Free meals and snacks daily.
  • In-office wellness experiences and reimbursement for wellness expenses.
  • Company-sponsored sports teams and fitness events.
  • Opportunities for volunteer work and charitable giving.
  • Social events and celebrations year-round.
  • Workshops and continuous learning opportunities.
Full Job Description
Summary:

As part of Tower Research's Core Engineering team, you will bridge the gap between quantitative research and high-performance production systems, architecting inference pipelines that operate at the physical limits of hardware. Your objective will be to drive the speed, efficiency, and reliability of our ML inference pipelines to their absolute limits, ensuring our predictive models consistently achieve microsecond-level latency.

Responsibilities:
  • Benchmarking & Strategy:
    • Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide Tower's infrastructure deployment decisions.
  • System Architecture Optimization:
    • Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle.
  • Infrastructure & Deployment Feasibility:
    • Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes.
  • GPU Kernel Development:
    • Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon.
  • Model Optimization & Deployment:
    • Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements.
  • Cross-Functional Collaboration:
    • Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments.

Qualifications:
  • 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain.
  • ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation.
  • Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute).
  • GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM).
  • Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures.
  • Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs.
  • Prior experience in financial trading is not required.


Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.

Tower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.

Our benefits include:
  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

About Tower Research Capital, LLC

Tower Research Capital, LLC is a quantitative trading firm that was founded in 1998. The company uses advanced technology and algorithms to trade in multiple asset classes across global markets. Tower Research Capital, LLC is headquartered in New York City and has offices in North America, Europe, and Asia. The company is known for its innovative approach to trading and its use of cutting-edge technology to analyze market data and make trading decisions. Tower Research Capital, LLC is a privately held company and does not disclose its financial information to the public.
Learn more about Tower Research Capital, LLC
Size
1,000 employees
Industry
Founded
1998

Similar Jobs

More Jobs at Tower Research Capital, LLC

More Enterprise Technology Jobs

Find similar Machine Leaning Performance Engineer (Inference) jobs: