Site Reliability Engineer - Low-Latency Trading Systems

Ondo Finance

$150K — $180K *
US-AnywhereRemote in United States
Finance & Insurance
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in SRE or production engineering roles focusing on real-time systems
  • Strong programming skills in Go or Rust, with application code modification experience
  • Deep experience managing stateful, latency-sensitive workloads in Kubernetes and AWS
  • Expertise in observability tools such as Prometheus and structured log analysis
  • Solid understanding of Linux internals and network fundamentals
  • Ability to maintain composure and communicate effectively during production incidents

Responsibilities

  • Debug production incidents end to end, including stale data feeds and latency issues
  • Enhance market data ingestion from various providers while ensuring robust failover and replay mechanisms
  • Develop tooling for reconciliation and data integrity across multiple data sources and formats
  • Participate in an on-call rotation for market hours and 24/7 crypto trading

Benefits

  • Competitive compensation package including salary, token rights, and equity options
  • Comprehensive medical, vision, and dental benefits with flexible PTO
  • Remote-first work environment with a globally distributed team
  • Opportunity to join a team with top-tier colleagues from renowned companies
  • Backing from leading investors in the crypto space
Full Job Description
About the role

Ondo operates real-time trading systems that run around the clock across traditional and crypto venues. The platform spans low-latency Rust engines, a fleet of Go services for trading, execution, and PnL accounting, and a multi-region Kubernetes footprint on AWS.

We are looking for an SRE with strong systems programming skills to own the reliability, observability, and performance of this platform. This is a hands-on role: you will read and modify Go and Rust code, debug latency regressions down to the feed handler, run incident response during market hours, and build the automation that keeps a 24/7 trading system healthy with a small team.

Target outcomes
  • Own production reliability for real-time trading services: trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines
  • Operate and evolve our multi-region Kubernetes clusters on AWS (EKS), deployed via GitOps (Flux) with SOPS-encrypted secrets
  • Build and refine observability: Prometheus metrics and alerting, Datadog logs and dashboards, and the SLOs that catch degradation before it costs money
  • Improve deploy safety: progressive rollouts, config-reload behavior, and guardrails that prevent a bad push from touching live trading

Responsibilities
  • Debug production incidents end to end: stale market data feeds, exchange rate limits, WebSocket disconnects, order-lifecycle desyncs, and latency regressions in the trading path
  • Harden market data ingestion from providers such as Databento and venue-native feeds (REST and WebSocket), including staleness detection, failover, and replay
  • Build reconciliation and data-integrity tooling across live gauges, Postgres, and our S3 parquet data lake, so positions, fills, and PnL always agree
  • Participate in an on-call rotation covering US equity market hours and 24/7 crypto venues

Requirements
  • 5+ years in SRE, production engineering, or infrastructure roles, with meaningful time supporting real-time or latency-sensitive systems
  • Strong programming ability in Go or Rust, and willingness to work in both; this role changes application code, not just infrastructure
  • Deep, hands-on Kubernetes and AWS experience: you have run stateful, latency-sensitive workloads in production, not just stateless web services
  • Strong observability instincts: fluent PromQL, structured-log analysis, and experience designing alerts with high signal and low noise
  • Solid Linux internals and networking fundamentals: you can chase a p99 regression through the kernel, the NIC, or the GC
  • Sound judgment under pressure and clear written communication during and after incidents

Nice to haves
  • Experience operating trading systems, execution infrastructure, or market data infrastructure at a trading firm, exchange, or broker
  • Familiarity with market microstructure and order lifecycle (order books, order types, fills and reconciliation)
  • Experience with market data providers and protocols (Databento, SIP/prop equity feeds, venue WebSocket APIs)
  • Exposure to crypto venues and on-chain trading
  • Python for operational tooling and data analysis (pandas, parquet, BigQuery)
  • Experience with GitOps workflows, infrastructure as code, and secrets management at scale

Tech stack

Go, Rust, Python | Kubernetes (EKS), Flux, SOPS | AWS (multi-region), S3 parquet lake | Prometheus, Grafana, Datadog | Postgres, CockroachDB, BigQuery | Databento, venue WebSocket/REST feeds

What we offer
  • Competitive compensation including but not limited to salary, future token rights, and/or equity (according to your preferences) - We are well-funded and believe that great talent deserves great compensation.
  • Full benefits (medical, vision, and dental) and flexible vacation policy (PTO).
  • Remote-first team across many countries - You will be an early team member helping shape our vision, culture, and design practices.
  • A+ colleagues - Our team includes alumni from: Goldman Sachs, Blackrock, Two Sigma, Bridgewater, SpaceX, AWS, Meta, Google, McKinsey, Coinbase, Circle, Uniswap.
  • Best-in-class investors - We are proud to be backed by leading crypto experts and VCs, including Pantera Capital, Founders Fund and Coinbase Ventures.

Similar Jobs

More Jobs at Ondo Finance

More Finance & Insurance Jobs

Find similar Site Reliability Engineer - Low-Latency Trading Systems jobs: