Site Reliability Engineer

Sporttrade

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, or production engineering roles for 24/7 systems
  • Strong Linux fundamentals and scripting skills
  • Experience debugging Java applications in production
  • Solid TCP/IP networking knowledge to troubleshoot connectivity
  • Hands-on Kubernetes experience in production environments
  • Familiarity with observability tools like Datadog and Prometheus
  • Working knowledge of SQL and relational databases, preferably PostgreSQL

Responsibilities

  • Own daily operation of the exchange trading lifecycle
  • Lead incident response during on-call rotations
  • Operate and improve observability stack including dashboards and alerts
  • Manage hybrid infrastructure, including Kubernetes and cloud services
  • Automate infrastructure and operational procedures with tools like Terraform
  • Support data platform and disaster recovery processes
  • Contribute to process improvement and policy establishment

Benefits

  • 100% premium coverage for employee medical, dental, and vision insurance
  • Short- and long-term disability insurance
  • 401(k) retirement plan
  • Equity options available
  • Flexible time off policy
  • MacBook issued to all employees
Full Job Description
The Position

Sporttrade operates a regulated sports betting exchange that runs the way a financial exchange does. A successful Site Reliability Engineer at Sporttrade will be an operations-minded self-starter who treats the exchange like the production trading system it is: sessions open and close on schedule, every trade is captured and reported, and incidents are resolved quickly and documented thoroughly. In this role, you will own the day-to-day health of the exchange, spanning cloud infrastructure and on-premises datacenters, and you will automate away the toil that comes with operating a market that never wants to miss an open. This role works closely with the DevOps, Technical Operations, and Engineering teams to keep a live, regulated marketplace fast, observable, and compliant.

Duties
  • Own the daily operation of the exchange trading lifecycle - market startup and shutdown, enabling and disabling trading, pre- and post-session sanity checks, and capture of settlement, clearing, and trade-reporting artifacts.
  • Participate in an on-call rotation for a live regulated marketplace; lead incident response, drive incidents to resolution, author postmortems, and turn one-off fixes into runbooks and automation
  • Operate and improve our observability stack (Datadog, Wazuh, Prometheus, Grafana) - dashboards, alert quality, SLOs, and reducing time-to-detection for market-impacting issues
  • Run and maintain hybrid infrastructure: Kubernetes clusters (GKE and EKS) with an Istio service mesh, AWS and GCP accounts, and exchange servers in geographically distributed on-premises datacenters
  • Automate infrastructure and operational procedures with Ansible, Terraform, and Jenkins pipelines, with secrets managed in HashiCorp Vault
  • Support the data platform behind the exchange: PostgreSQL (Cloud SQL), Kafka (Confluent Cloud) change-data-capture and streaming pipelines, Redis, and backup/restore and disaster-recovery procedures - including proving that backups actually restore
  • Support market maker and partner connectivity (site-to-site VPNs and datacenter cross-connects), as well as conformance testing and onboarding support for partners joining the exchange.
  • Contribute to ongoing process improvement and the establishment of new policy and procedure for monitoring, incident management, change control, and exchange operations

Your portfolio
  • 5+ years of experience in a Site Reliability Engineering, DevOps, production engineering, or technical operations role supporting a 24/7 production system
  • Strong Linux fundamentals and scripting ability
  • Experience supporting and debugging Java applications in production - reading stack traces and thread dumps, working with JVM memory and garbage-collection behavior, and diagnosing service issues from logs and metrics
  • Solid working knowledge of TCP/IP networking - comfortable reasoning about connections, ports, routing, and firewalls to debug connectivity between exchange components, partners, and datacenters
  • Hands-on experience operating Kubernetes in production and managing infrastructure as code (Terraform, Ansible) with CI/CD pipelines (Jenkins or similar)
  • Experience with modern observability tooling (Datadog, Prometheus, Grafana, or equivalent) and a track record of being on-call for systems that matter
  • Working knowledge of SQL and relational databases (PostgreSQL preferred); experience with Kafka or other streaming platforms a plus
  • Self-starter who can deliver results with minimal guidance
  • Comfortable working independently and with a team
  • Excellent communication and organizational skills - especially written incident communication and documentation
  • Background and interest in trading, capital markets, exchange operations, or sports betting a plus; familiarity with exchange protocols a plus
  • Previous experience in a regulated industry (gaming, finance) a plus
  • Startup experience preferred but not required

Perks
  • Medical, Dental, and Vision Benefits: Company pays 100% Employee premium and 50% Spouse & Dependent premiums
  • Short- & Long-Term Disability; Group Term Life and AD&D; Voluntary Life and AD&D
  • 401(k) Plan
  • Equity Options
  • Flexible time off
  • MacBooks issued to all employees

Are you ready to Trade Up?

Similar Jobs

More Jobs at Sporttrade

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: