Site Reliability Engineer

Databento

• $130K — $155K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in SRE, DevOps, or backend engineering, ideally in a trading firm or tech startup.
  • Proficient in observability tools like Prometheus and Jaeger for logging and metrics.
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Strong Python skills, particularly in application development and performance tuning.
  • Familiarity with Linux debugging tools like strace and gdb.
  • Demonstrated impact in previous roles, such as performance improvements or cost savings.
  • Good communication skills for remote collaboration.

Responsibilities

  • Own uptime, SLAs, and SLOs for API and platform services.
  • Establish reliability best practices for developers without hindering their workflow.
  • Build and maintain observability across logging, metrics, and tracing.
  • Design and implement high-availability deployment strategies.
  • Profile and optimize Python applications for performance and cost efficiency.
  • Debug production issues at the OS level using advanced tools.
  • Enhance deployment and CI/CD workflows.

Benefits

  • Flexible remote work environment.
  • Opportunity to work on cutting-edge technology in a transparent culture.
  • Engagement with a wide range of complex problems, from data processing to SaaS.
  • Chance to lead incident response and improve operational practices.
Full Job Description
We9re looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You9ll own the uptime, performance, and observability of our platform, and help set the standard for how the rest of engineering builds, ships, and monitors software.

The range of problems is wide. You9ll work on petabyte-scale data processing, SaaS problems like customer management and billing, and the query systems that power our APIs. Unlike most finance firms, we9re open about our tech and how we build it.

Responsibilities
  • Own uptime, SLAs, and SLOs across our API and platform services.
  • Set reliability and operational best practices for other developers without slowing them down.
  • Build and maintain observability across logging, metrics, and tracing.
  • Design and run high-availability deployment and containerization strategies.
  • Profile and optimize Python applications for throughput, latency, and cost.
  • Debug production issues down to the OS level using tools like strace, perf, eBPF, ss, and gdb.
  • Improve deployment and CI/CD workflows.
  • Join the on-call rotation, lead incident response, and run post-incident reviews.
  • Find what needs fixing on your own, then take projects from idea to completion.

Preferred background
  • Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup.
  • Hands-on experience with observability tooling for logging, metrics, and tracing (e.g. Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector).
  • Experience with containerization and high availability deployment (e.g. Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s).
  • Strong proficiency in Python, including application development and performance optimization.
  • Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb.
  • A track record of measurable impact in a recent role, such as improving performance by X%, speeding something up Nx, or saving $Y per year.
  • Experience with alerting and incident response best practices is a plus.
  • Familiarity with configuration management or infrastructure-as-code tools (Ansible, Terraform) is also helpful.
  • Bonus if you9ve done HTTP benchmarking, load testing, and capacity planning.
  • Database schema design and query optimization skills are nice to have.
  • Good communication skills and work ethic for a remote workplace.
  • An interest in financial data or algorithmic trading.
Notice about phishing scams

Be cautious of phishing scams impersonating Databento that may offer a job interview and request that you make a purchase through a phishing link. All official Databento emails come from databento.com or, occasionally, us.greenhouse-mail.io (as Greenhouse.io is our ATS). Any other domains-such as databento-careers.com, databento.online, databento.io, databento.us, etc.-are fake.

Our recruiting data suggests that underrepresented applicants often downplay their skills. Even if your experience doesn9t exactly match the qualifications listed, we still want to hear from you. Please apply!

Similar Jobs

More Jobs at Databento

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: