Senior Site Reliability Engineer - Linux Systems & Application Observability

tastytrade$180K — $200K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Site Reliability Engineering or similar role.
  • Deep expertise in designing fault-tolerant distributed systems.
  • Experience running gap analyses on observability and telemetry systems.
  • Proficient with OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux internals and networking principles knowledge.
  • On-call experience with production systems and blameless post-incident reviews.
  • Strong programming skills in languages like Python, Ruby, or Java.

Responsibilities

  • Build self-healing and fault-tolerant infrastructure.
  • Conduct gap analysis on observability stacks to identify blind spots.
  • Manage scalability across the HashiCorp Nomad service fabric.
  • Enhance observability stack with necessary instrumentation.
  • Establish SLOs and error budgets for critical brokerage functions.
  • Mentor engineers to foster a culture of site reliability across teams.

Benefits

  • Performance bonuses for individual and company success.
  • Stock purchase options for employee investment.
  • Comprehensive medical, vision, and dental benefits.
  • Generous vacation and sick leave policies.
  • Gym membership reimbursement for wellness support.
  • Charitable donation matching to encourage community involvement.
  • Daily catered lunches and a fully stocked kitchen for in-office workers.
Full Job Description
Role: Senior Site Reliability Engineer - Linux Systems & Application Observability

Location: Chicago, IL (Hybrid, 3 days/week in office)

Role Summary
Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform
that active options, futures, and equities traders rely on every market day. As our first Senior Site
Reliability Engineer, you'll harden the systems behind order execution and market data delivery -
designing for fault tolerance, closing gaps in our telemetry, and making sure our HashiCorp Nomad-based
service fabric scales cleanly as trading volume grows. You'll work embedded alongside our infrastructure
and application engineering teams, contributing directly to our Ruby, Java, and Elixir services. This is a
rare opportunity to shape a practice and a culture from day one, on a platform where every order, quote,
and position has to be right, because real client capital is on the line.

What You'll Do (Job Responsibilities)
• Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive
operational work and reduces toil for Platform and Application teams.
• Run gap analysis across our observability stack to find blind spots in telemetry, logging, and alerting
coverage, then close them so failures surface before customers feel them.
• Own scalability work across our HashiCorp Nomad service fabric: capacity planning, load testing,
and identifying architectural bottlenecks before they become incidents.
• Extend our observability stack (Prometheus, Honeycomb, OpenTelemetry) with the instrumentation
needed to actually see the failure modes above.
• Set SLOs and error budgets with multi-window burn-rate alerting for critical brokerage flows, once
the fault-tolerance and telemetry foundation is in place.
• Mentor engineers across teams to build a culture of site reliability champions so the practice outlives
any one person.

Who You Are (Skills Needed)
• Hands-on experience designing fault-tolerant, self-healing distributed systems - not just describing
the patterns, but having shipped them.
• Deep understanding of one or more: distributed systems, Linux systems, cloud-native architectures,
containerization.
• Experience running gap analyses on observability/telemetry systems: identifying what's not
instrumented, not alerted on, or not visible until it's too late.
• A track record scaling systems under real production load, including capacity planning and
architectural bottleneck identification.
• Hands-on experience with OpenTelemetry, Prometheus, and Grafana, with the ability to instrument
services directly.
• Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet
capture, and flow analysis.
• On-call experience on production systems and comfort building a blameless post-incident review
process.
• Working knowledge of SLOs and error budgets as a tool, not the job description; HashiCorp Nomad,
Consul, or Vault experience is a strong plus.
• Strong programming skills in a language such as Python, Ruby, Java, or similar.

Company Perks + Benefits:
  • Performance Bonuses
  • Stock Purchase Options
  • Medical/Vision/Dental Benefits
  • 401k Plan
  • 20 Paid Vacation Days (plus an additional paid vacation day the month of your birthday!)
  • 10 Paid Sick Days
  • Gym Membership Reimbursement
  • Commuter Benefits
  • Pet Insurance
  • Wellness & Mental Health Programs
  • Charitable Donation Matching
  • Two Paid Volunteer Days Off
  • Daily catered lunch when in the office
  • Full kitchen with snacks and beverages
  • In-building gym
  • Shuttle to/from Metra

Base Salary Range: $180,000-$200,000 The actual salary offered will be based on the candidate's level of experience and qualifications.

Discretionary Performance Bonus: 15-20% of base salary based on individual and company performance.

Location: Our office is in the West Loop - Chicago's growing center of tech, great cuisine, and high-end bars.

*Don't meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they have every single qualification. Our team is dedicated to building a diverse, inclusive, and authentic workplace, so if you're excited about this role, but your experience doesn't align perfectly, we encourage you to apply anyway. You may be just the right candidate for this or other roles!

About tastytrade

tastytrade is a financial media company that provides online trading education, research, and financial news. The company was founded in 2011 by Tom Sosnoff and Kristi Ross, and has since grown to become one of the leading online trading platforms in the United States. tastytrade offers a wide range of financial products and services, including options trading, futures trading, and stock trading. The company is known for its innovative approach to trading education, which emphasizes the importance of risk management and probability-based trading strategies. tastytrade has received numerous awards and recognitions for its contributions to the financial industry.
Learn more about tastytrade
Size
200 employees
Industry
Founded
2011

Similar Jobs

More Jobs at tastytrade

More Information Technology Jobs

Find similar Senior Site Reliability Engineer - Linux Systems & Application Observability jobs: