Staff Site Reliability Engineer

NinjaTrader

$160K — $210K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in DevOps, Site Reliability Engineering, or Platform Engineering roles
  • Expertise with Kubernetes, Docker, and container orchestration
  • Experience with CI/CD tools like GitHub Actions
  • Proficiency in programming languages such as Python, Bash, or Go
  • Hands-on knowledge of cloud platforms like AWS, GCP, or Azure
  • Familiarity with monitoring tools like Prometheus and Grafana
  • Strong communication and leadership skills for team collaboration

Responsibilities

  • Lead the SRE function, setting direction and mentoring engineers
  • Analyze and resolve production issues to maintain system uptime
  • Participate in a weekly 12x7 on-call rotation for incident response
  • Conduct root cause analysis for production incidents
  • Automate repetitive tasks and deployments
  • Develop monitoring and alerting systems with Product and QA teams
  • Collaborate with cross-functional teams to meet project deadlines and performance standards
  • Implement security and compliance best practices throughout the software delivery lifecycle

Benefits

  • Generous PTO policy
  • 7 Paid Holidays + 5 Conditional Holidays
  • 401K with 3.5% Company Match
  • Paid Parental Bonding Leave
  • Comprehensive Health, Vision, and Dental Coverage
  • 100% Coverage for Life and Disability Insurance
Full Job Description
Disclaimer: Please be advised that the most accurate and up-to-date information about our open roles-including job descriptions, compensation, and benefits-can only be guaranteed on our official job board. For the latest listings and details, please visit: https://job-boards.greenhouse.io/ninjatrader.

What you'll do:

Join our Platform Engineering team, where you'll ensure the availability, performance, and scalability of the trading systems that power NinjaTrader. You'll deploy single- and multi-cluster services to Kubernetes, partnering with Product Engineering and Technical Operations to strengthen monitoring, reduce operational toil, and keep our revenue-generating systems running for traders around the world. Working alongside technical leadership, you'll help maintain 99.95% uptime and high performance across the trading platform.

In this role you will:
  • Serve as the technical lead for the SRE function, setting technical direction and mentoring engineers across reliability initiatives
  • Analyze, troubleshoot, and remediate production issues with a systematic problem-solving approach to keep revenue-generating systems running
  • Participate in a weekly 12x7 on-call rotation, including weekend deployments and checkouts before markets open on Sundays
  • Perform initial root cause analysis and remediation of production incidents
  • Build tools that automate repetitive tasks, deployments, and incident responses with the goal of minimal human involvement
  • Design reliable monitoring and alerting systems with Product and QA teams, establishing and tracking SLIs/SLOs across web, mobile, desktop, and trading platforms
  • Leverage Infrastructure as Code tools like Terraform to automate provisioning, scaling, and management across all platforms
  • Collaborate with Product Engineering, Operations, and other cross-functional teams to deliver features on time while meeting scalability, security, and performance requirements
  • Implement security and compliance best practices (e.g., SOC 2, PCI DSS) throughout the software delivery lifecycle for web, mobile, and desktop deployments

What you'll need:
  • 8+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering roles
  • Expertise with Kubernetes, Docker, and container orchestration
  • Hands-on experience with CI/CD tools (GitHub Actions or equivalent)
  • Proficiency in programming languages (e.g., Python, Bash, or Go) and automation tools such as Ansible, Terraform, or Helm
  • Hands-on experience with AWS, GCP, or Azure, including in-depth knowledge of networking, security, and identity management in cloud environments
  • Knowledge of monitoring and observability tools such as Prometheus, Grafana, Datadog, or similar
  • Strong collaboration, communication, and leadership skills, with the ability to influence technical decisions across teams and mentor junior engineers

Bonus points for:
  • Trading industry experience
  • Contributions to open-source projects

Compensation:

The salary range for this role will be $160,000.00 - $210,000.00 USD. In addition, this position will also receive an annual target bonus of 12%. Bonus pay at NinjaTrader is based on individual performance (50%) as well as company/team performance (50%).

Salary and bonus earnings are only two components of the total compensation package offered by NinjaTrader. NinjaTrader offers a 401K plan through ADP under which the company will match up to 3.5% of employee contributions. Annual paid time off allowance accrues at a rate of 23 days per year plus seven paid holidays.

Location:

This role is based in Chicago, IL. We are not open to remote candidates for this role

Hybrid:

For Chicago-based employees, we follow a hybrid work schedule: In-office Tuesday through Thursday, with remote work on Mondays and Fridays. In addition to these weekly remote days, we offer:
  • 20 additional flex remote days annually
  • 5 Company Wide Office-Optional weeks tied to major holidays


Our Core Benefits Include:
  • Generous PTO
  • 7 Paid Holidays Annually + 5 Conditional Holidays Annually
  • 1 Service Day Annually
  • 401k with 3.5% Company Match
  • Paid Parental Bonding Leave
  • Health, Vision, Dental Coverage
  • Life and Disability Insurance Covered 100% by NinjaTrader

Similar Jobs

More Jobs at NinjaTrader

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: