We're looking for our next Senior Principal Site Reliability Engineer. Could It Be You? The Senior Principal Site Reliability Engineer is directly responsible for the stability, resiliency, and scalability of business-critical brokerage back-end applications running across a hybrid on-premises and cloud architecture.
This individual drives reliability engineering practices - SLOs/SLIs, observability, incident response, and capacity planning - while also making hands-on, code contributions directly into multiple applications across different technology stacks using an inner-source model. The role blends deep technical execution with cross-team influence: this person is expected to identify systemic reliability risks, fix them where they appear, and raise the operational bar for every team they touch.
This position is a strong fit for a hands-on senior principal engineer who is energized by fixing production reliability at the source, is comfortable navigating multiple codebases and cloud/on-prem environments, and wants to have an outsized impact on the resiliency of a regulated, high-availability brokerage platform.
Need more details? Keep reading...Application Stability, Reliability & Growth- Own the end-to-end reliability posture of critical brokerage back-end applications, driving measurable improvements in availability, latency, and error budgets across on-premises and cloud environments.
- Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets in partnership with application teams; use them to prioritize reliability work over feature work when warranted.
- Lead root cause analysis and blameless post-incident reviews for high-severity production incidents; drive remediation items to closure and identify systemic patterns across applications.
- Establish and mature observability practices (metrics, logging, tracing, alerting) so that failures are detected proactively and diagnosed quickly across a heterogeneous, multi-stack estate.
- Build capacity planning, load testing, and chaos/failure-injection practices to validate resilience before incidents occur.
- Champion a culture of operational excellence, toil reduction, and AI and automation-first thinking across engineering teams.
Cross-Stack Engineering & Inner Sourcing- Make hands-on, code contributions directly into multiple applications spanning different languages, frameworks, and stacks, using an inner-source model to fix reliability defects, add instrumentation, and improve resiliency patterns.
- Partner with individual application teams to raise pull requests, follow their contribution standards, and pair with owning engineers so fixes land safely and are properly reviewed and owned long-term.
- Identify recurring reliability anti-patterns across codebases (e.g., missing timeouts/retries, unbounded queues, improper connection pooling) and drive standardized, reusable fixes or shared libraries.
- Contribute to and help govern internal reliability tooling, shared SDKs, and common patterns (circuit breakers, backoff/retry, health checks) that can be inner-sourced across teams.
Cloud Scalability & Hybrid Architecture- Design and advise on cloud scalability strategies (auto-scaling, load balancing, multi-region/multi-AZ failover, caching, queuing) for workloads that span on-premises data centers and public cloud.
- Guide capacity and cost-aware scaling decisions, balancing performance, resiliency, and cloud spend across hybrid deployments.
- Evaluate and recommend cloud-native and hybrid resiliency patterns (e.g., disaster recovery, active-active/active-passive architectures, data replication strategies) appropriate for regulated brokerage workloads.
Organizational Awareness & Risk- Bring strong organizational awareness of the operational, financial, regulatory, and reputational risk that production incidents pose to a brokerage business, and factor that into prioritization.
- Participate in risk assessments related to system reliability, availability, and disaster recovery, partnering with Risk, Compliance, and Information Security as needed.
- Contribute to change management and release governance practices that reduce the likelihood and blast radius of production incidents.
- Promptly identify, escalate, and help remediate reliability or security-related incidents in accordance with company policy.
Leadership & Influence- Act as a technical reference and mentor for reliability engineering practices, coaching application teams on operational excellence without formal direct reports.
- Influence architecture and design decisions across multiple teams by bringing a reliability and scalability lens to reviews and planning.
- Document and evangelize reliability standards, runbooks, and best practices; lead or contribute to internal tech talks and communities of practice.
- Partner with engineering leadership to define the reliability roadmap and report on progress against stability goals
So are YOU our next Senior Principal Site Reliability Engineer.? You are if you...- Bachelor's or Master's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent combination of education and experience.
- 8+ years of software engineering and/or site reliability engineering experience, including production ownership of business-critical applications; financial services or brokerage experience strongly preferred.
- Demonstrated ability to read, debug, and make minor-to-moderate code changes across multiple languages/stacks (e.g., Java, .NET, Node.js/TypeScript, Python) in an inner-source or cross-team contribution model.
- Deep experience with cloud scalability strategies on one or more major providers (AWS, Azure, GCP), including auto-scaling, load balancing, multi-region resiliency, and cost-aware capacity planning.
- Experience operating and supporting hybrid architectures spanning on-premises data centers and cloud environments.
- Strong background in observability tooling (e.g., Prometheus/Grafana, Datadog, Splunk, ELK, AppDynamics, Dynatrace) and building actionable alerting and dashboards.
- Practical experience defining and operating against SLOs/SLIs/error budgets and running blameless post-incident reviews.
- Experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, Ansible, CloudFormation) in support of reliable, repeatable deployments.
- Solid understanding of microservices architecture, distributed systems failure modes, and resiliency patterns (circuit breakers, retries/backoff, bulkheads, timeouts).
- Familiarity with relational and NoSQL data stores and their operational/scaling characteristics.
- Experience with incident management and on-call practices (e.g., PagerDuty, Opsgenie) including leading major incident response
- Knowledge of security, audit, and regulatory considerations relevant to brokerage / financial services production systems.
- Excellent communication skills, with the ability to influence engineers and stakeholders across many teams without direct authority.
- Strong documentation, analytical, and problem-solving skills
Compensation Information:- Base salary range: $150,000 - $190,000
- The final compensation package will be commensurate with the successful candidate's experience, skills, and geographic location (Canada). It includes a comprehensive benefits plan and a competitive incentive (bonus) program for Full-Time Permanent roles.
Sounds like you? Click below to apply!#LI-Hybrid