About The Role
Location: Vancouver, BC
Employment type: Full-Time, Permanent
Reports to: Site Reliability Lead, Platform Engineering
We're looking for Site Reliability Engineers to help build and operate highly available, secure, and scalable SaaS
systems at PayByPhone. This role is open across the experience spectrum - whether you're a few years into an SRE
or ops career or a seasoned specialist, we'll shape scope, mentorship, and support around where you're starting
from.
This role calls for a motivated, quality- and results-oriented person who enjoys collaborating with cross-functional
teams of skilled developers. The focus of the role is to:
• Help ensure the PayByPhone platform meets its availability and stability requirements
• Drive continuous improvements in our software quality assurance processes, practices, and culture
Working under the Site Reliability Lead (SRL), this role helps drive reliability by ensuring deliverables are
implemented on time and on budget, and are operating at the expected levels of 24/7/365 at four nines of
availability (target 99.99%) and within specific customer SLAs. This role also supports the SRL's agenda of incident
prevention, improved incident response, and timely remediation
Key Responsibilities
- Help ensure platform reliability and stability deliverables are implemented on time and on budget, and are operating at 24/7/365 at four nines of availability, within specific customer SLAs
- Standardize our quality assurance plans and templates for cross-team projects
- Assist in quality assurance tooling selection and operational procedures
- Help ensure continuity across all critical business transactions
- Generate and report on reliability metrics to various stakeholders
- Work collaboratively as a member of the team to define, refine, and execute the Platform Reliability / DevOps and SRE roadmap
- Support PayByPhone in meeting security and compliance requirements by collaborating with the Security & Compliance team and implementing security and compliance tooling, processes, and policies as part of the CI/CD process
- Work collaboratively with release & support management leaders on compliant and secure processes for triage, investigation, resolution, and release of software
- Assist the SRL in the quality assurance process within the SDLC, including automation language selection and usage
- Help build a strong sense of ownership and accountability across teams and individual contributors, reflected at the code level and in implementation and operations
- Contribute to high-quality execution and technical and operational excellence
- Gather and analyze metrics from operating systems and applications to support performance tuning and fault-finding
- Participate in system design consulting, assisting the Platform team where needs overlap with reliability and availability
- Assist in scoping, developing, and testing the ongoing disaster recovery plan under the SRL's leadership
- Help own observability, monitoring, and alerting tools under the SRL's leadership - proactively looking for ways to improve, standardize, document, and train
- Provide operational support: documentation and debugging of production issues, including -
- Being available to join Sev-1 pages
- Assisting the SRL with responsibilities such as post-mortems and on-call bootcamps
- Assist with CI/CD toolset setup and support (GitLab runners) from a reliability standpoint
- Support cost reduction, budget setting, and monitoring across the platform and the tools used within it
- Support software reliability practices (logging standards, use of reliable libraries, SLA/SLO goals)
- Participate in on-call responsibilities when needed
- Maintain a personal data plan to support your on-call responsibilities
Key Requirements
- 3+ years of experience in software development, delivery, or site reliability / operations for large, complex software systems spanning both legacy and modern stacks. Equivalent experience gained through non-traditional paths is welcome.
- Bachelor's or higher degree in Computer Science, Computer Engineering, or a related technical field is preferred; equivalent hands-on experience will also be considered.
- Experience working with high-performing SRE, Ops, or Dev teams
- Solid grounding in quality assurance discipline, software quality management, and related frameworks and tools
- Working experience with Amazon Web Services (AWS) solution architectures and technologies
- Experience building verification and validation practices into end-to-end delivery pipelines, from business development through the delivery phase, to accelerate product launch to market
- Familiarity with testing techniques such as unit, functional requirement, performance, GUI, regression, integration, system load, vulnerability assessment, security testing, and test automation
- Understanding of SaaS multi-tenant and distributed / micro-service architectures
- Understanding of DevOps principles, processes, and tools (e.g. IaC, CI/CD, and orchestration)
- Understanding of cloud computing architecture, services, and platforms
- Understanding of web and/or mobile development technologies and programming/scripting languages
- Ability to program (structured and OOP) using one or more high-level languages, such as Python or JavaScript/TypeScript
- Experience with distributed storage technologies such as NFS, HDFS, and Amazon S3, as well as dynamic resource management frameworks
- Experience with Infrastructure as Code (IaC)
- A proactive approach to identifying problems, performance bottlenecks, and areas for improvement
- Comfortable working with and supporting cross-functional teams
- Strong written communication, including technical documentation and training materials
What We Offer
Compensation: The expected salary range for this role is $110,000 - $120,000 CAD. Final compensation will be based on factors such as experience, skills, qualifications, and internal equity.
Retirement Savings Program: Access to our retirement savings program (RRSP for Canada / 401(k) for U.S.-based employees).
Vacation: All permanent full-time employees start with 4 weeks of vacation per year.
Work from Anywhere: Up to 15 days of work from anywhere subject to management and IT Security approval.
Personal Days: We provide 5 personal days annually, in addition to paid sick days, to support flexibility and work-life balance.
Comprehensive medical & dental coverage:
Employee Assistance Program (EAP): Access to confidential support services and resources for you and your family.
Career Growth & Learning Support: Opportunities for professional development, continuous learning, and career progression.