DevOps EngineerPosition SummaryWe are seeking a highly skilled DevOps Engineer to design, build, automate, secure, and operate our Digital Products Platform. This role will be responsible for managing cloud infrastructure, CI/CD pipelines, infrastructure as code, platform reliability, monitoring, security automation, and operational excellence across multiple digital products and services.
The successful candidate will collaborate closely with Software Developers, Solution Architects, Product Owners, and Technology Leadership to enable rapid, secure, and reliable software delivery while maintaining platform stability and scalability.
Key ResponsibilitiesCloud Platform Engineering- Design, implement, and maintain cloud infrastructure supporting enterprise digital products and services.
- Manage production and non-production environments ensuring high availability, resiliency, and scalability.
- Support cloud resource optimization, capacity planning, and cost management initiatives.
Infrastructure as Code (IaC)- Develop and maintain Infrastructure as Code using automation frameworks and declarative configuration tools.
- Standardize infrastructure deployment patterns and reusable platform components.
- Establish environment provisioning and configuration management best practices.
CI/CD & Release Automation- Design, build, and maintain automated CI/CD pipelines.
- Implement automated testing, deployment, and release management processes.
- Support blue/green, rolling, and zero-downtime deployment strategies.
- Improve deployment frequency while maintaining security and reliability standards.
Platform Reliability & Operations- Monitor platform health, availability, and performance.
- Develop automated alerting, logging, and operational dashboards.
- Support incident management, root cause analysis, and post-incident reviews.
- Drive continuous improvement of platform reliability and operational maturity.
Security & Compliance- Implement DevSecOps practices throughout the software delivery lifecycle.
- Manage vulnerability scanning, secrets management, access controls, and infrastructure security.
- Ensure platform compliance with organizational security standards and regulatory requirements.
Observability & Monitoring- Implement enterprise monitoring, logging, tracing, and performance management solutions.
- Create operational metrics, SLAs, SLOs, and dashboards.
- Proactively identify and resolve performance bottlenecks.
Collaboration & Enablement- Work closely with development teams to improve deployment automation and platform adoption.
- Provide technical guidance on cloud-native architecture and operational best practices.
- Create technical documentation, standards, and operational runbooks.
- Mentor development and operational teams on DevOps practices.
Required Qualifications- 5+ years of experience in DevOps, Site Reliability Engineering, Cloud Engineering, or Infrastructure Automation.
- Experience managing enterprise cloud environments.
- Strong experience implementing CI/CD pipelines and deployment automation.
- Experience with Containerization & Elastic Container Service (AWS)
- Hands-on experience with Infrastructure as Code tools.
- Strong scripting and automation skills.
- Experience with monitoring, logging, and observability platforms.
- Experience with provisioning and managing database services.
- Understanding of networking, identity management, security, and cloud governance.
- Experience supporting production environments and incident response activities.
Preferred Qualifications- Experience supporting digital product platforms and distributed microservices architectures.
- Experience with Azure cloud services.
- Experience with AWS cloud services and Amazon AI/ML services.
- Experience implementing DevSecOps practices.
- Experience with AWS Lambda and Azure Functions.
- Experience with version control ie. GitHub, GitLab, Azure Repos, etc.
- Experience integrating enterprise applications, APIs, and data platforms.
- Experience with SaaS platform integration and enterprise authentication technologies.
- Knowledge of ITIL, Site Reliability Engineering (SRE), and operational excellence frameworks.
- Relevant cloud certifications (Azure, AWS, Terraform, etc.).
Success Measures- Improved deployment frequency and release reliability.
- Reduced incident volume and faster recovery times.
- Increased platform availability and performance.
- Enhanced security posture and compliance adherence.
- Greater automation across infrastructure and software delivery processes.
- Reduced operational overhead through self-service platform capabilities.
Why choose us?* Salary range: $110,000 to $135,000
* Opportunity to work in a collaborative and supportive environment.
* Competitive salary and benefits package.
* Room for growth and professional development.
** We thank all applicants for their interest. Only those selected for an interview will be contacted. We are committed to employment equity and welcome applications from all qualified individuals. **