Job Title:- Senior DevOps EngineerWork Type:- Full-TimeWork Mode:- HybridTru Inc is seeking a Senior DevOps / Site Reliability Engineer . We have one (1) permanent full-time position available for an immediate start.
Position Overview :- We are looking for a Senior DevOps / Site Reliability Engineer to help build, operate, monitor and scale our new enterprise AI platform. This role will be responsible for the DevOps and reliability capabilities supporting the platform across all environments including Development, QA and Production. The environment is currently focused and manageable consisting of approximately 10-30 containers but is expected to grow as the AI platform expands in 2027 and beyond.
This is an opportunity to work with a collaborative team on a new platform where the focus is not simply deploying infrastructure. The most important part of the role will be creating intelligent dashboards, monitoring, alerting, automated scaling and AI-assisted platform management. The platform is hosted primarily in Microsoft Azure with limited exposure to AWS.
Location:-This role is hybrid, providing some flexibility to work from home. However, candidates must reside in the Greater Toronto Area (GTA) and be prepared to attend in-person meetings as required. Occasional travel to client sites or workshops may also be necessary.
Key Responsibilities:-Azure Infrastructure and Platform Deployment:-- Design, deploy, configure and maintain infrastructure within Microsoft Azure.
- Deploy and manage virtual machines, containers, Kubernetes clusters, networking, storage and supporting platform services.
- Support infrastructure across Development, QA and Production environments.
- Establish repeatable and reliable deployment processes using Infrastructure as Code and CI/CD automation.
- Maintain secure, resilient and appropriately sized platform environments.
Kubernetes and Container Management:-- Deploy, configure and operate containerized applications using Kubernetes.
- Manage container lifecycle, configuration, secrets, networking, storage and application dependencies.
- Monitor container and cluster health, resource consumption, capacity and performance.
- Troubleshoot deployment, networking, configuration and runtime issues.
- Establish appropriate standards for container deployment and Kubernetes operations.
Performance, Load Management, and Scaling:-- Monitor platform demand, workload patterns, resource utilization and application performance.
- Configure horizontal and vertical scaling policies for containers and supporting infrastructure.
- Develop intelligent scaling approaches based on workload, queue depth, response time, resource utilization and business demand.
- Conduct capacity planning and identify potential performance bottlenecks before they affect production.
- Help introduce predictive or AI-assisted scaling and platform management capabilities.
Dashboards and Platform Visibility:-- Design and build advanced operational dashboards using tools such as Grafana, Kibana, Azure Monitor, Application Insights and similar technologies.
- Create clear executive, operational, application and infrastructure views of platform health.
- Build dashboards covering availability, performance, capacity, errors, latency, traffic, container health, AI workloads and service dependencies.
- Establish meaningful service-level indicators, service-level objectives and reliability metrics.
- Continuously improve dashboards so that issues, trends and risks can be quickly identified.
- Advanced dashboard design and dashboard-building experience is a core requirement for this role.
Monitoring and Alerting:-- Implement monitoring and alerting across infrastructure, applications, containers, integrations and AI platform services.
- Configure actionable alerts that identify real production risks while minimizing unnecessary alert noise.
- Establish thresholds, anomaly detection, health checks, synthetic monitoring and automated remediation where appropriate.
- Create operational runbooks and troubleshooting guidance.
- Work with development and architecture teams to improve platform observability.
Production Reliability and Support:-- Support the stability, availability and operational readiness of the production AI platform.
- Investigate and resolve platform, deployment, infrastructure, monitoring and performance issues.
- Participate in root-cause analysis and implement preventative improvements.
- Ensure that production support processes, documentation and escalation paths are established before platform usage increases.
- Provide very light production support during 2026, with no regular after-hours support currently anticipated.
- Help prepare the operating model for increased platform adoption and support requirements expected in 2027.
Requirements:-- Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure or platform engineering.
- Advanced hands-on experience with Microsoft Azure.
- Strong experience deploying and operating Kubernetes environments.
- Strong knowledge of containerization technologies such as Docker.
- Experience deploying and supporting containerized applications in Development, QA and Production environments.
- Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights or comparable tools.
- Strong experience implementing monitoring, observability, logging, alerting and operational health checks.
- Experience managing application load, infrastructure capacity, performance and automated scaling.
- Experience with CI/CD pipelines and automated application deployment.
- Experience with Infrastructure as Code tools such as Terraform, Bicep or ARM templates.
- Strong troubleshooting skills across applications, containers, infrastructure, networking and cloud services.
- Ability to work independently while collaborating closely with developers, architects, AI engineers and platform stakeholders.
Preferred Experience:-- Experience supporting AI, machine learning, data or high-compute platforms.
- Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues or model performance.
- Experience implementing automated remediation, predictive monitoring or AI-assisted platform operations.
- Familiarity with AWS services and cloud operations.
- Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus or similar observability technologies.
- Experience defining service-level indicators, service-level objectives and reliability standards.
- Experience with security, identity, secrets management and cloud governance within Azure.
What Makes This Role Different:-This is not a large-scale high-pressure production support environment. The initial platform footprint is relatively focused with approximately 10-30 containers and very limited production support expected during 2026.
The role offers the opportunity to establish the platform correctly from the beginning, introduce modern DevOps and SRE practices and experiment with intelligent monitoring, automated scaling, advanced dashboards and AI-assisted platform management. As platform adoption increases the responsibilities and operational scope are expected to grow throughout 2027.
Benefits:-Salary Range:- $100,000-$120,000