[8SN] Site Reliability Engineer (SRE) - UI/UX

Software Mind

$110K — $130K *
US-AnywhereRemote in Montreal, QC
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Support, or a related role.
  • Hands-on experience with Kubernetes for production service deployments and operations.
  • Proficient in Splunk for log analysis and incident debugging.
  • Demonstrated experience in incident response and root-cause analysis.
  • Strong troubleshooting and analytical skills with proficiency in problem-solving.
  • Ability to collaborate effectively with engineering and cross-functional teams.
  • Fluent written and verbal English, at least B2 level.

Responsibilities

  • Support the deployment and ongoing maintenance of production services on Kubernetes.
  • Monitor service health, availability, and performance metrics.
  • Troubleshoot production incidents using log analysis and monitoring tools.
  • Perform log analysis and debugging using Splunk to resolve issues.
  • Identify service problems and work with engineering teams for resolution.
  • Participate in incident response and production support activities.
  • Conduct first-level debugging of UI-related issues involving Web Components.

Benefits

  • Competitive salary and laptop provided.
  • Professional development and training opportunities available.
  • Opportunity to work with cutting-edge cloud and container technologies.
  • Flexible work arrangements within a collaborative team environment.
  • Contribution to impactful digital transformation initiatives.
Full Job Description
About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

Contract Duration: Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance.

Job Description

About the Role

We are looking for a Site Reliability Engineer (SRE) - UI/UX to support the deployment, operations, and ongoing maintenance of a production UI service running on Kubernetes.

This role focuses on monitoring service health, troubleshooting production issues, investigating incidents, and ensuring reliable service delivery. You will work closely with engineering and client teams to support production operations and complete work based on a client-directed backlog.

While this role supports a UI-based service, it is not a frontend development position. Working knowledge of Web Components is required to perform first-level debugging of UI-related issues, but deep frontend development expertise is not expected.

What You'll Do
  • Support the deployment, operations, and ongoing maintenance of production services running on Kubernetes.
  • Monitor service health, availability, and performance.
  • Investigate and troubleshoot production incidents using logs, monitoring, and debugging tools.
  • Perform log analysis and incident debugging using Splunk.
  • Identify service issues and collaborate with engineering teams to support timely resolution.
  • Participate in incident response and production support activities.
  • Perform first-level debugging of UI-related issues involving Web Components.
  • Support service reliability and continuous improvement initiatives.
  • Assist with CI/CD pipelines and cloud-native application operations when needed.
  • Work effectively within a client-directed backlog and established priorities.


Qualifications

Required Qualifications
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Support, or a related role.
  • Hands-on experience supporting the deployment, operations, and ongoing maintenance of production services running on Kubernetes.
  • Experience monitoring service health, troubleshooting production issues, and supporting service reliability.
  • Proficiency with Splunk for log analysis and incident debugging.
  • Experience participating in production incident response and root-cause analysis.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Experience collaborating with software engineering and cross-functional teams.
  • Ability to work independently and effectively within a client-directed backlog.
  • Excellent written and spoken English, at least B2 level.


Additional Information

Preferred Qualifications
  • Experience supporting CI/CD pipelines.
  • Familiarity with multi-tenant services.
  • Experience with cloud-native application operations.
  • Experience supporting high-availability enterprise or SaaS platforms.
  • Familiarity with additional monitoring and observability tools.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Familiarity with container and deployment technologies such as Docker and Helm.
    Working knowledge of Web Components and the ability to perform first-level debugging of UI-related issues.

What We Offer
  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Similar Jobs

More Jobs at Software Mind

More Information Technology Jobs

Find similar [8SN] Site Reliability Engineer (SRE) - UI/UX jobs: