Tangentia

Senior Site Reliability Engineer / Devops

Tangentia$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience in Site Reliability Engineering or related positions
  • Familiarity with Apigee Hybrid, Google Distributed Cloud, Azure, GCP, and Kubernetes
  • Strong knowledge of CI/CD, automation, monitoring, and infrastructure as code
  • Skill in certificate management scripting and automation
  • Proficient in Ansible for configuration management and orchestration
  • Experience using APM tools like Dynatrace and Splunk
  • Programming experience in Python

Responsibilities

  • Oversee reliability, availability, and performance of Apigee Hybrid and Google Distributed Cloud environments
  • Manage and automate certificate management processes
  • Plan and execute upgrades and maintenance for Apigee Hybrid and cloud infrastructure
  • Implement and maintain monitoring solutions using Dynatrace and Splunk
  • Troubleshoot production incidents and conduct root cause analysis
  • Develop automation scripts and Ansible playbooks for operational efficiency
  • Collaborate with cross-functional teams for security and compliance
  • Mentor team members in SRE methodologies

Benefits

  • Hybrid work environment
  • Opportunity to mentor and develop team members
  • Exposure to cutting-edge technologies like Google Distributed Cloud
  • Participation in incident handling and root cause analysis
  • Focus on continuous improvement and operational excellence
Full Job Description
Role: Senior DevOps & Site Reliability Engineer
Location: Toronto, ON
Interview Mode - Virtual
Hybrid

Key Responsibilities:
  • Oversee the reliability, availability, and performance of Apigee Hybrid and Google Distributed Cloud environments, ensuring robust SRE practices.
  • Manage and automate certificate management processes, including renewals, deployments, and compliance checks.
  • Plan and execute upgrades and maintenance activities for Apigee Hybrid and distributed cloud infrastructure, minimizing downtime and ensuring seamless transitions.
  • Implement and maintain monitoring solutions using Dynatrace and Splunk, proactively identifying and resolving issues to ensure system health and performance.
  • Troubleshoot complex production incidents, perform root cause analysis, and drive incident resolution to restore service quickly and prevent recurrence.
  • Develop and maintain automation scripts and Ansible playbooks for operational efficiency, including tasks such as Kubernetes context retrieval, proxy configuration, and container management.
  • Collaborate with cross-functional teams to ensure security, compliance, and best practices are followed across all SRE activities.
  • Mentor and guide team members in SRE methodologies, fostering a culture of continuous improvement and operational excellence.
Required Skills for this role:
  • 3+ years of experience in Site Reliability Engineering or related roles.
  • Experience with Apigee Hybrid, Google Distributed Cloud, Azure, GCP, and Kubernetes.
  • Advanced DevOps and SRE skills: CI/CD, automation, monitoring, infrastructure as code.
  • Certificate management scripting and automation.
  • Proficiency with Ansible for configuration management and orchestration.
  • Experience with APM tools such as Dynatrace, Splunk
  • Programming experience with python
Ansible (Software), Apigee Hybrid, API Management, Azure Kubernetes Service (AKS), CI/CD, Dynatrace APM, Google Anthos, Kubernetes, Public Key Infrastructure, Python (Programming Language), Red Hat Enterprise Linux (RHEL), Site Reliability Engineering, Splunk, Terraform, VMware

Similar Jobs

More Jobs at Tangentia

More Information Technology Jobs

Find similar Senior Site Reliability Engineer / Devops jobs: