Reporting to DevOps Lead, Production Operations, you will join a growing practice around Agile, DevOps, Continuous Integration/Development, and Infrastructure as Code. This is a multi-discipline systems engineering position within our Level 2 Application Support team supporting the continued growth and successes within our various business lines. The main focus will being the Level 2 Support for many of the application in production and helping agile teams ability to deliver reliable and secure technology.
This role is hybrid (2 - 3 days/week in the office) - based in Montreal, QC. Key Accountabilities:- Be part of a Centre of Excellence in the practice of DevOps and to help spread a culture of innovation driven by experimentation
- Off hours support, flexible working hours as required
- Identify opportunities to make systems more resilient and efficient using automation
- Interact with the Agile Development team and internal customer support teams to find resolution quickly for production incidents
- Collaborate regularly with colleagues both inside and outside the application support team
- Facilitate CI/CD in infrastructure using DevOps and Agile methodologies
- Debug systems and identify issues quickly and effectively by analyzing metrics
- Proactively identify problem areas for the team both technically and procedurally and recommend improvements
- Identify system resource bottlenecks and work with other teams to test and tune systems for optimal performance
- Deliver on various business initiatives by building and supporting innovative solutions using PaaS, IaaS, and traditional on-prem technology stacks
Must Haves:- 1 year of hands-on experience Agile methodology
- 1 to 3 years of experience in maintaining and operating mission critical distributed systems
- Must be able to develop script/code using Python and Shell from scratch
- Knowledge of Kubernetes, Helm, go templating, ArgoCD and Rancher
- Knowledge of at least one of the Cloud Providers such as Azure, GCP or AWS
- Must be well-versed with Linux OS and Windows
- Ability to quickly identify issues and performance bottlenecks by recognizing patterns in data analytics
- Strong experience using orchestration tools such as Ansible
- Excellent communication and technical documentation skills
- An inherent desire to automate everything
- Demonstrated ability to be an active contributor in a collaborative environment and be willing to take on problem ownership within the team
- The ability to work under pressure and with tight deadlines
Nice to Haves:- Knowledge of FIX protocol, order management system and/or market data is highly desirable
- Experience with systems benchmarking and performance testing
- Perl, JavaScript, Golang
- Familiarity with application and systems performance monitoring, data visualization and logging tools such as Splunk, ELK, Grafana, Graphite, Nagios