- Deploy cloud infrastructure and services on Google Cloud Platform
- Assist in developing, maintaining and monitoring production 24 * 7 * 365 infrastructure.
- Produce systems that are secure, scalable, automated, and well-documented, and train others to take on operations
- Collaborate with Engineering teams to deploy and support customer-facing web applications
- Ensure backup functionality and consistency, including restoration process and data recovery plans for disaster recovery
- Provide DBA support as assigned, doing installations, upgrades, database loading, migrations, testing
- Troubleshooting and resolution of database/server problems
- Backup and recovery of local and remote replicated databases
- Create and verify data backups including configuration data in critical environments
- Communicate timelines, service dependencies, resource constraints and progress with key stakeholders quickly and effectively
- Maintain the technical infrastructure for business continuity plan (BCP) and DR (Disaster Recovery)
- Organize and coordinate Disaster Recovery drills and Tabletops
- Develop and improve infrastructure and application performance monitoring and alerting systems
- Provide platform support as required, on 24x7 on-call rotation
What You Will Bring:Must have- 10+ years experience with design, implementation and management of highly available (HA) system architectures
- 7+ years experience administering and troubleshooting Linux and Windows systems
- 5+ years Containerization
- 5+ years Kubernetes experience
- 3+ years Database Administration experience (e.g. Postgresql, MSSQL, CloudSQL)
- Hands-on experience with cloud infrastructure such as GCP, AWS, Azure, etc.
- Experience managing large-scale infrastructure using automation for provisioning, software deployment and configuration management tool (e.g. Puppet, Ansible, Terraform, etc.).
- Experience handling on-call shifts for mission critical systems.
Nice to have- Strong scripting skills in Bash, Ruby, Python, or similar languages, with a focus on automation
- Understanding of DevSecOps principles and practices
- Exposure to CI/CD platforms such as ArgoCD, Cloud Build, and Cloud Run
- Knowledge of monitoring and observability tools such as Datadog and New Relic
- Experience with metrics collection and visualization using Prometheus and Grafana
- Knowledge of centralized logging solutions such as Loki and Graylog
This is a full-time, remote position with an expected salary range of
$130,000-$150,000 CAD annually, depending on experience and qualifications. Think Research does not require Canadian work experience for this role.
Think Research may use artificial intelligence-enabled tools to support certain administrative aspects of the recruitment process, such as transcribing or summarizing interviews and meetings. AI tools are not used to screen, assess, or make decisions about candidates. All hiring decisions are made by our team.