YOUR IMPACTAs a Senior Site Reliability Engineer, you will play a key hands-on role within a globally distributed SRE team responsible for the reliability, performance, and stability of the data services that power our customer-facing SaaS products. You will work on well-defined components of our distributed systems stack - such as Kafka, Elasticsearch, Cassandra, Solr, Redis, and OpenSearch - across both on-premises and public cloud environments (AWS, Azure, GCP). This role is ideal for engineers who enjoy solving challenging operational problems, contributing to automation and reliability improvements, and collaborating across teams to support high-quality services. You'll deepen your expertise in distributed systems while helping the team execute on operational excellence.
WHAT THE ROLE OFFERS- Operate, maintain, and scale distributed data services including Kafka, Elasticsearch, Cassandra, Solr, Redis, and OpenSearch.
- Build, enhance, and support infrastructure across on-prem and public cloud environments (AWS, Azure, GCP).
- Develop and maintain Infrastructure-as-Code (IaC) using Terraform and Ansible.
- Apply patches, perform routine maintenance, and ensure systems remain secure and compliant with internal standards.
- Participate in the design, deployment, and monitoring of data platforms in collaboration with SRE and engineering teams.
- Support incident response activities and participate in the on-call rotation for critical services.
- Assist with capacity planning, performance tuning, and health assessments of data services.
- Create and maintain documentation, including operational procedures, change plans, and incident reports.
- Contribute to automation and reliability initiatives that improve service performance and reduce manual work.
- Support service requests and help ensure SLA/OLA commitments are met.
- Participate in team knowledge-sharing and training activities.
- May require shift work and participation in a 24x7 on-call rotation.
WHAT YOU NEED TO SUCCEED- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field - or equivalent practical experience.
- 4+ years of experience in Information Technology supporting large-scale enterprise systems.
- 2+ years operating or supporting distributed data platforms (e.g., Kafka, Elasticsearch, Cassandra, Solr, Redis, OpenSearch).
- 2+ years working with automation and configuration tools such as Terraform and Ansible.
- Strong knowledge of Linux systems administration.
- Experience working with public cloud infrastructure (AWS, Azure, or GCP).
- Solid troubleshooting skills and ability to resolve complex technical problems.
- Excellent written and verbal communication skills.
- Self-driven, detail-oriented, and able to manage multiple tasks in a fast-moving environment.
- Familiarity with ITIL processes; certification is a plus.
- Experience with observability tools (Prometheus, Zabbix, Grafana, New Relic, etc.) is a plus.
Compensation: At OpenText, we offer a thoughtfully designed benefits package that supports your physical, emotional, and financial wellbeing. As you move through the hiring process, we're happy to provide more details about our compensation programs, including variable and commission compensation opportunities for eligible roles, vacation entitlement, and paid time off.
Salary Range:$93,320-$138,480; Depending on the candidate's education, experience, skills, geographical location, and alignment with internal equity and external market, actual salary may vary and be higher or lower than the range posted.