Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | Hybrid
About the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient enough to meet business goals.
- Investigate issues raised by customers, database administrators, and application support and suggest short term and long term solutions.
- Provide day-to-day technical support in maintaining the information system, including responsibility for ensuring processes and outputs and complete and error-free.
- Work with Operation Team for deployment and validation of changes to production during release/deployment/Change Request Implementations.
- Participate in production issue bridge and work with different R&D and Operation teams to resolve customer issues.
- Responding to and resolving escalated incidents for customer issues or monitoring alerts.
- Building diagnostic and analytical tools that improve the MTTA, MTTD, and MTTR.
- Configure and integrate commercially available monitoring tools into the production systems to improve availability, scalability, and latency.
- Participate in on-call rotation and handling time critical issues during on-call schedule.
How You Will Make an Impact - Responds to and resolves escalated incidents for customer issues or monitoring alerts within SLA
- In-depth analysis of incident RCA;
- Working with R&D and architecture teams on defects and runtime inefficiencies identified in the production environment;
- Building diagnostic and analytical tools that improve the MTTA, MTTD, and MTTR.
- Building systems/site monitoring tools for system health and APIs to ensure smooth operations of production systems
- Validate and Verify software deliverables for production readiness.
- Risk assessment and mitigation of changes to the production systems
- Conducts project planning, cost analysis and vendor comparisons (POC/POV) and works on project implementation.
- Works with development teams to enhance and improve system operability.
- Conducts tests of network redundancy, resilience and failover of network elements to ensure up-time standards are fully achieved.
- May be required to provide on-call service coverage with other department employees.
Required Experience - BS or MS in Computer Science or similar discipline
- Strong work experience in Unix/Linux
- Strong knowledge of Java Web-based enterprise applications, Python, or Bash to automate tasks and build tooling.
- Strong work experience and troubleshooting skills in Kubernetes and Docker.
- Working experience of Azure, AWS with (CloudWatch, EKS, EFS, S3, RedShift and other AWS services) and Infrastructure as Code (IaC) tools like Terraform or Ansible.
- Working Knowledge in basic networking and various application and transport protocols. HTTP(s), JMS etc. TCP, UDP etc
- Experience working with one or more of the following: Splunk, Datadog, Dynatrace, Zabbix, Prometheus, etc.
- Working Experience on Akamai/Cloudflare DNS, CDN , DataStream & WAF.
- Experience working with one or more of the following: Oracle, PostgreSQL, MongoDB
- Experience working with messaging subsystems: RabbitMQ, Interconnect, AMQ.
- Familiar with AI tools.
What Sets You Apart (preferred qualifications)
- Minimum 7 years of experience in developing Software projects and/or DevOps/SRE.
Join SS&C, where innovation meets global opportunities. Click here to apply.
#LI-PE1#LI-HYBRIDSS&C Technologies offers a comprehensive total rewards package designed to support your wellbeing, growth, and future. Our benefits include medical, dental, and vision coverage; a 401(k) plan with company match; paid time off, holidays, and parental leave; and professional development reimbursement opportunity.
Actual base salary will vary based on several factors, including but not limited to relevant skills, prior experience, education, demonstrated performance, and geographic location.
Massachusetts: The expected base salary for the position is between 130,000 USD to 140,000 USD.
Applications will be accepted on an ongoing basis until the position is filled.