Senior Site Reliability Engineer

SimCorp A/S

$100K — $140K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or related field (Master's is a plus)
  • 3+ years in Site Reliability, DevOps, or Cloud Engineering roles
  • Expertise with Microsoft Azure Cloud
  • Experience with Infrastructure as Code (IaC) using Bicep, ARM, and Terraform
  • Knowledge of monitoring and logging tools (Azure Monitor, Application Insights, DataDog, Log Analytics)
  • Hands-on experience with IdP onboarding and integrating IdP solutions (Azure Entra ID, Okta, etc.)
  • Proficiency in IT service management (ITSM) frameworks like ITIL

Responsibilities

  • Support operational enhancements of mission-critical Cloud Native products & services
  • Collaborate with product development teams to improve service reliability and performance
  • Deeply understand systems at the code level through collaboration with engineering teams
  • Manage and improve infrastructure deployment pipelines while troubleshooting issues
  • Drive capacity planning for resilient and scalable platforms
  • Build automation tools to reduce manual efforts and enhance developer experience
  • Define and manage service level objectives (SLOs) in partnership with engineering teams

Benefits

  • Global hybrid work policy allowing remote work
  • Inclusive and diverse company culture
  • Focus on work-life balance
  • Empowerment to influence work processes
  • Opportunities for individualized career growth and professional development
Full Job Description

WHY THIS ROLE IS IMPORTANT TO US

As a Senior Site Reliability Engineer, you will be working on Cloud Native Products & Services, taking ownership of various responsibility domains like monitoring, observability, release management, vulnerability management, cost management, audit & compliance etc. You will work closely with DevOps engineers, clients, and stakeholders to ensure reliability, performance, and automation for both existing and new cloud native products & services. Onboard and long-running clients on them. Your contributions will drive stability, continuous improvement, and operational excellence in our Azure-based environments. This role blends hands-on engineering, incident response, platform configuration, and service quality, - guided by ITIL and SRE best practices.

WHAT YOU WILL BE RESPONSIBLE FOR

  • Support the operational and enhancement of mission-critical environments for both new and existing Cloud Native products & services

  • Collaborate with product development teams to enhance monitoring, observability, reliability, and performance of these services.

  • Collaborate deeply across engineering teams to understand systems at the code level.

  • Manage & improve our infrastructure deployment pipelines and troubleshoot onboarding and operational issues

  • Drive capacity planning efforts to ensure our platform is resilient and scalable as we grow.

  • Build tools and automation to eliminate manual TOIL, improve engineering velocity, developer experience, and improve system reliability.

  • Define and manage SLOs and error budgets in partnership with Engineering teams.

  • Contribute to incidents, problems, and change management processes.

  • Execute disaster recovery, configuration management, and platform readiness tasks.

  • Flexible working in regular & evening shift on rotational basis and provide weekend or On-Call support as needed.

  • Collaborate with Agile teams and take part in design discussions with clients, vendors, and stakeholders.

  • Contribute to knowledge sharing across multiple Product Areas.

  • Leverage a strong foundation in ITIL practices, including problem, change, and incident management.

WHAT WE VALUE

  • Bachelor6s degree in Computer Science or related field (Master6s is a plus)

  • 3+ years in Site Reliability, DevOps, or Cloud Engineering roles

  • Must have expertise with Microsoft Azure Cloud.

  • Expertise in Infrastructure as Code (IaC) using Bicep, ARM and Terraform.

  • Solid experience in monitoring and logging tools (Azure Monitor, Application Insights, DataDog, Log Analytics).

  • Hand-on experience in IdP Onboarding and integrating, configuring IdP solutions like Azure Entra ID, Okta, KeyCloak or PingFederate.

  • Experience in centralizing authentication, managing user identities, and implementing secure access protocols (SAML, OAuth, OIDC)

  • Experience working with observability frameworks like Open Telemetry and distributed tracing systems

  • Experience working with application reliability platforms like Checkly or equivalent

  • Experience setting up synthetic monitoring using Playwright or equivalent

  • Knowledge of AI/ML-based anomaly detection, log aggregation and analysis tools like Microsoft Azure Anomaly Detector or equivalent

  • Experience working with Microsoft Defender Suite (EDR, XDR) and Sentinel. Proficient in KQL for threat hunting and improving compliance scores using Defender for Cloud. Able to identify and remediate vulnerabilities

  • Understanding of networking, containerization (Kubernetes, Docker)

  • Good understanding of APIs, scripting languages like PowerShell, Bash, Kusto and databases like SQL, Cosmos DB and Postgres SQL

  • Familiarity with SimCorp Dimension & Sales force is a plus

  • Proficiency in IT service management (ITSM) frameworks like ITIL, focusing on incident, change, and problem management to improve operational efficiency

  • Experience managing both onboarding projects and live production operations

  • Collaborative mindset and ability to work in cross-functional teams

  • Interest in continuous learning and growth within your Product Area

Benefits

  • Global hybrid work policy - We ask you to work 2 days a week from the office. If you choose you can work remotely the other days. Of course, you are welcome at the office if that is your preference.

  • Culture 6 Inclusive and diverse company culture

  • Work-life balance 6 We believe that an equilibrium between professional responsibilities makes us all the best version of ourselves, both in private life and as colleagues in the workplace

  • Empowerment 6 We believe that all voices are valuable and must be heard. You will be involved in shaping our work processes

  • Career & Growth 6 Simcorp does offer opportunities for professional development: there is never just only one route - we offer an individual approach to professional development to support the direction you want to take.

NEXT STEPS

Please send us your application in English via our career site as soon as possible, we process incoming applications continually. Please note that only applications sent through our system will be processed. At SimCorp, we recognize that bias can unintentionally occur in the recruitment process. To uphold fairness and equal opportunities for all applicants, we kindly ask you to exclude personal data such as photos, age, or any non-professional information from your application. Thank you for aiding us in our endeavor to mitigate biases in our recruitment process.

We are eager to continually improve our talent acquisition process and make everyone6s experience positive and valuable. Therefore, during the process we will ask you to provide your feedback, which is highly appreciated.

For Toronto City only: The annual base salary range for this position is 100 000,00 - 140 600,00 CAD. Additionally, employees are eligible for an annual discretionary bonus, and benefits including health care, leave, and retirement plans.

Your total compensation may vary based on role, location, department and individual performance.

#Li-Hybrid

Similar Jobs

More Jobs at SimCorp A/S

More Information Technology Jobs

Find similar Senior Site Reliability Engineer jobs: