About the RoleWe are seeking a Systems Reliability Engineer to join the Technical Operations team supporting our Embedded Finance (EmFi) platform. In this role, you will own the reliability and resiliency of a large-scale enterprise platform, ensuring that our services remain highly available, performant and secure.
Job TitleSRE
Key Responsibilities- Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity.
- Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform.
- Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health.
- Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams.
- Own and advance the root cause analysis (RCA) process - investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence.
- Ensure timely creation and management of incident tickets (e.g., ServiceNow) and accurate incident tracking, aging and reporting.
- Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency.
- Collaborate with engineering, product and risk stakeholders to embed reliability best practices into the platform lifecycle.
Minimum Qualifications- Hands-on experience with monitoring, observability and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog.
- Proven experience operating and supporting a large-scale enterprise platform environment.
- Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes.
- Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices.
- Experience with ticketing and incident management workflows (e.g., ServiceNow).
- Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams.
Desired Qualifications- Experience in financial services, payments or embedded finance environments.
- Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash).
- Familiarity with cloud platforms, containerization and CI/CD pipelines.
- Experience defining and managing SLOs, SLIs and error budgets.
LocationJacksonville, FL, Berkeley Heights, NJ , Alpharetta, GA , Toronto, ON (Onsite)
Salary and Other Compensation:The annual [salary/hourly rate] for this position is between [$110K- $135K annually]/[$55/hr -65 per hour]. Factors which may affect pay within this range may include geography/market, skills, education, experience and other qualifications of the successful candidate.
Benefits:The Company offers the following benefits for this position, subject to applicable eligibility requirements: [medical insurance] [dental insurance] [vision insurance] [401(k) retirement plan] [long-term disability insurance] [short-term disability insurance] [5 personal days accrued each calendar year. The Paid time off benefits meet the paid sick and safe time laws that pertains to the City/ State] [10-15 days of paid vacation time] [6 paid holidays and 1 floating holiday per calendar year] [Ascendion Learning Management System]
Want to change the world? Let us know.Tell us about your experiences, education, and ambitions. Bring your knowledge, unique viewpoint, and creativity to the table. Let's talk!
Preferred Skills CI/CD
Job details Job ID 333116
Role System Reliability Engineer
Location Jacksonville, Florida, US
Job type Direct Hire
Recruiter Kanchan
Email
[email protected]