American Bureau of Shipping

Sr. Manager, SRE, Operations & Product Support

American Bureau of Shipping • $125K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent experience.
  • 10+ years in site reliability engineering, production engineering, or cloud operations, with leadership experience in incident response.
  • Proven track record in operating production web or SaaS services and enhancing reliability through automation.
  • Experience in establishing observability and service health measures for production systems.
  • Familiarity with AI-assisted tools for technical operations and incident management, with a focus on validation before production use.
  • Experience with Azure cloud services is preferred.

Responsibilities

  • Define service health measures and reliability objectives for the Fleet Management System (FMS).
  • Establish end-to-end observability across applications and infrastructure, leveraging AI for anomaly detection.
  • Lead the technical incident-response model, including coordination and post-incident reviews.
  • Collaborate with engineering teams to design for resilience and performance under growth expectations.
  • Implement operational-readiness criteria for beta and production releases, ensuring effective monitoring and support handoffs.
  • Develop a technical support escalation model in partnership with customer support and engineering teams.
  • Support customer migrations and post-launch stabilization through effective monitoring and response plans.

Benefits

  • Opportunity to lead and shape the production reliability model for a next-generation Fleet Management System.
  • Collaborative work environment with cross-functional teams including engineering, QA, and customer support.
  • Access to AI-assisted tools to enhance operational efficiency and incident management.
  • Potential for career growth as the SRE and production-operations capability scales.
  • Engagement in innovative practices that combine traditional methods with cutting-edge technology.
Full Job Description
Job Description

ABS Group Digital Solutions is seeking a Lead, Site Reliability Engineering, Operations & Product Support to establish the production reliability and support operating model for its next-generation Fleet Management System (FMS). The role will help take FMS from development through beta, customer migration, and scaled SaaS operations. It combines hands-on reliability engineering with leadership of incident response, operational readiness, technical support escalation, and post-launch improvement.

Working with the Platform Engineering & Cloud Architecture lead, application engineering, QA, security, and customer support teams, this person will ensure the product is observable, recoverable, and supportable. The lead will use conventional and AI-assisted automation to streamline operations while establishing clear boundaries between technical SRE ownership and customer-facing support ownership.

What You Will Do:
  • Define service health measures and reliability objectives for FMS, including availability, latency, error rates, recovery expectations, and appropriate service-level indicators and objectives.
  • Establish end-to-end observability across applications, infrastructure, integrations, and AI-enabled services through actionable logs, metrics, traces, dashboards, health checks, and alerts; assess where AI-assisted anomaly detection can improve signal quality.
  • Lead the technical incident-response model, including severity definitions, on-call and escalation practices, incident coordination, recovery procedures, and post-incident reviews; use AI-assisted summarization and evidence gathering where it improves response without replacing human judgment.
  • Work with engineering and platform teams to design for resilience, performance, capacity, backup and recovery, and safe operation under expected customer and data growth.
  • Define and implement operational-readiness criteria for beta and production releases, including monitoring, runbooks, rollback plans, ownership, support handoffs, and post-release validation.
  • Establish the technical support escalation model and partner with customer support, product, and engineering to resolve issues and turn recurring incidents and tickets into permanent fixes; evaluate AI-assisted ticket categorization and knowledge retrieval to speed technical triage.
  • Support customer migrations, go-lives, and post-launch stabilization by preparing technical monitoring and response plans, triaging production issues, and incorporating lessons into repeatable procedures.
  • Automate routine operational tasks, health checks, deployment verification, incident triage, and recovery; use AI where it demonstrates improved speed or accuracy, with access controls, auditability, and human approval for production-impacting actions.
  • Track reliability, incident, supportability, and operational-efficiency trends; communicate risks, corrective actions, and progress to engineering and program leadership, including evidence of whether AI-assisted workflows reduce toil or improve outcomes.
  • Help build and mentor an SRE/production-operations capability as FMS moves from initial releases to scaled customer use.

What You Will Need:

Education and Experience
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent relevant experience.
  • 10+ years of relevant experience in site reliability engineering, production engineering, cloud operations, or software operations, including experience leading incident response or operational improvement across teams.
  • Demonstrated experience operating production web or SaaS services and improving their reliability through software engineering and automation.
  • Experience establishing observability, on-call practices, runbooks, and service-health or reliability measures for production systems.
  • Experience partnering with software engineering, platform engineering, and customer-facing support teams during releases, incidents, and customer go-lives.
  • Experience applying AI-assisted tools or workflows to technical operations, incident triage, monitoring analysis, support knowledge retrieval, or operational automation, with an understanding of how to validate results before use in production.
  • Experience with Azure cloud preferred.

Knowledge, Skills, and Abilities
  • Strong command of SRE practices, including SLIs/SLOs, incident response, root-cause analysis, performance, capacity, resilience, and disaster recovery.
  • Experience operating production SaaS applications on Azure, including compute, networking, identity, storage, containers, databases, integrations, and security.
  • Ability to build effective observability and on-call practices using telemetry, logs, metrics, traces, and tools such as Azure Monitor, Application Insights, and Log Analytics-without creating unnecessary alert noise.
  • Ability to automate secure deployments and operational workflows using scripting, APIs, CI/CD, infrastructure as code, managed identities, and Key Vault.
  • Sound judgment on release risk, rollback, customer impact, and the responsible use of AI-assisted operations, including data protection and human oversight of production-impacting actions.
  • Clear communication and collaborative leadership across engineering, support, and business stakeholders, including the ability to drive improvements without direct ownership of every team or system.

Reporting Relationships:

Reports to the Sr Director, Platform Engineering & Cloud Architecture. The role will initially lead cross-functional operational practices; direct-report scope will be determined as the SRE and production-operations capability scales. Customer-facing support teams retain ownership of routine customer communications and frontline support.

About American Bureau of Shipping

The American Bureau of Shipping (ABS) is a classification society, with a mission to promote the security of life, property and the natural environment, primarily through the development and verification of standards for the design, construction and operational maintenance of marine-related facilities. ABS is the largest classification society globally, with more than 200 offices in 70 countries. The organization was founded in 1862 and is headquartered in Spring, Texas.
Learn more about American Bureau of Shipping
Size
5,000 employees
Industry
Founded
1862

More Jobs at American Bureau of Shipping

More Information Technology Jobs

Find similar Sr. Manager, SRE, Operations & Product Support jobs: