Campaign Monitor

Manager, Product Engineering

Campaign Monitor$108K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or related field.
  • 6+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • 2+ years of leadership or direct engineering management experience.
  • In-depth understanding of SRE practices and cloud platforms (AWS/Azure).
  • Strong skills in production incident management and automated infrastructure management (Terraform).
  • Experience with monitoring tools like Datadog, Dynatrace, or Grafana.
  • Ability to lead teams during critical production outages.

Responsibilities

  • Lead and develop a high-performing team of SRE and DevOps engineers.
  • Establish Service Level Objectives (SLOs) and architect monitoring strategies.
  • Lead major incident response efforts and implement post-mortem analysis.
  • Collaborate with IT Security for bot mitigation and DDoS defense implementation.
  • Oversee operational readiness and administrative duties.

Benefits

  • Cruise and Travel Privileges for You and Your Family
  • Health Benefits
  • 401(k) plan
  • Employee Stock Purchase Plan
  • Training & Professional Development opportunities
  • Tuition & Professional Certification Reimbursement
  • Rewards & Incentives programs.
Full Job Description
Job Description

The Manager, Web and Mobile Site Reliability Engineering (SRE) leads the engineering team responsible for ensuring maximum uptime, high availability, performance, and resilience for enterprise web applications, mobile app backends, and public API endpoints. This role defines reliability standards, oversees 24/7 incident response, manages edge infrastructure and bot mitigation, and drives automated deployment and observability pipelines.

Here's a summary of what Princess is looking for in a Manager, Web and Mobile Site Reliability Engineering. Is this you?

Responsibilities:
  • Team Leadership & SRE Operations: Lead and develop a high-performing team of SRE and DevOps engineers supporting 24/7 high-volume web and mobile systems. Manage on-call rotations, incident command protocols, and operational readiness.
  • Reliability & Observability Governance: Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. Architect end-to-end monitoring, tracing, and alerting strategies using tools like Datadog, Dynatrace, or Grafana.
  • Incident Management & Remediation: Lead major incident response efforts, drive blameless post-mortems, and collaborate with engineering teams to prioritize root-cause fixes and architectural resiliency improvements.
  • Traffic, Edge & Security Management: Partner with IT Security (PCL IT Security) and CDN providers (Akamai) to implement bot mitigation strategies, DDoS defense, WAF rules, and edge caching for key APIs and digital endpoints.
  • Administrative: Perform all other administrative and organizational duties as required (time keeping, training, travel, collaboration and correspondence, etc.)


Knowledge & Skills:
  • Scope: Direct management of SRE and DevOps engineers. Operational oversight for consumer-facing web platforms, mobile backend APIs, edge routing networks, and cloud deployment pipelines.
  • Problem Solving: Rapidly diagnoses and mitigates complex system outages, performance bottlenecks, traffic anomalies, bot campaigns, and infrastructure failures in high-volume production environments.Resolves highly complex, enterprise-scale operational challenges that impact guest operations, maritime services, revenue-generating systems, regulatory requirements, and technology service availability. Anticipates emerging operational risks, evaluates competing business priorities, establishes governance frameworks, and makes decisions where significant operational, financial, service, and reputational consequences may exist. Develops innovative solutions to improve enterprise resilience, scalability, and operational effectiveness.
  • Impact: Directly ensures continuous operational availability, system security, optimal site performance, and guest trust across web and mobile touchpoints.
  • Leadership: The role requires strong leadership skills. Requires strong incident command leadership, strategic operational decision-making, calm under pressure, and collaborative mentorship.
  • Knowledge: In-depth understanding of Site Reliability Engineering practices, cloud platforms (AWS/Azure), containerization (Kubernetes, Docker), Akamai/CDN edge routing, bot detection, and CI/CD pipelines (GitLab).
  • Skills: Production incident management, automated infrastructure management (Terraform), performance tuning, distributed tracing, metrics-driven SLI/SLO establishment.
  • Abilities: Ability to lead teams during critical production outages, drive cross-functional engineering accountability for reliability, and automate operational workflows.


Essential/Minimum Qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, System Administration, or equivalent experience.
  • 6+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • 2+ years of leadership or direct engineering management experience.


Travel: Less than 25% with shoreside travel likely

Work Conditions: Work primarily in a climate-controlled environment with minimal safety/health hazard potential.

Physical Demands: Remain in a stationary position at a desk and/or computer for extended periods of time; reasonable accommodations will be offered.

**This position is classified as "hybrid." As an in-office role, it requires employees to work from a designated Princess location Mondays through Thursdays. On Fridays you can work from home.

Princess provides comprehensive and innovative benefits to meet your needs, including:

What You Can Expect
  • Cruise and Travel Privileges for You and Your Family
  • Health Benefits
  • 401(k)
  • Employee Stock Purchase Plan
  • Training & Professional Development
  • Tuition & Professional Certification Reimbursement
  • Rewards & Incentives


About Campaign Monitor

Campaign Monitor is a global email marketing and automation software company. The company was founded in 2004 and is headquartered in Sydney, Australia. Campaign Monitor provides a platform for businesses to create, send, and optimize email campaigns. The company has offices in San Francisco, London, and Nashville, and serves over 250,000 customers worldwide. Campaign Monitor's customers include small businesses, non-profits, and Fortune 500 companies.
Learn more about Campaign Monitor
Size
250 employees
Industry
Founded
2004

Similar Jobs

More Jobs at Campaign Monitor

More Information Technology Jobs

Find similar Manager, Product Engineering jobs: