OverviewThis is an office-based position.
Responsibilities
Job Summary:
The Senior Manager, Platform Engineering manages the activities of first-level managers and/or group leaders and individual contributors responsible for designing, building, and evolving the Company's cloud platform, CI/CD pipelines, IaC frameworks, observability architecture, and engineering standards. This leader is responsible for day-to-day operations and leading teams, projects, or functional area and focuses on achieving specific goals, driving performance, and ensuring team cohesion. The Senior Manager champions and enables change that results in impactful business outcomes while advancing platform maturity, migration execution, and next-generation cloud capabilities.
Platform Engineering:
- Manage multi-team delivery of IaC, CI/CD, and automation capabilities, ensuring cross-team coordination, consistent engineering standards, and alignment with the platform roadmap.
- Set team-level quality standards for IaC module development and CI/CD pipeline reliability. Drive adoption measurement and feedback loops with consuming teams (including Cloud Ops).
- Oversee CI/CD pipeline architecture decisions across teams, ensuring consistency, security-by-default integration, and alignment with release engineering patterns.
- Coordinate container platform delivery across teams, ensuring consistent cluster standards, namespace policies, and registry governance.
- Champion automation initiatives across managed teams, setting targets for toil reduction and tracking delivery velocity improvements.
Observability Suite Architecture:
- Manage multi-team delivery of observability platform components, ensuring cross-team consistency in instrumentation standards and data pipeline architecture.
- Coordinate observability platform delivery with Cloud Ops requirements, ensuring the platform supports their operational alerting, dashboarding, and on-call needs.
- Ensure observability-by-default is embedded across platform capabilities: every IaC module, CI/CD pipeline, and platform service includes baseline telemetry.
- Coordinate AI/ML observability delivery with Data Science and Cloud Ops requirements, ensuring model-specific telemetry integrates into the broader observability platform.
Migration Factory & Cloud Transformation:
- Manage multi-team migration execution across workload streams, ensuring coordination of dependencies, risk management, and adherence to the migration program timeline.
- Lead cross-team workload discovery and dependency analysis, identifying migration risks and escalating blockers to Director-level leadership.
- Coordinate cutover planning and execution across teams, partnering with Cloud Ops on operational readiness validation for migrated workloads.
Greenfield / Next-Gen Cloud Platform Development:
- Manage multi-team delivery of next-gen platform capabilities, coordinating dependencies and ensuring architectural consistency across platform components.
- Sponsor and oversee proof-of-concept initiatives across teams, evaluating new technologies for strategic fit and scalability.
- Establish processes for capturing operational requirements from Cloud Ops and Engineering stakeholders, embedding supportability and operability criteria into platform design standards.
Platform Reliability Engineering & SRE Enablement:
- Manage multi-team delivery of SRE enablement capabilities across the platform, ensuring tooling supports the SRE framework governed by Cloud Ops.
- Ensure reliability and resilience testing is embedded by default into platform CI/CD pipelines and deployment workflows.
Developer Experience & Platform Enablement:
- Lead developer experience initiatives across managed teams: establish friction logs, measure onboarding lead times, and coordinate platform office hours and enablement sessions.
Platform Lifecycle Management:
- Coordinate platform lifecycle activities across managed teams, ensuring consistent patching cadence, upgrade planning, and deprecation timelines.
Platform Engineering KPIs:
- Analyze platform effectiveness across multiple teams, identifying cross-team patterns in engineering productivity, pipeline performance, and platform adoption. Present analysis and recommendations to Director-level leadership.
Platform Tooling & Vendor Management:
- Manage vendor relationships for platform tooling within scope and provide recommendations on vendor performance and contract renewals to Director-level leadership.
Platform Cost Management / FinOps Partnership:
- Enforce cost allocation tagging across managed teams' platform resources. Contribute utilization analysis to FinOps reporting for platform shared services.
Security Integration:
- Ensure security-by-default is embedded across all managed platform capabilities, coordinating with Security teams on control requirements and compliance evidence.
- Ensure audit readiness across all managed platform services, coordinating with Security and Compliance on evidence requirements and addressing gaps.
DR Infrastructure Capabilities:
- Manage multi-team delivery of DR infrastructure capabilities, coordinating with Cloud Ops on requirements alignment and testing schedules.
- Ensure DR infrastructure components are code-driven, version-controlled, and testable across all managed environments.
AI/ML Platform Infrastructure:
- Coordinate delivery of AI/ML platform capabilities across managed teams, ensuring consistent patterns for model training, deployment, and artifact management.
- Ensure AI governance infrastructure is consistently implemented across managed platform services, coordinating with Security and Compliance on requirements for the Life Insurance and Wealth Management industries.
- Lead evaluation and integration of AI-assisted developer tools across managed teams — including AI-powered code review, IaC generation, intelligent test scaffolding, and documentation automation — tracking measurable productivity improvements.
- Coordinate AI gateway delivery across managed teams, ensuring consistent model access patterns and integration with security controls.
Cross-Functional Collaboration & Change Leadership:
- Improve effectiveness and efficiencies by initiating new approaches. Partner with Cloud Ops, Engineering, Security, and Product managers to align platform delivery with operational and business requirements.
People Leadership:
- Manage the activities of first-level managers and/or group leaders, and individual contributors. Responsible for day-to-day operations and leading teams, projects, or functional area. Focus on achieving specific goals, driving performance, and ensuring team cohesion. Identify high-potential team members and provide them with development and growth opportunities.
- Responsible for hiring, firing, performance management, appraisals, and pay reviews. Identify talent and provide growth opportunities. Create and execute on employee engagement strategies, engage team on focus areas, and monitor progress. Foster a culture of engineering excellence, automation-first thinking, and continuous learning across multiple teams.
Qualifications
- Typically requires 12+ years of relevant experience including leadership experience in software engineering, DevOps engineering, platform engineering, cloud architecture, or infrastructure engineering, with at least 3 years in a management role leading teams of engineers.
- Strong knowledge of AWS services and cloud-native architectures. Experience managing teams that build platform infrastructure for multi-region, high-availability environments.
- Strong proficiency with IaC tools and module library management. Ability to set team-level standards for IaC development, versioning, and publishing.
- Strong experience leading teams that build and maintain CI/CD platform capabilities. Ability to set pipeline architecture standards and evaluate new pipeline tooling.
- Strong experience leading teams that build observability platforms. Ability to set instrumentation standards and coordinate observability architecture with Cloud Ops operational needs.
- Proven experience leading multi-team migration programs, including wave planning, dependency management, and risk mitigation.
- Solid understanding of platform-as-a-product principles. Ability to gather requirements from consuming teams and translate them into platform features.
- Solid knowledge of security-by-default practices across platform capabilities. Ability to coordinate security control implementation with Security teams.
- Strong understanding of reliability engineering. Ability to lead teams that build SRE enablement tooling (chaos engineering, self-healing, resilience testing) aligned with Cloud Ops SRE framework requirements.
- Solid knowledge of compliance and regulatory considerations for platform engineering in the Life Insurance and Wealth Management industries.
- Solid knowledge of compliance and regulatory considerations for platform engineering in the Life Insurance and Wealth Management industries.
- Proven experience leading teams that build and maintain automated DR infrastructure at scale.
- Solid knowledge of AI-enabled platform engineering practices: integrating AI capabilities into CI/CD pipelines (intelligent test selection, AI-powered code review, pipeline failure prediction), IaC workflows (AI-assisted generation and validation), and observability (AI-enhanced anomaly detection).
- Ability to evaluate AI-assisted developer tools for team adoption, assess productivity impact, and ensure compliance with data handling and security policies. Developing familiarity with MLOps pipeline patterns and AI governance tooling concepts.
- Experience operating in regulated industries (Life Insurance, Financial Services, or Wealth Management) is preferred.
Benefits
We offer a competitive compensation and benefits package, opportunities for career growth, an employee stock purchase plan, 401(k), generous time off and flexible work/life balance, company-matched retirement packages, an employee wellness program, and an awards and recognition program – all in a creative, fast-growing, and innovative company.