Production Support Engineers are responsible for availability of systems and business capabilities including applications, infrastructure operations and data, business process execution which may include batch scheduling and processing, production of target business outputs, transmission of data including messages and files business needs requiring IT based solutions.
Key Responsibilities and Duties- Identify gaps in processing and production of outputs and availability of systems and outputs/outcomes. Monitor and alert the applications and services.
- Manage environments and technology processing response to incidents and crisis management.
- Responsible for triaging, recovery and remediation of incidents and problems including root cause analysis, defect and data analysis and impact assessment (i.e., end users, business processes, data, applications, and devices.
- Roles at this level should be managing People Leaders with span of control of at least 7 FTE direct reports. Typically manages one or more groups that has a moderate to high impact on business results. Implements policies and procedures for the areas they manage; may participate with other areas in establishing broader policies and processes. Large contribution to strategic planning within an area of expertise. Provides advice to internal and/or external clients on implications of business trends, issues, operating environment changes and firm or business unit strategy. Manages the performance of direct reports through regular, timely feedback as well as the formal performance review process to ensure the delivery of projects and engagement, motivation and development of the team.
- Demonstrates ability to apply understanding of concepts in complex situations at mastery level, serving as a key resource in providing advice, leadership, coaching and mentoring to others in the following:
Think critically, analyze complex data, and develop long-term plans that align with the organization's vision and mission.
Make sound decisions based on available information and data, and to effectively communicate those decisions to stakeholders.
Inspire, motivate, and guide others towards achieving common goals and objectives.
Communicate effectively with all levels of the organization, including the board of directors, senior management, and front-line employees.
Effectively manage budgets, financial reporting, and overall financial performance of the organization.
Lead and manage organizational change initiatives, including strategic planning, restructuring, and mergers and acquisitions.
Attract, retain, and develop top talent within the organization, and to create a positive and productive work environment.
Drive innovation and creativity within the organization, and to identify and pursue new growth opportunities.
Recognize and manage one's own emotions, as well as to understand and respond effectively to the emotions of others.
Understand and appreciate diverse cultural perspectives and to create a culture of inclusion and equity within the organization. - If you are part of an Agile team, you may be asked to perform the functions of Analyst (Tech) , Dev, Lead Dev , or Engineer.
Educational Requirements- University (Degree) Preferred
Work ExperienceCareer Level11PL
POSITION SUMMARY:This role establishes and operates a
Service Reliability Platform Organization, integrating command center operations, IT service management (ITSM), production operations, and ServiceNow platform capabilities into a unified
control tower operating model.The Managing Director ensures operational excellence through
strong governance, automation, risk management, and continuous improvement, while delivering a resilient, scalable, and client-centric production environment.
Position Scope & Reporting:- Reports to: Head of Workplace Experience & Production Support
- Leadership Level: Managing Director (Executive Band)
- Span of Control: Multi-functional global organization (Command Center, ITSM, Production Ops, Platform Engineering, Reliability Engineering, Reporting)
- Geographic Coverage: Global (24x7 follow-the-sun model)
KEY RESPONSIBILITIES: 1. Service Reliability & Executive Accountability - Serve as the single accountable executive for enterprise service stability, client-impacting incidents, and operational resilience outcomes.
- Lead a closed-loop service health model covering detection, governance, execution, and prevention.
- Establish KPIs and performance discipline aligned to enterprise priorities including Availability, MTTR, incident reduction, and client experience.
2. Command Center Leadership (Enterprise Control Tower) - Lead the Operational Command Center (OCC/NOC) responsible for real-time monitoring, detection, and incident command.
- Operate a control tower model delivering integrated oversight across detection, triage, and response workflows.
- Ensure rapid decision-making and coordinated response for major incidents with clear ownership and escalation protocols.
3. IT Service Management (Governance & Risk Discipline) - Oversee enterprise Incident, Problem, and Change Management processes with strong adherence to governance, audit, and regulatory expectations.
- Ensure root cause elimination and risk mitigation via effective problem management practices.
- Maintain ITSM as the governance spine of production operations.
4. Production Operations (Execution & Stability) - Lead global production operations ensuring availability, performance, and operational integrity of the enterprise platform.
- Oversee execution of operational workflows including patching, batch processing, and recovery activities.
- Ensure strong alignment between change governance and operational execution.
- Work in the coordination and collaboration with other Ops functions sitting in the CIO and Shared Service orgs.
5. ServiceNow Platform Strategy & Automation - Lead ServiceNow as a strategic enterprise platform, operating as a system of record, workflow engine, and intelligence layer.
- Drive adoption of automation, AI, and workflow orchestration to improve scale, efficiency, and responsiveness.
- Ensure strong data integrity (CMDB, service mapping) supporting operational decision-making.
6. Reliability Engineering & Continuous Improvement - Establish proactive reliability engineering capabilities focused on trend analysis, service stability, and incident prevention.
- Drive continuous improvement initiatives to reduce recurring failures and optimize client-facing services.
- Enable transition from reactive operations to automation-driven resilience.
7. Client Experience, Reporting & Executive Communications - Deliver transparent, data-driven reporting on SLA/SLO performance and client impact.
- Provide executive-facing dashboards and insights supporting strategic decisions.
- Lead communications strategy during high-severity incidents to senior stakeholders and impacted clients.
Risk, Control & Governance Responsibilities (TIAA-Specific Emphasis) - Ensure alignment with enterprise risk management, regulatory requirements, and audit standards
- Maintain strong change control, incident documentation, and traceability across all workflows
- Partner with Risk, Compliance, and Audit teams to ensure control effectiveness and remediation tracking
- Promote a culture of accountability, transparency, and operational discipline
Leadership & Talent Expectations - Build and lead a high-performing, globally distributed organization
- Drive a culture of ownership, continuous improvement, and client-first thinking
- Attract and develop talent across operations, platform engineering, and reliability disciplines
- Maintain appropriate span of control across functional leads (typically 5-9 direct leaders)
Key Performance Measures - Enterprise MTTR and incident response effectiveness
- Reduction in incident volume and recurrence
- Percentage of automated resolutions and workflow efficiency
- Change success rate and operational risk reduction
- Client-impacting incident metrics and service stability
REQUIRED QUALIFICATIONS:- 10+ years of experience in enterprise technology operations, service management, or reliability engineering
- Executive leadership experience managing large-scale, mission-critical production environments
- Proven ability to lead global 24x7 operations and incident management functions
- Deep expertise in ITSM frameworks, operational governance, and risk management
- Experience leading platform-driven transformation (ServiceNow or equivalent)
PREFERRED QUALIFICATIONS:- 15+ years of experience in enterprise technology operations, service management, or reliability engineering
- Experience in financial services or other highly regulated environments
- Background in building SRE or service reliability organizations
- Experience with automation, AI-enabled operations, and digital transformation initiatives
Strategic Impact - This role is central to advancing TIAA's operational model from fragmented accountability to a unified Service Reliability Platform Organization, delivering:
- End-to-end ownership of service outcomes
- Improved client experience and responsiveness
- Automation-driven operational scale
- Enhanced governance, auditability, and transparency
Related Skills
Debugging, Prioritizes Effectively, Problem Solving, Systems Design/Analysis
Anticipated Posting End Date:2026-07-30
Base Pay Range: $206,000/yr - $309,000/yr
Actual base salary may vary based upon, but not limited to, relevant experience, time in role, base salary of internal peers, prior performance, business sector, and geographic location. In addition to base salary, the competitive compensation package may include, depending on the role, participation in an incentive program linked to performance (for example, annual discretionary incentive programs, non-annual sales incentive plans, or other non-annual incentive plans).