About this role:
Wells Fargo is seeking a Lead Site Reliability Engineer (SRE) / Lead Systems Operations Engineer within the Consumer Technology (CT) organization. This role will provide technical leadership for operational excellence, platform reliability, resiliency, observability, and support readiness across critical consumer-facing applications and platforms.
The Lead SRE will serve as a senior technical leader responsible for driving reliability engineering practices, reducing operational risk, improving service availability, and enabling scalable platform operations. This role will partner closely with Application Development, Platform Engineering, Infrastructure teams, Shared Services, and External Vendors to ensure highly resilient, supportable, and observable solutions.
The ideal candidate combines deep technical expertise with strong operational leadership and will play a critical role in advancing Site Reliability Engineering practices across the organization.
In this role, you will support:
Reliability Engineering & Platform Stability
Lead reliability initiatives across critical business platforms and customer journeys.
Establish and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budget practices.
Improve platform resilience through automation, self-healing capabilities, capacity planning, and fault-tolerant designs.
Identify and eliminate single points of failure across applications, infrastructure, vendor integrations, and customer flows.
Champion engineering solutions that improve availability, scalability, recoverability, and operational maturity.
Incident Management & Operational Excellence
Serve as a technical lead during major production incidents, providing coordination, technical direction, and recovery leadership.
Drive improvements in Mean Time to Detect (MTTD), Mean Time to Diagnose (MTTDiag), and Mean Time to Recover (MTTR).
Lead Root Cause Analysis (RCA) efforts and ensure corrective actions are implemented and tracked to completion.
Identify recurring operational patterns and develop preventive solutions to reduce production incidents.
Develop and maintain incident playbooks, recovery procedures, and operational readiness standards.
Observability & Monitoring
Lead enterprise observability initiatives leveraging Splunk, Grafana, GCP Monitoring, AppDynamics, and related platforms.
Define monitoring standards, alerting strategies, dashboards, and customer journey observability solutions.
Partner with application and infrastructure teams to improve telemetry, logging, tracing, and synthetic monitoring capabilities.
Develop actionable operational metrics and executive-level reliability reporting.
Automation & Engineering Excellence
Drive automation strategies that reduce manual effort and improve operational consistency.
Design and implement self-service operational capabilities and automated recovery solutions.
Utilize AI-assisted tools and engineering practices to improve incident detection, diagnosis, and remediation workflows.
Promote Infrastructure-as-Code (IaC), CI/CD best practices, and platform engineering principles.
Vendor & Dependency Management
Partner with internal and external service providers to improve reliability, support responsiveness, and recovery performance.
Evaluate vendor operational performance and contribute to service improvement initiatives.
Establish and monitor operational readiness expectations for critical vendor dependencies.
Drive resilience planning and support strategies for third-party integrations.
Technical Leadership
Provide technical leadership and mentorship to SREs, Systems Operations Engineers, and Platform Support Engineers.
Lead technical reviews, operational readiness assessments, and production support governance activities.
Influence architecture decisions to ensure supportability, resiliency, observability, and operational sustainability.
Collaborate with engineering leaders to establish and mature Site Reliability Engineering practices across Consumer Technology.
Required Qualifications:
5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
5+ years of Site Reliability Engineering, Platform Engineering, Production Support, or equivalent experience demonstrated through work experience, military experience, training, or education.
5+ years supporting mission-critical production applications in large enterprise environments.
3+ years leading major incident management, operational support, or reliability engineering initiatives.
3+ years of experience with observability and monitoring platforms such as Splunk, Grafana, AppDynamics, Dynatrace, GCP Monitoring, or similar technologies.
2+ years of experience driving automation, operational improvements, and reliability initiatives.
3+ years of experience supporting distributed systems, cloud-based platforms, infrastructure, networking, and application architectures.
1+ year of experience supporting highly regulated or customer-facing financial services platforms
Desired Qualifications:
Consumer Technology, Credit Card, Payments, Lending, or Digital Banking experience.
Experience implementing Site Reliability Engineering (SRE) principles, SLOs, SLIs, and Error Budgets.
Experience with DevOps, CI/CD, Infrastructure-as-Code, and cloud-native architectures.
Experience with AI-assisted engineering, incident management automation, or observability platforms.
Strong executive communication and stakeholder management skills.
Experience leading cross-functional technical teams without direct authority.
Experience supporting vendor governance and third-party operational readiness initiatives.
Job Expectations:
Relocation assistance is not provided for this position
Visa sponsorship is not available for this position
Position requires onsite presence at one of the posted Wells Fargo locations.
Locations:
401 W. Las Collinas Blvd, Irving, Texas
300 S. Brevard St. Charlotte, North Carolina
2600 S. Price Rd. Chandler, Arizona
Posting End Date:
10 Sep 2026
*Job posting may come down early due to volume of applicants.