Bank of America Corporation

Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America Corporation$125K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in infrastructure, cloud, platform engineering, or site reliability engineering (SRE).
  • 5+ years managing Kubernetes and/or OpenShift in production environments.
  • Proven experience with large-scale distributed systems and 24/7 operational responsibilities.
  • Strong knowledge of Kubernetes, OpenShift, Rancher, and VKS orchestration platforms.
  • Familiarity with Infrastructure as Code tools and CI/CD platforms.

Responsibilities

  • Own platform reliability objectives for availability, resiliency, and operational health.
  • Lead incident response and problem management for critical issues.
  • Develop and maintain operational runbooks and recovery procedures.
  • Execute platform upgrades and cluster modernization efforts.
  • Design enterprise observability solutions with monitoring and alerting frameworks.
  • Drive automation for operational processes using Infrastructure-as-Code.
  • Conduct capacity planning and trend analysis to optimize resource consumption.

Benefits

  • Collaborative work environment across diverse technical teams.
  • Opportunities for continuous improvement and professional growth.
  • Focus on operational excellence with a commitment to quality.
  • Engagement with cutting-edge technologies and practices.
  • Support for strategic initiatives related to platform performance and reliability.
Full Job Description

Position Summary:

The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:

Reliability & Operations

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Kubernetes & OpenShift Platform Engineering

  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption.

Observability & Automation

  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.

Capacity & Performance Engineering

  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.

Security & Compliance

  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.

Required Qualifications:

Education / Experience

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Technical Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.

Desired Qualifications

  • BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications.
  • Experience running large-scale enterprise container platforms.
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
  • Experience implementing cloud-native security controls and platform governance.
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
  • Experience with vulnerability management and container security scanning solutions.
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicate effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.

Job Description:
This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads.Key responsibilities includeensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units.Job expectations includeadvancing efficient solution delivery practices andpromoting exceptionaldesign, engineering, and organizational practices.

Responsibilities:

  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Participates regularly in architecture community of practice meetings and communication via other channels
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations

Skills:

  • Automation
  • Collaboration
  • Influence
  • Production Support
  • Result Orientation
  • Analytical Thinking
  • Application Development
  • Architecture
  • Solution Design
  • Stakeholder Management
  • Adaptability
  • DevOps Practices
  • Project Management
  • Risk Management
  • Solution Delivery Process

Shift:

1st shift (United States of America)

Hours Per Week:

40

About Bank of America Corporation

Bank of America Merrill Lynch is the corporate and investment banking division of Bank of America. It provides services in mergers and acquisitions, equity and debt capital markets, lending, trading, risk management, research, and liquidity and payments management. It was formed through the combination of the corporate and investment banking activities of Bank of America and Merrill Lynch following the acquisition of the latter by the former in January 2009. Bank of America completed the acquisition of Merrill Lynch & Co on 1 January 2009. Bank of America began rebranding all of its corporate and investment banking activities under the Bank of America Merrill Lynch name in September 2009. In April 2010, Bank of America Merrill Lynch appointed Christian Meissner as head of investment banking for Europe, Middle East and Africa. In April 2011, Bank of America Merrill Lynch integrated its corporate and investment banking operations into a single division. In October 2013, Bank of America Merrill Lynch was recognised as the Most Innovative Investment Bank of the Year in The Banker's Investment Banking Awards.

Bank of America Corporation Careers

Join the dynamic team at Bank of America Corporation, a premier global financial institution where innovation, leadership, and growth go hand in hand. As one of the largest banks in the world, we offer unparalleled job opportunities and a culture that values diversity, inclusion, and professional growth. Work You’ll Do At Bank of America Corporation, you’ll be part of a team that’s dedicated to making a real difference. Whether you’re helping families buy their first home, advising businesses on expansion, or developing cutting-edge financial technologies, your work will have an impact. Our commitment to leadership in the financial industry has never been stronger, and we need passionate, skilled professionals to lead our journey. Explore a World of Opportunities From entry-level positions to leadership roles, Bank of America Corporation offers a variety of career paths in areas such as investment banking, technology, marketing, and risk management. Our job opportunities span the globe, providing the chance to work alongside the best in the industry and develop skills that will propel your career forward. Internship Programs Kickstart your career with Bank of America Corporation’s internship programs. These opportunities provide hands-on experience and a chance to engage in meaningful work that complements your academic studies. Interns gain invaluable networking opportunities, receive mentorship from seasoned professionals, and learn about the culture and operations of a global financial leader. Benefits and Growth Bank of America Corporation is committed to the well-being and continuous professional development of our team members. We offer a competitive benefits package that supports the health, financial stability, and work-life balance of our employees. Our training programs and development initiatives ensure that every team member has the opportunity to grow and advance within the company. Inclusive Culture We believe our strength lies in our diversity. Bank of America Corporation fosters an inclusive environment where all employees can thrive. Through diversity training and a commitment to equal opportunities, we cultivate leadership and innovation that reflect the wide-ranging communities we serve. Join Our Team Are you ready to advance your career at a company that’s at the forefront of the financial industry? Explore the positions available at Bank of America Corporation and find where your skills and interests align with our needs. We are continuously hiring and looking for individuals who are curious, creative, and eager to drive change. Stay Connected Keep up to date with the latest from Bank of America Corporation Careers by subscribing to our job alert emails. Tailor your subscription to receive updates that match your career interests and get insider tips that can help you during your application and interview process. Bank of America Corporation is not just a company—it’s a place where you can shape your future and the future of finance. Join us and be part of a team that’s redefining what a bank can be.
Learn more about Bank of America Corporation
Size
208,000 employees
Market Cap
$260.3 billion
Industry
Net Income
$17.8 billion
Founded
1998
5 Year Trend
-1.4%
NASDAQ

Similar Jobs

More Jobs at Bank of America Corporation

More Information Technology Jobs

Find similar Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP) jobs: