Truist Financial

Site Reliability Engineer

Truist Financial$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor’s degree in Computer Science or related field
  • 7+ years of professional experience in software development
  • Deep knowledge of software architecture and design principles
  • Expertise in software development lifecycle and security practices
  • Preferred: Advanced degree and professional certifications in technical disciplines
  • Strong experience in cloud-native architectures and DevOps
  • Proficiency with automation and scripting languages like Python and Go

Responsibilities

  • Lead major incident response efforts and drive resolution across teams
  • Implement automation solutions that improve service resilience and reduce downtime
  • Standardize observability practices and enhance telemetry coverage
  • Collaborate with cross-functional teams to embed reliability in designs
  • Develop and enforce incident response playbooks and operational documentation
  • Mentor junior SREs and provide technical guidance
  • Evaluate new technologies that enhance the reliability framework

Benefits

  • Professional growth opportunities and mentorship
  • Collaboration with senior engineers and cross-functional teams
  • Engagement in enterprise-wide initiatives for operational excellence
  • Access to advanced tools and technologies in SRE
  • Involvement in shaping company-wide reliability policies
Full Job Description
The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams. Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime. The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks. Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management. ESSENTIAL DUTIES AND RESPONSIBILITIES Following is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time. 1. Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results. 2. Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery. 3. Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities. 4. Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow. 5. Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives. 6. Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing. 7. Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities. 8. Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience. 9. Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision. Qualifications Required Qualifications The requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. 1. Bachelor’s degree in Computer Science, Software Engineering, or related field. 2. Minimum of 7 years of professional experience in software development. 3. Deep knowledge of multiple programming languages, software architecture, and design principles. 4. Deep understanding of software development lifecycle, testing, deployment, and security practices. Preferred Qualifications 1. Advanced degree in Computer Science or related technical discipline. 2. Professional certifications such as Certified Software Development Professional (CSDP) or equivalent. 3. Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps. 4. Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management. 5. 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations. 6. Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling. 7. Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible). 8. Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring. 9. Proven leadership in major incident management and cross-team technical coordination. 10. Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns. 11. Excellent communication skills, including executive-level situational awareness during critical incidents. 12. Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices. Preferred Qualifications Financial services or regulated industry experience. Experience enabling large-scale SRE transformations or modernization initiatives. Familiarity with chaos engineering, resilience assessments, and service failure modeling. Exposure to hybrid-cloud and multi-cloud operational frameworks. Experience contributing to or leading Center for Enablement functions or Communities of Practice. Key Responsibilities Incident & Problem Management Leadership Lead major and high-severity incident response efforts,focusing on diagnosing technical rootcausestherein, and driving multi-team technical resolution. Drive problem management to closure, ensuring systemic fixes replace recurring operational risks. Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks. Reliability Engineering & Automation Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Implement intelligent alerting, anomaly detection, and event correlation leveragingAI andAIOps tools. Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization. Observability & Operational Excellence Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk. Define and standardize enterprise observability practices, dashboards, and KPIs. Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation. Cross-Functional Leadership & Influence Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution. Act as a change agent to elevate operational maturity and drive transformative improvements acrossWholesale. Lead workshops, maturity assessments, and enablement sessions through the SRE C4E and Communities of Practice. Standardization & Documentation Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns. Contribute to enterprise SRE frameworks, templates, and maturity models. Promote consistent adoption of best practices across domains and lines of business. Mentorship & Technical Development Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline. Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering. Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.

About Truist Financial

Truist Financial Careers

Join the dynamic team at Truist Financial, a leader in the financial services sector, and propel your career to new heights. At Truist Financial, we offer more than just job opportunities; we provide a platform for professional growth and innovation in an environment that values diversity and leadership.

Why Truist Financial?

At Truist Financial, we are committed to building a diverse and inclusive workplace where every team member is empowered to contribute their unique skills and perspectives. We believe that our strength lies in our diversity, and we are dedicated to fostering a culture that embraces the differences that make each of us unique.

Explore a World of Opportunities

Whether you're seeking your first internship or a seasoned professional looking to advance your career, Truist Financial offers a range of employment opportunities across various disciplines. Our team is growing, and we are constantly looking for talented individuals who are eager to make an impact.

Innovate and Lead

Join us and be part of a culture of innovation where your ideas can help shape the future of banking. At Truist Financial, you’ll work alongside industry leaders and have access to cutting-edge resources that foster continuous professional development and innovation.

Develop Your Career

Truist Financial is deeply invested in the career progression of our employees. We offer robust training programs, including leadership development and diversity training, to ensure you have the tools needed to succeed. Our commitment to your growth is reflected in our comprehensive benefits package, designed to support you both professionally and personally.

Networking and Professional Development

Enhance your professional network and connect with like-minded colleagues through our various networking events and community engagement initiatives. At Truist Financial, we believe in the power of connections and the impact they can have on your career.

Join Our Team

Ready to take the next step in your career? Explore the current job openings at Truist Financial. We are hiring across multiple departments, looking for passionate, curious, and innovative individuals to join our team. Check out our available positions and find the one that best matches your skills and interests.

Prepare for Your Interview

Make a great first impression. Visit our Careers page for tips on how to craft a compelling resume and succeed in your interview at Truist Financial. We are excited to see how you can contribute to our team and help us drive the future of banking.

Stay Connected

Don’t miss out on future opportunities or insights into our company culture and industry trends. Subscribe to our job alert emails and stay informed about new positions and career tips directly from our professionals. At Truist Financial, we’re not just offering jobs; we’re building careers. Join us and discover how you can make a difference and fuel your future.

SEARCH TRUIST FINANCIAL JOBS

READ CAREERS BLOG

Learn more about Truist Financial
Size
50,283 employees
Market Cap
$56.6 billion
Industry
5 Year Trend
+14.3%
NASDAQ

Similar Jobs

More Jobs at Truist Financial

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: