The Goldman Sachs Group, Inc

Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group, Inc$150K — $200K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of hands-on experience in Site Reliability Engineering with enterprise-level system management.
  • Proven programming proficiency in languages like Java, Python, or Go for building scalable software.
  • Extensive experience with cloud platforms (AWS, GCP) and deep knowledge of container technologies (Docker, Kubernetes).
  • Mastery of Infrastructure as Code (IaC) tools (Terraform, CloudFormation) and configuration management (Puppet, Chef, Ansible).
  • Advanced experience in Prompt Engineering and Retrieval-Augmented Generation (RAG) for SRE workflows.
  • Strong background in Linux internals, networking, distributed systems, and performance tuning.
  • Excellent problem-solving and analytical skills, capable of tackling complex technical challenges.

Responsibilities

  • Drive the strategic focus on availability, scalability, and performance of critical applications and services.
  • Lead the design and implementation of resilient infrastructure and application architectures.
  • Develop advanced platforms and automation solutions to enhance operational workflows.
  • Manage critical incident responses and perform root cause analyses for systemic issues.
  • Embed reliability in application design through partnership with development teams and lead capacity planning.
  • Implement advanced monitoring and logging strategies for actionable application insights.
  • Provide technical vision and mentorship to senior engineers, ensuring best practices in project execution.

Benefits

  • Access to advanced training and development programs for continuous learning.
  • Opportunity to work with cutting-edge technology and innovative practices.
  • Engagement in a collaborative and dynamic work environment.
  • Participation in a global on-call leadership team that supports critical system incidents.
Full Job Description
Job Description

Site Reliability Engineer - Vice President

Site Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run scalable, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for improving the availability and reliability of the firm's most critical platform services and ensures they meet the requirements of our internal and external users. It is also responsible for firmwide policies and standards focused on firm's digital resilience. We are looking for engineers who are motivated to collaborate with our businesses to build and run sustainable production systems, which can evolve and adapt to changes in our fast-paced, global business environment.

The SRE team develops and maintains platforms and tools which help other Engineering teams in Goldman Sachs to build and operate reliable and resilient systems. These systems span on-premises datacenters and multiple public cloud environments. The platforms we offer include central logging, monitoring, agents and alerting and we provide tools to drive adoption and improvements to capacity planning, operational readiness assessments, production incident postmortems, SLIs / SLOs, and deployment automation including canary releases.

The products and services we provide to our internal customers are used by thousands of engineers every day. We believe that reliability is the most important feature of any system, and we are devoted to giving our engineers the platforms and tools they need to build and operate reliable products.
Role Overview
As a Site Reliability Engineer (SRE) at Goldman Sachs, you will be a pivotal leader in ensuring the availability, reliability, and scalability of the firm's most critical platform applications and services. You will combine deep software and systems engineering expertise to architect, build, and run large-scale, massively distributed, fault-tolerant systems. This role involves providing technical leadership, mentoring senior engineers, and collaborating closely with internal teams and executive stakeholders to build and operate sustainable production systems that can adapt to our dynamic global business environment. You will drive a culture of continuous improvement, championing the adoption of advanced SRE principles and best practices across the organization.

Responsibilities
  • Strategic Reliability & Performance: Drive the strategic direction for availability, scalability, and performance of mission-critical applications and platform services, ensuring alignment with firm-wide objectives.
  • Architectural Leadership: Lead the design, build, and implementation of highly available, resilient, and scalable infrastructure and application architectures.
  • Advanced Automation & Tooling: Architect and develop sophisticated platforms, tools, and automation solutions to eliminate toil, optimize operational workflows, and enhance deployment processes across the enterprise.
  • Complex Incident Management & Post-Mortem Analysis: Lead critical incident response, conduct in-depth root cause analysis for systemic issues, and implement long-term preventative measures to significantly enhance system stability and resilience.
  • System Design & Capacity Planning: Partner with development teams to embed reliability into application design from inception, provide expert system design consulting, and lead comprehensive capacity planning initiatives for future growth.
  • Observability & Insights: Define and implement advanced monitoring, high volume logging with multi-user query capabilities, and tracing strategies to provide deep, actionable insights into application performance, infrastructure health, and user experience.
  • Technical Vision & Mentorship: Provide technical vision, lead complex technical projects, conduct rigorous code reviews, enforce SDLC best practices, and actively mentor and develop senior and staff-level engineers.
  • Technology Evaluation & Adoption: Stay at the forefront of industry trends and advancements, evaluating and integrating cutting-edge tools and frameworks to significantly improve operational efficiency and reliability.
  • On-Call Leadership: Participate in and lead on-call rotations, providing expert guidance and hands-on support for critical system incidents.
Qualifications
  • Experience: Minimum of 6+ years of hands-on experience in Site Reliability Engineering, with a proven track record in architecting, designing, building, and maintaining highly available, scalable, and fault-tolerant systems at an enterprise level.
  • Technical Proficiency:
    • Exceptional programming skills in one or more major languages such as Java, Python, Go with a focus on building robust, scalable software.
    • Extensive hands-on experience with cloud platforms (e.g., AWS, GCP) and deep expertise in containerization and orchestration technologies (e.g., Docker, Kubernetes).
    • Mastery of Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation) and configuration management tools (e.g., Puppet, Chef, Ansible).
    • Advanced proficiency in Prompt Engineering and Retrieval-Augmented Generation (RAG) architectures to automate complex SRE workflows, such as the generation of Infrastructure as Code (IaC), dynamic runbooks, and incident response summaries.
    • Profound understanding of Linux internals, networking, distributed systems, and advanced system performance tuning.
    • Expertise in designing and implementing comprehensive monitoring, alerting, logging and tracing solutions (e.g., Prometheus, Grafana, ELK stack, Datadog, PagerDuty).
    • Deep experience with CI/CD tools and practices (e.g., Jenkins, GitLab, Maven).
    • Strong foundation in databases and distributed systems.
    • Exceptional problem-solving abilities and analytical skills, with a track record of resolving complex technical challenges.
  • Preferred Experience:
    • Experience with Distributed Databases like Elastic Search
    • Experience with working on GCP Big Query
    • Experience with messaging Systems Like Kafka
  • Education: Advanced degree (Bachelor's or Mas ter's or PhD) in Computer Science or a related technical field involving coding and/or systems engineering, or equivalent practical experience.
  • Soft Skills: Superior communication, collaboration, and interpersonal skills, with the ability to influence technical direction, lead cross-functional initiatives, and effectively engage with global teams and executive leadership. Proven ability to work independently, manage multiple complex stakeholders, and drive significant organizational change.

About The Goldman Sachs Group, Inc

The Goldman Sachs Group, Inc. provides investment banking, securities, and investment management services, as well as financial services to corporations, financial institutions, governments, and high-net-worth individuals worldwide. Its Investment Banking segment offers financial advisory services, including advisory assignments concerning mergers and acquisitions, divestitures, corporate defense, risk management, and restructurings and spin-offs; and underwriting services comprising public offerings and private placements of a range of securities, loans, and other financial instruments, and derivative transactions. The company’s Institutional Client Services segment provides client execution services, such as fixed income, currency, and commodities client execution related to making markets in interest rate products, credit products, mortgages, currencies, and commodities; and equities related to making markets in equity products, as well as executes and clears institutional client transactions on stock, options, and futures exchanges. This segment also engages in the securities services business providing financing, securities lending, and other brokerage services to institutional clients, including hedge funds, mutual funds, pension funds, and foundations. Its Investing and Lending segment originates longer-term loans; and invests in debt securities, loans, public and private equity securities, real estate, consolidated investment entities, distressed assets, currencies, commodities, and power generation facilities. The company’s investment management segment provides investment products and services, as well as offers wealth advisory services, including portfolio management and financial counseling, and brokerage and other transaction services.

The Goldman Sachs Group, Inc. Careers

Join the prestigious team at The Goldman Sachs Group, Inc., a global leader in finance and investments, and propel your career to new heights. Our firm is renowned for its commitment to excellence, innovation, and leadership in the financial sector.

Work You’ll Do

At Goldman Sachs, you will be part of a dynamic environment that fosters growth, diversity, and professional development. Engage in transformative projects that redefine the landscape of global finance, driven by a culture of high performance and continuous improvement.

Lead with Innovation and Leadership

Step into a role where your skills will be honed and your leadership capabilities enhanced. The Goldman Sachs Group, Inc. is at the forefront of merging financial expertise with technological innovation, creating a platform for you to lead impactful initiatives.

Join Our Diverse and Inclusive Team

Diversity and inclusion are at the core of our company culture. With employees from various backgrounds, Goldman Sachs thrives on the rich ideas and perspectives that this diversity brings. Here, every team member’s contribution is valued, and every voice is heard.

Explore Job Opportunities and Internships

Whether you’re seeking full-time employment or looking into internship opportunities, Goldman Sachs offers a range of positions that cater to different skills, experiences, and career aspirations. From analyst to executive positions, discover how you can contribute to our legacy of success.

Networking and Professional Growth

Goldman Sachs is not just a workplace. It is a community where you can build lasting relationships with colleagues and industry leaders. Our networking events, mentorship programs, and professional groups foster connections that can accelerate your career trajectory.

Benefits and Career Development

Invest in your future with Goldman Sachs’ unparalleled employee benefits and career development programs. From comprehensive health benefits to personalized career coaching and leadership training, we ensure that our team is equipped for success, both professionally and personally.

Prepare for Your Future

Ready to take the next step? Prepare your resume, sharpen your interview skills, and explore the vast array of job opportunities at The Goldman Sachs Group, Inc. Our hiring process is designed to identify and nurture talent, helping you to achieve your career goals.

Stay Connected

Keep up to date with the latest from Goldman Sachs Careers by subscribing to our job alert emails. Tailor your preferences to receive updates that match your career interests and stay ahead in the competitive financial sector.

Join The Goldman Sachs Group, Inc.

Embark on a rewarding journey with The Goldman Sachs Group, Inc. where your potential is limitless. Explore positions, read about our company culture, and apply today to become part of a team that values growth, leadership, and innovation.

SEARCH GOLDMAN SACHS JOBS

READ CAREERS BLOG

At The Goldman Sachs Group, Inc., your career is just the beginning – it’s an opportunity to excel and lead in the global financial industry. Join us and make your mark!
Learn more about The Goldman Sachs Group, Inc
Size
45,100 employees
Market Cap
$115.8 billion
Industry
Net Income
$9.4 billion
Founded
1869
5 Year Trend
+11.3%
Revenue
$53.4 billion
NASDAQ

Similar Jobs

More Jobs at The Goldman Sachs Group, Inc

More Information Technology Jobs

Find similar Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas jobs: