Engineer, Site Reliability

Vanguard Group, Inc.

$112K — $135K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Minimum of eight years of relevant experience, including two years in development.
  • Undergraduate degree or equivalent experience; graduate degree preferred.
  • Experience in designing and operating production-facing platforms with measurable reliability outcomes.
  • Deep knowledge of distributed systems architecture and operational resilience.
  • Strong technical leadership skills for driving architecture decisions and aligning teams on engineering direction.
  • Proficient in Java or JavaScript for developing in cloud-native environments.
  • Hands-on experience with AWS and automation using Python or similar languages.

Responsibilities

  • Lead the technical strategy and evolution of reliability engineering platforms across numerous applications.
  • Design and build software that enhances reliability outcomes with automated detection and remediation.
  • Drive observability capabilities for consistent metrics and insights in cloud-native technology stacks.
  • Define resilient patterns for applications to ensure graceful degradation and automated recovery.
  • Influence engineering teams to incorporate reliability from the design phase of projects.
  • Resolve complex production issues by identifying risks and engineering durable solutions.
  • Engage in special projects and additional duties as required.

Benefits

  • Opportunity for leadership in shaping the technical landscape of reliability engineering.
  • Collaborative environment influencing multiple teams and technology leaders.
  • Exposure to cutting-edge technologies in cloud-native and microservices architecture.
  • Focus on measurable outcomes, allowing impact on operational performance.
  • Engagement in special projects that may enhance skills and visibility within the organization.
Full Job Description
Core Responsibilities
  • Lead the technical strategy, architecture, and evolution of PWTech reliability engineering platforms and capabilities, ensuring they scale across hundreds of applications and critical client-facing systems.
  • Design and build production-grade software and platforms that improve reliability outcomes, including automated incident detection, diagnostics, remediation, resiliency engineering, and operational intelligence.
  • Drive enterprise observability and diagnostics capabilities, enabling consistent telemetry, distributed tracing, metrics, and operational insights across cloud-native technologies and application stacks.
  • Define and codify resilient application and platform patterns, such as graceful degradation, circuit breakers, load shedding, fault isolation, failover, and automated recovery, driving adoption through reusable software, frameworks, and engineering standards.
  • Influence engineering teams and technology leaders across the organization, establishing technical standards and ensuring reliability is designed into systems from inception.
  • Lead the resolution of complex reliability and production challenges, identifying systemic risks, driving root cause analysis, and engineering durable solutions that improve long-term resilience.
  • Participates in special projects and performs other duties as assigned.


Qualifications
  • Minimum of eight years related experience, with at least two years of development experience.
  • Undergraduate degree or equivalent combination of training and experience. Graduate degree preferred.


Preferred Skills
  • Experience designing, building, and operating production-facing platforms or engineering capabilities that achieve broad adoption and deliver measurable reliability, operational, or business outcomes.
  • Deep expertise in distributed systems architecture, including scalability, availability, resiliency, fault tolerance, performance optimization, and production operations at scale.
  • Strong technical leadership and influence skills, with a demonstrated ability to drive architecture decisions, establish technical standards, and align multiple teams on engineering direction.
  • Deep expertise in Java or JavaScript, with hands-on experience developing and operating software in modern cloud-native and microservices environments.
  • Demonstrated ability to diagnose and resolve complex production issues, perform root cause analysis, and engineer durable solutions that prevent recurrence.
  • Hands-on experience with AWS and modern cloud architecture patterns.
  • Experience with observability and telemetry platforms, including metrics, logging, distributed tracing, and production diagnostics. Experience with OpenTelemetry is strongly preferred.
  • Proficiency with Python or similar scripting languages to develop automation, tooling, and operational workflows.
  • Strong software engineering fundamentals, systems thinking skills, experience solving complex reliability challenges, and the ability to influence across teams and drive engineering best practices

Special Factors

Sponsorship
Vanguard is not offering visa sponsorship for this position.

Similar Jobs

More Jobs at Vanguard Group, Inc.

More Information Technology Jobs

Find similar Engineer, Site Reliability jobs: