Platform Reliability & Automation Analyst

Vanguard Group, Inc.

• $88K — $105K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in IT, Computer Science, Engineering, Business, or related field; equivalent experience considered.
  • Preferred training in systems administration, cloud technologies, automation, and IT operations.
  • Certifications in Microsoft, AWS, ITIL, observability, automation, or cybersecurity preferred.
  • Strong grasp of technology operations, system reliability, and operational support practices.
  • Experience with monitoring, logging, reporting, and service management platforms.

Responsibilities

  • Evaluate health, availability, and performance of technology services for improved visibility.
  • Support incident triage and troubleshooting through analysis of alerts, logs, and telemetry data.
  • Develop and enhance operational runbooks, recovery procedures, and operational readiness standards.
  • Leverage automation and AI tools like Microsoft Copilot to enhance productivity and reduce manual effort.
  • Support disaster recovery and operational readiness efforts; identify and recommend actions to mitigate risks.
  • Provide Help Desk and operational support coverage during high-demand periods or colleague absences.
  • Maintain operational documentation and engage in continuous improvement initiatives.

Benefits

  • Opportunities for professional development through training and coursework.
  • Engagement with innovative tools like Microsoft Copilot and automation technologies.
  • Work within a dynamic and collaborative team environment focused on operational excellence.
  • Involvement in continuous improvement initiatives that enhance service delivery.
Full Job Description
JOB SUMMARY

Supports the reliability, resilience, performance, and continuous improvement of Vanguard Charitable's technology services. The role combines operational engineering, monitoring, automation, technical incident response, and service support responsibilities while helping the organization mature toward a more proactive and data-driven operating model. Working closely with service owners, engineering teams, vendors, and the IT Service Management & Support Lead, the analyst helps improve observability, reduce operational risk, strengthen disaster recovery readiness, support technical troubleshooting, and automate repetitive operational activities. The position also serves as a catalyst for operational innovation by applying Microsoft Copilot, Claude, automation technologies, and modern operational practices to improve productivity, reduce manual effort, strengthen knowledge sharing, and enhance business outcomes.

ESSENTIAL JOB FUNCTIONS

  • Evaluate the health, availability, performance, and reliability of technology services; develop and improve monitoring, alerting, dashboards, observability practices, and operational reporting that improve visibility and early issue detection. (20%)
  • Support incident triage, troubleshooting, technical investigation, and service restoration activities by analyzing alerts, logs, telemetry, configuration data, and system dependencies in partnership with engineering teams, vendors, and service owners. (15%)
  • Develop, maintain, and improve operational runbooks, recovery procedures, synthetic testing, resilience validation activities, and technical readiness standards that improve supportability and operational consistency. (15%)
  • Leverage Microsoft Copilot, Claude, automation tools, scripting, and workflow technologies to improve productivity, documentation, reporting, troubleshooting, knowledge management, and operational effectiveness; develop reusable prompts, templates, and automation solutions that reduce manual effort across Platform Operations. (20%)
  • Support disaster recovery, resilience testing, recovery validation, capacity planning, and operational readiness efforts; identify reliability risks and recommend actions that strengthen recovery capabilities and service continuity. (10%)
  • Provide secondary Help Desk and operational support coverage, support service-request fulfillment, assist with escalation triage, maintain support knowledge, and help ensure continuity during colleague absences, periods of elevated demand, and business-critical events. (15%)
  • Maintain operational documentation, dependency records, support information, and configuration data while participating in special projects and continuous-improvement initiatives. (5%)


QUALIFICATIONS

Education & Training

  • Bachelor's degree in Information Technology, Computer Science, Engineering, Business, or a related discipline required; equivalent experience may be considered.
  • Training or coursework related to systems administration, cloud technologies, platform operations, automation, reliability engineering, AI, or IT operations preferred.


Licensure & Certification

  • Microsoft, AWS, ITIL, observability/monitoring, automation, cybersecurity, AI, or related certifications preferred.


Knowledge, Skill, and Ability

  • Strong understanding of technology operations, system reliability, monitoring, troubleshooting, and operational support.
  • Working knowledge of observability, incident response, automation, disaster recovery, and operational resilience concepts.
  • Experience using monitoring, logging, reporting, and service management platforms.
  • Ability to analyze technical issues, identify patterns, and recommend practical improvements.
  • Experience leveraging Microsoft Copilot, Claude, automation tools, scripting, or similar technologies to improve productivity and operational effectiveness.
  • Strong written and verbal communication skills with the ability to explain technical concepts to technical and non-technical audiences.
  • Curious, self-directed learner with a passion for continuous improvement, automation, and operational innovation.
  • Strong organizational skills with the ability to manage multiple priorities and operate effectively during incidents and business-critical events.


Experience

  • Four or more years of experience in technology operations, technical support, systems administration, platform operations, reliability engineering, or related disciplines.
  • Experience supporting production technology services and participating in incident resolution activities.
  • Experience developing operational documentation, runbooks, monitoring solutions, dashboards, or automation capabilities.
  • Experience working with vendors, managed service providers, and cross-functional technology teams preferred.
  • Experience applying AI capabilities, automation, or workflow improvements to real business or operational challenges preferred.
  • Experience participating in disaster recovery, resilience testing, platform monitoring, or operational readiness activities preferred.

Similar Jobs

More Jobs at Vanguard Group, Inc.

More Information Technology Jobs

Find similar Platform Reliability & Automation Analyst jobs: