Site Reliability Engineer

OneAZ Credit Union

$96K — $120K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • High School Diploma required; Bachelor's degree in IT, Computer Science, or related field preferred.
  • 5-8 years of experience in enterprise infrastructure and monitoring platforms.
  • Deep knowledge of networking, cloud technologies, Windows Server, and virtualization.
  • Familiarity with disaster recovery and business continuity initiatives.
  • Experience with SolarWinds in a large enterprise setting is preferred.
  • Strong troubleshooting, analytical, and problem-solving skills.

Responsibilities

  • Administer and optimize enterprise monitoring platforms, including SolarWinds.
  • Develop monitoring dashboards and performance metrics to ensure system health.
  • Identify and mitigate infrastructure reliability issues proactively.
  • Lead disaster recovery planning and testing activities.
  • Collaborate with technology teams to enhance system resiliency and operational readiness.
  • Investigate and resolve complex incidents while conducting root cause analyses.
  • Maintain thorough documentation of recovery processes and disaster recovery exercises.

Benefits

  • Generous paid time off including holidays and personal days.
  • Low-cost Medical, Dental & Vision plans.
  • Paid childcare assistance and tuition reimbursement.
  • Award-winning 401K plan and student loan repayment options.
  • Gym fee reimbursement and additional perks as detailed in the Benefits Booklet.
Full Job Description
What You9ll Do This position will be located at our Corporate Office: 2355 W Pinnacle Peak Rd, Phoenix, AZ 85027

The Site Reliability Engineer is responsible for ensuring the reliability, availability, performance, and recoverability of OneAZ9s technology platforms and infrastructure. This position combines infrastructure engineering, automation, monitoring, and resiliency practices to maintain highly available systems supporting associates and members. The engineer designs and implements automation solutions using PowerShell and other scripting technologies, administers enterprise monitoring platforms, leads disaster recovery testing activities, and partners with technology teams to improve operational resilience, service reliability, and recovery readiness across on-premises and cloud environments. This role serves as a key contributor to incident response, infrastructure modernization, and continuous improvement initiatives focused on reducing operational risk and improving system uptime.

The Site Reliability Engineer works closely with Infrastructure, Information Security, Application Support, Enterprise Architecture, and business teams to identify operational risks, strengthen recovery capabilities, and improve the overall resilience of technology services that support associates and members.

Essential Functions
  • Administer, maintain, and optimize SolarWinds and other enterprise monitoring platforms.
  • Develop and maintain monitoring dashboards, alerts, reports, and performance metrics.
  • Proactively monitor and improve infrastructure health, availability, capacity, and performance, identifying and mitigating potential issues before they impact service reliability or user experience.
  • Proactively identify infrastructure risks, reliability concerns, and opportunities for improvement.
  • Lead disaster recovery planning, testing, failover exercises, and recovery validation activities.
  • Develop and maintain disaster recovery documentation, runbooks, recovery procedures, and test plans.
  • Coordinate and execute system failover and failback activities for critical applications and infrastructure.
  • Partner with Infrastructure, Security, and Application teams to improve system resiliency and operational readiness.
  • Support business continuity planning initiatives and technology recovery efforts.
  • Lead the investigation and resolution of complex infrastructure and service reliability incidents, conduct root cause analyses (RCAs), identify underlying systemic issues, and drive corrective and preventive actions to improve availability, performance, resiliency and operational excellence.
  • Track, document, and report on disaster recovery testing results, remediation activities, and recovery readiness metrics.
  • Maintain infrastructure diagrams, recovery documentation, and operational procedures.
  • Support audits, examinations, and compliance activities related to disaster recovery, resiliency, and infrastructure operations.
  • Research and recommend technologies and best practices that improve reliability, monitoring, and recoverability.
  • Participate in infrastructure maintenance activities, upgrades, and projects.
  • Develop and maintain automation scripts using PowerShell, Python, Bash, or similar tools to streamline infrastructure operations, automate remediation of common issues, and enhance service availability and performance.

What You Bring
  • High School Diploma Required
  • Bachelor9s Degree in Information Technology, Computer Science, Information Systems, Engineering, or a related technical field; or equivalent combination of education and experience. Required
  • 5-8 years similar or related experience of experience supporting enterprise infrastructure environments and administering infrastructure monitoring platforms.. Required
  • Deep knowledge of networking and cloud technologies, Windows Server, virtualization, and storage with hands on experience supporting and optimizing enterprise infrastructure environments.
  • Experience with disaster recovery testing, failover planning, or business continuity initiatives.
  • Experience administering SolarWinds in a large enterprise environment. Preferred
  • Experience within a financial institution, credit union, or regulated industry. Preferred
  • Experience supporting hybrid cloud environments including Microsoft Azure. Preferred
  • Experience coordinating disaster recovery exercises and recovery readiness assessments. Preferred
  • Familiarity with ITIL service management principles. Preferred
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Ability to manage multiple priorities and projects in a fast-paced environment.
  • Strong knowledge of infrastructure monitoring and alerting best practices.
  • Understanding of disaster recovery, business continuity, and operational resiliency concepts.
  • Ability to analyze system performance trends and identify potential issues before business impact occurs.
  • Strong technical documentation and organizational skills.
  • Ability to coordinate activities across multiple technology teams.
  • Strong verbal and written communication skills.
  • Ability to remain organized and effective during critical incidents and recovery activities.
  • Commitment to continuous improvement and operational excellence.
  • Industry certifications: SolarWinds Certified Professional (SCP), Microsoft Azure Administrator Associate, Microsoft Azure Solutions Architect, Google Professional Site Reliability Engineer, ITIL Foundation, CBCP, or related certifications.

Compensation & Benefits
  • Generous paid time off: paid holidays, floating holidays, personal days, vacation days, plus sick time
  • Low-cost Medical, Dental & Vision plans
  • Paid childcare assistance
  • Award-winning 401K
  • Gym fee reimbursement
  • Tuition Reimbursement
  • Student loan repayment
  • ...and much more. Explore all the details in our comprehensive Benefits Booklet
  • Target hiring range $96,664.36 - $120,830.45 3 (Depending on experience and prior to any incentives this position is eligible for)

Similar Jobs

More Jobs at OneAZ Credit Union

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: