Early Warning Services

Principal Site Reliability Engineer - Paze

Early Warning Services$194K — $237K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 15+ years of experience in relevant technical disciplines
  • Proficiency in software development or scripting in modern programming languages
  • Strong understanding of software engineering principles and distributed systems
  • Experience with public cloud technologies, especially AWS
  • Excellent analytical, problem-solving, communication, and collaboration skills

Responsibilities

  • Improve service reliability, resilience, and scalability using software engineering and DevOps practices
  • Utilize data and engineering analysis to identify reliability risks and inform decisions
  • Define and enhance SLIs, SLOs, and other service-health metrics
  • Enhance observability through monitoring and alerting systems
  • Drive continuous improvement of CI/CD processes and operational readiness
  • Translate operational experiences into actionable improvements in engineering practices
  • Lead incident response efforts and foster a culture of blameless post-incident learning

Benefits

  • Comprehensive healthcare coverage including medical, dental, and vision plans
  • 401(k) retirement plan with a 100% company match on initial deferrals
  • Generous paid time off including flexible time for exempt employees
  • Paid parental leave of 12 weeks, with support for family planning
  • Encouraging work-life balance with a paid volunteer day
Full Job Description

Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.

Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.

Role Summary

The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle.

The roleoperatesat enterprise scope,establishingtechnical direction and applying evidence-driven engineering, technical rigor, sound judgment, automation, and broadsystemsexpertiseacross organizational boundaries.

Core Responsibilities

  • Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed,observed, operated, and recovered.

  • Use data, evidence, experimentation, and rigorous engineering analysisappropriate tothe level toidentifyreliability risks, test assumptions, and guide technical decisions.

  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.

  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.

  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.

  • Identifyrecurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.

  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout thedevelopmentlifecycle.

  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learningappropriate tothe level.

  • Provides enterprise-level technical leadership for critical production incidents andestablishesor influences engineering practices that improve incident response, escalation, servicerestorationand sustainable on-call operations across the organization.

  • Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.

Leveling Intent

Principalrepresentsdomain-level technical leadership and organizational impact. Deep individualexpertiseis expected, but Principal-level impact comes fromidentifyingsystemic risk,establishingtechnical direction,influencing engineering practices and architecture, and multiplying the capability of the broader engineering organization.

Level Expectations

  • Acts as an enterprise force multiplier, raising the effectiveness and technical capability of engineers and teams across the organization while building sustainable organizational capability rather than individual dependency.

  • Demonstrates software engineering, systems thinking, troubleshooting, andproductionreliability capabilities appropriate to the level.

  • Applies evidence-driven reasoning and technical rigor to distinguish observed facts from assumptions and make defensible engineering recommendations.

  • Shares knowledge and contributes to sustainable engineering capability rather than creating dependency on individualexpertise.

  • Operates with significant autonomy across the organization's most consequential reliability challenges.

  • Establishes enterprise technical direction, develops senior technical leaders, anddemonstratesimpact well beyond systems personally touched.

Minimum Qualifications

  • Typically15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.

  • Experience with software development or scripting using one or more modern programming languages.

  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observabilityappropriate tothe level.

  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architecturesappropriate tothe level.

  • Demonstrated analytical, problem-solving, communication, and collaboration skillsappropriate tothe scope of the role.

Preferred Qualifications

  • Hands-on experience with AWS is preferred, or comparable experience withanother major cloud platformsuch as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).

  • Experience developing, deploying,operating, or improvinghighly availableproduction software or distributed systems.

  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation.

  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readinessappropriate tothe level.

  • Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness.

  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.

The base pay scale for this position in:
Phoenix, AZ/ Chicago, IL / Washington, DC in USD per year is: $194,000 - $237,000.
New York, NY/ San Francisco, CA in USD per year is: $232,000 - $284,000.


Additionally, candidates are eligible for a discretionary incentive plan and benefits.

This pay scale is subject to change and is not necessarily reflective of actual compensation that may be earned, nor a promise of any specific pay for any specific candidate, which is always dependent on legitimate factors considered at the time of job offer. Early Warning Services takes into consideration a variety of factors when determining a competitive salary offer, including, but not limited to, the job scope, market rates and geographic location of a position, candidates education, experience, training, and specialized skills or certification(s) in relation to the job requirements and compared with internal equity (peers). The business actively supports and reviews wage equity to ensure that pay decisions are not based on gender, race, national origin, or any other protected classes.

Physical Requirements

Early Warning works together in a highly collaborative office environment.Working conditions consist of a normal office environment. Work is primarily sedentary and requires extensive use of a computer and involves sitting for periods of approximately four hours. Work may require occasional standing, walking, kneeling, and reaching. Must be able to lift 10 pounds occasionally and/or negligible amount of force frequently. Requires visual acuity and dexterity to view, prepare, and manipulate documents and office equipment including personal computers. Requires the ability to communicate with internal and/or external customers.

Employee must be able to perform essential functions and physical requirements of position with or without reasonable accommodation.

Candidates responding to this posting must independently possess the eligibility to work in the United States at the date of hire.

Some of the Ways We Prioritize Your Health and Happiness

  • Healthcare CoverageCompetitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.

  • 401(k) Retirement PlanFeaturing a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.

  • Paid Time Off Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.

  • 12 weeks of Paid Parental Leave

  • Maven Family Planning provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.

AndSOmuch more! We continue to enhance our program, so be sure tocheck our Benefits page herefor the latest. Ourteamcan share more during the interview process!

About Early Warning Services

Early Warning Services is a financial services company that provides fraud prevention and risk management solutions to banks, credit unions, and other financial institutions. The company was founded in 1990 and is headquartered in Scottsdale, Arizona. Early Warning Services offers a variety of products and services, including identity verification, account verification, and payment authentication. The company's solutions are designed to help financial institutions reduce fraud and improve the customer experience.
Learn more about Early Warning Services
Size
1,000 employees
Industry

Similar Jobs

More Jobs at Early Warning Services

More Information Technology Jobs

Find similar Principal Site Reliability Engineer - Paze jobs: