Job Description As a Reliability Engineer, your role will be a combination of supporting production applications and proactively looking for ways to automate your discoveries, eliminate incidents from recurring and/or reduce the time it takes to get our customers back up and running. In addition, you'll focus on improving the following for our applications: availability, latency, performance, efficiency, and effective proactive monitoring. The reliability engineer interfaces with business users, development teams and system administrators to ensure systems perform to meet their business needs and specifications. Responsibilities include: Developing, coordinating, and conducting technical reliability studies on engineering designs to assess the likelihood that a product/process performs its intended function over the intended lifecycle. Measuring and analyzing the reliability of the design, materials, processes, cost, and final products of production. Recommending design or test methods and statistical process control procedures for achieving required levels of product reliability. Completing risk analysis studies of new designs and processes. Undertaking testing and analysis on failures, proposing changes in design or formulation to improve system and/or process reliability.
Basic Qualifications
- Bachelor's degree, or equivalent work experience
- Five to seven years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development
Preferred Skills/Experience
- Expertise in Site Reliability Engineering (SRE), or Reliability Engineering.
- Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring.
- Demonstrated ability to understand stakeholder needs and guide the development of reliability requirements for large, complex multi-system products.
- Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks.
- Proficiency with Datadog, , Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry.
- Experience building, standardizing, and tuning operational dashboards and actionable alertsthat communicate service health, customer impact, dependency health, performance trends, failure conditions, severity, ownership, routing, and runbook linkage.
- Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes.
- Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements.
- Excellent stakeholder management, communication, and technical leadership skills.
- Hands on experience with ServiceNow.
Location expectations
This role requires working from a U.S. Bank location three (3) or more days per week.
Benefits:
Our approach to benefits and total rewards considers our team members9 whole selves and what may be needed to thrive in and outside work. That's why our benefits are designed to help you and your family boost your health, protect your financial security and give you peace of mind. Our benefits include the following:
Healthcare (medical, dental, vision)
Basic term and optional term life insurance
Short-term and long-term disability
Pregnancy disability and parental leave
401(k) and employer-funded retirement plan
Paid vacation (from two to five weeks depending on salary grade and tenure)
Up to 11 paid holiday opportunities
Adoption assistance
Sick and Safe Leave accruals of one hour for every 30 worked, up to 80 hours per calendar year unless otherwise provided by law
Review our full benefits available by employment status here.
The salary range reflects figures based on the primary location, which is listed first. The actual range for the role may differ based on the location of the role. In addition to salary, U.S. Bank offers a comprehensive benefits package, including incentive and recognition programs, equity stock purchase 401(k) contribution and pension (all benefits are subject to eligibility requirements). Pay Range: $105,400.00 - $124,000.00