Senior Infrastructure Engineer, SRE

Rocket Money

$150K — $185K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in cloud or infrastructure engineering with a focus on reliability and production operations
  • Experience defining SLIs and SLOs for production services with tangible outcomes
  • Proficient in observability platforms; preference for Datadog
  • Coding skills in Python, Go, TypeScript, or similar for tooling and automation
  • Hands-on experience with Terraform in AWS and troubleshooting production issues
  • Background in developing and executing disaster recovery plans
  • Past on-call experience for services developed personally, with insights on alert management

Responsibilities

  • Build and enhance the reliability and resiliency of systems and services
  • Establish and regularly review SLIs, SLOs, and error budgets for critical services
  • Own and advance the disaster recovery strategy with clear objectives and exercises
  • Collaborate with product engineering teams to enable service ownership and user-focused metrics
  • Evolve the observability platform and set standards across metrics, tracing, and logs
  • Enhance the incident response process through tuning alerts and maintaining documentation
  • Contribute to daily Cloud Infrastructure tasks while focusing on reliability initiatives

Benefits

  • Health, Dental & Vision Plans
  • 401k Matching
  • Unlimited PTO
  • Daily lunch (in-office only)
  • Snacks & Coffee (in-office only)
  • Commuter benefits (in-office only)
Full Job Description


We're looking to expand our Cloud Infrastructure team with a Senior Infrastructure Engineer, SRE to lead the reliability and operational evolution of our platform. We run hundreds of services in production, which enable us to process billions of transactions, consume multiple terabytes of data, and produce hundreds of millions of logs per day, and our reliability practice needs to evolve to match our growing scale. This includes:
  • Building and improving the reliability and resiliency of our systems and services
  • Establishing SLIs, SLOs, and error budgets for our most critical services and user journeys, and reviewing them regularly with the teams that own them
  • Owning and evolving our disaster recovery strategy: recovery objectives, failover and restore paths, and regular exercises that prove they work
  • Partnering with product engineering teams so they can own and operate their own services, with metrics that reflect real user experience
  • Evolving our observability platform and standards across metrics, tracing, and logs: including instrumentation paved roads, alert quality, and observability cost
  • Strengthening our incident practice: tuning paging thresholds, keeping runbooks current, and following through on postmortem action items
  • Contributing to day-to-day Cloud Infrastructure work alongside your reliability specialty - infrastructure build-outs, platform backlog, and a shared on-call rotation (1 week out of every 6 weeks)

You'll join the Cloud Infrastructure team and partner with engineering and internal support teams to drive this work.

We support millions of people to improve their financial lives, and this role ensures we can continue to do so reliably and at scale.

ABOUT YOU
  • You have 5+ years of hands-on cloud or infrastructure engineering experience, with substantial time spent on reliability and production operations at scale
  • You have defined SLIs and SLOs for real production services, and can talk about what changed as a result. What got fixed, what got deprioritized, and what you got wrong the first time
  • You have hands-on experience with an observability platform in production; Datadog strongly preferred
  • You're comfortable writing code (Python, Go, TypeScript, or similar) for internal tooling, production debugging, and automation
  • You write production Terraform and are comfortable in AWS, and when production breaks you can find the problem and fix it
  • You have built or operated a disaster recovery plan: you set the recovery goals, wrote the failover and restore steps, and ran the drills that proved it works
  • You have been on-call for services you helped build, and you have opinions about what makes an alert worth waking someone for
  • You prefer giving teams paved roads and good defaults over mandates, so they can own their own instrumentation
BONUS POINTS
  • You have led a reliability or observability modernization project where you defined the vision, approach, and delivered the implementation
  • You have built internal tooling, libraries, or instrumentation standards that made it easier for other teams to operate their services well
  • You have run game days, chaos experiments, or DR exercises, and fixed the problems they uncovered
  • You have cut observability spend while keeping the coverage you needed


WE OFFER
  • Health, Dental & Vision Plans
  • Competitive Pay
  • 401k Matching
  • Unlimited PTO
  • Lunch daily (in-office only)
  • Snacks & Coffee (in-office only)
  • Commuter benefits (in-office only)

Additional information: Salary range of $150,000 - $185,000/year + bonus + benefits. Base pay offered may vary depending on job-related knowledge, skills, and experience.

This job description is an outline of the primary responsibilities of this position and may be modified at the discretion of the company at any time.

Similar Jobs

More Jobs at Rocket Money

More Information Technology Jobs

Find similar Senior Infrastructure Engineer, SRE jobs: