Skip to content
← Back to job listings

Senior Infrastructure Engineer (SRE)

Rocket Money · Washington, United States

RemoteImported listingfull-time8 days ago

About The Role

Join our Cloud Infrastructure team as a Senior Infrastructure Engineer (SRE) to lead the reliability and operational evolution of our platform. You will be responsible for building and improving the reliability and resiliency of our systems and services, establishing SLIs, SLOs, and error budgets, owning and evolving our disaster recovery strategy, and partnering with product engineering teams. You will also contribute to day-to-day Cloud Infrastructure work and ensure we can continue to support millions of people in improving their financial lives reliably and at scale.

  • Lead the reliability and operational evolution of the platform, ensuring systems and services are resilient and reliable.
  • Establish and review Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services.
  • Own and evolve the disaster recovery strategy, including recovery objectives, failover and restore paths, and regular exercises.
  • You prefer giving teams paved roads and good defaults over mandates, so they can own their own instrumentation
  • You have been on-call for services you helped build, and you have opinions about what makes an alert worth waking someone for
  • You have defined SLIs and SLOs for real production services, and can talk about what changed as a result. What got fixed, what got deprioritized, and what you got wrong the first time
  • You write production Terraform and are comfortable in AWS, and when production breaks you can find the problem and fix it
  • You have built or operated a disaster recovery plan: you set the recovery goals, wrote the failover and restore steps, and ran the drills that proved it works
  • You're comfortable writing code (Python, Go, TypeScript, or similar) for internal tooling, production debugging, and automation
  • You have 5+ years of hands-on cloud or infrastructure engineering experience, with substantial time spent on reliability and production operations at scale
  • You have hands-on experience with an observability platform in production; Datadog strongly preferred
  • You have led a reliability or observability modernization project where you defined the vision, approach, and delivered the implementation
  • You have built internal tooling, libraries, or instrumentation standards that made it easier for other teams to operate their services well
  • You have run game days, chaos experiments, or DR exercises, and fixed the problems they uncovered
  • You have cut observability spend while keeping the coverage you needed

This is an external listing. JobSpring does not represent or verify the employer. Report this listing