Skip to content
← Back to job listings

Senior Site Reliability Engineer

Spotme · United Kingdom

RemoteImported listingfull-time4 days ago

About The Role

As a Senior Site Reliability Engineer, you will be responsible for maintaining and optimizing the platform's infrastructure, ensuring its reliability and scalability. You will work closely with the engineering and product teams, and your time will be divided between infrastructure development, platform optimization and resilience, and support and observability. You will have the opportunity to work remotely, enjoy a minimum of 25 days of paid time off, and benefit from a comprehensive health and retirement package.

  • Infrastructuurontwikkeling: Ontwikkelen en implementeren van schaalbare infrastructuur met behulp van Terraform en cloud-native AWS-services.
  • Platformoptimalisatie en veerkracht: Optimaliseren van de cloudinfrastructuur van het platform voor hoge beschikbaarheid en kostenefficiëntie.
  • Ondersteuning en observabiliteit: Deelnemen aan de on-call infrastructuurrotatie, reageren op incidenten en deze snel oplossen.
  • Observability and operations. You instrument systems for visibility and act on what they tell you, using tools such as Datadog, Pingdom, and Elastic to find and resolve issues before they affect end users
  • Ownership and collaboration. You take full ownership of reliability, you communicate clearly with engineering and product, you can persuade engineers outside your own team to prioritise reliability work, and you raise the standard of the systems and the teams you work with
  • Infrastructure automation and delivery. You automate and manage infrastructure with Terraform, and you build and maintain CI/CD pipelines with Jenkins as well as GitHub actions, including image builds with Packer and Docker. You write real code to do it: Python is essential, and experience with JavaScript, Node.js, or Go is an asset
  • Reliability engineering at scale. You bring around five or more years in a site reliability role, built on earlier experience as a system administrator or software developer, and a track record of keeping large-scale, 24/7 SaaS platforms running. You diagnose and resolve complex system issues under pressure, and you design for resilience before incidents happen
  • Cloud-native, distributed systems. You are hands-on with cloud-native architectures, distributed systems, and high-availability platforms, and you are comfortable across both document-oriented and relational databases. You have strong, production-grade experience with AWS, and knowledge of Azure is a bonus
  • AI-assisted development. You use AI coding tools such as Anthropic's Claude as a natural part of building and operating infrastructure, you apply good judgment about where they help and where they don't, with concrete examples of both from your own infra work, and you push the team to get more out of them
  • We are looking for a senior engineer who has spent years keeping large, business-critical SaaS platforms running, and who wants to take full ownership of reliability rather than simply keeping the lights on. In practice, that means:

This is an external listing. JobSpring does not represent or verify the employer. Report this listing