← Back to job listings
SP
Senior Site Reliability Engineer
Spotme · United Kingdom
About The Role
As a Senior Site Reliability Engineer, you will be responsible for maintaining and optimizing the platform's infrastructure, ensuring its reliability and scalability. You will work closely with the engineering and product teams, and your time will be divided between infrastructure development, platform optimization and resilience, and support and observability. You will have the opportunity to work remotely, enjoy a minimum of 25 days of paid time off, and benefit from a comprehensive health and retirement package.
- Infrastructuurontwikkeling: Ontwikkelen en implementeren van schaalbare infrastructuur met behulp van Terraform en cloud-native AWS-services.
- Platformoptimalisatie en veerkracht: Optimaliseren van de cloudinfrastructuur van het platform voor hoge beschikbaarheid en kostenefficiëntie.
- Ondersteuning en observabiliteit: Deelnemen aan de on-call infrastructuurrotatie, reageren op incidenten en deze snel oplossen.
- Observability and operations. You instrument systems for visibility and act on what they tell you, using tools such as Datadog, Pingdom, and Elastic to find and resolve issues before they affect end users
- Ownership and collaboration. You take full ownership of reliability, you communicate clearly with engineering and product, you can persuade engineers outside your own team to prioritise reliability work, and you raise the standard of the systems and the teams you work with
- Infrastructure automation and delivery. You automate and manage infrastructure with Terraform, and you build and maintain CI/CD pipelines with Jenkins as well as GitHub actions, including image builds with Packer and Docker. You write real code to do it: Python is essential, and experience with JavaScript, Node.js, or Go is an asset
- Reliability engineering at scale. You bring around five or more years in a site reliability role, built on earlier experience as a system administrator or software developer, and a track record of keeping large-scale, 24/7 SaaS platforms running. You diagnose and resolve complex system issues under pressure, and you design for resilience before incidents happen
- Cloud-native, distributed systems. You are hands-on with cloud-native architectures, distributed systems, and high-availability platforms, and you are comfortable across both document-oriented and relational databases. You have strong, production-grade experience with AWS, and knowledge of Azure is a bonus
- AI-assisted development. You use AI coding tools such as Anthropic's Claude as a natural part of building and operating infrastructure, you apply good judgment about where they help and where they don't, with concrete examples of both from your own infra work, and you push the team to get more out of them
- We are looking for a senior engineer who has spent years keeping large, business-critical SaaS platforms running, and who wants to take full ownership of reliability rather than simply keeping the lights on. In practice, that means:
Similar roles you might like
See all →WC
East Locality Family Safeguarding Team Manager - Ref: CH07126
Walsall Council
Salary not disclosedPosted today
HQ
Maintenance Fitter Nights
Hanson Quarry Products Europe Limited
£54,340Posted today
LV
Veterinary Surgeon/Lead Vet Opportunity
Linnaeus Veterinary Limited
Up to £65,000/yrPosted today
SS
Lead Test Engineer
Saab Seaeye Ltd
Salary not disclosedPosted today
BG
Software Engineering & Innovation Summer Internship 2027
Baillie Gifford & Co
Salary not disclosedPosted today
BG
Cloud, Infrastructure & Security Summer Internship 2027
Baillie Gifford & Co
Salary not disclosedPosted today
BE
Director, Forward Deployed Engineer
BNY External Career Site
Salary not disclosedPosted today
PB
Engineering Team Leader & Contractor Supervisor
Polypipe Building Services
Salary not disclosedPosted today
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
