← Back to job listings
SC
Senior Site Reliability Engineer
Stellar Cyber · Spain
About The Role
Join our team as a Senior Site Reliability Engineer (SRE) and drive reliability, scalability, and efficiency across our production systems. You will have deep expertise in cloud infrastructure, Kubernetes administration, observability, and incident management. As a senior member of the SRE team, you will influence architecture, tooling, and best practices to ensure operational excellence.
- Administer and maintain container orchestration platforms and containerized workloads, ensuring high availability and resilience.
- Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms.
- Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments, implementing Infrastructure as Code (IaC) practices.
- Expertise in operating data platforms (Elasticsearch, MongoDB, Spark, Kafka, Redis)
- Deep understanding of Infrastructure as Code (Terraform, Helm)
- Proficiency with public cloud services (AWS, Azure, GCP, or OCI)
- Hands-on experience with observability tools: Prometheus, Grafana, Loki, and Alertmanager
- Strong programming and automation skills in Python and Bash
- Certifications in AWS, GCP, Observability, Linux or Kubernetes are a plus
- 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles
- Bachelor's degree in Computer Science, Engineering, or a related technical field
- Knowledge in chat-based operations interfaces and/or auto-remediation controllers using AI agentic framework
- Proven success leading large-scale production systems in cloud environments (AWS, GCP, Azure, or OCI)
- Understanding of AI agents for Auto-triaging alerts, correlate signals and suggest/root-cause hypotheses
- Strong experience with production on-call operations and incident management
- Experience with CI/CD pipelines (GitHub Actions, Bitbucket, ArgoCD)
- Excellent problem-solving, communication, and leadership abilities
- Advanced proficiency in Kubernetes administration and troubleshooting
- Strong technical background in distributed systems, databases, networking, and Linux administration
- Demonstrated leadership in driving incident response, on-call best practices, and reliability-focused culture
Similar roles you might like
See all →1G
Key Account Manager Oncología - Madrid & Extremadura
1725 GlaxoSmithKline S.A.
Salary not disclosedPosted 2 days ago
1G
Beca Claims
12_GRUPO GENERALI ESPAÑA, A.I.E.
Salary not disclosedPosted 2 days ago
9D
Operations Specialist- Performance Measurement AVP
930G DWS International GmbH, Madrid Branch
Salary not disclosedPosted 2 days ago
4D
Risk Data Analyst (f/m/x)
4599 DB Operaciones y Servicios Interactivos, S.L.U.
Salary not disclosedPosted 2 days ago
GS
Rental Sales Agent
Goldcar Spain S.L.U
Salary not disclosedPosted 2 days ago
E
Consultant, Global Market Access & Pricing (French & English speaking)
EVERSANA
Salary not disclosedPosted 2 days ago
RE
Beca Ing. Logística Inbound
RENAULT ESPAÑA, S.A.
Salary not disclosedPosted 2 days ago
RE
Beca - Introducción a la tecnología de baterías y Electrónica de Potencia
RENAULT ESPAÑA, S.A.
Salary not disclosedPosted 2 days ago
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
