← Back to job listings
S
Senior Site Reliability Engineer / Platform Engineer
SimScale · Munich, Germany
About The Role
Join SimScale, a leading browser-based simulation platform, as a Senior Site Reliability Engineer / Platform Engineer. In this hands-on role, you will own and improve the cloud infrastructure, drive organization-wide adoption of observability tools, shape multi-region architecture, and own cloud cost and efficiency at scale. You will work closely with a small infrastructure team supporting 50+ engineers across the company. Enjoy benefits such as mobile working, competitive health benefits, discounted gym membership, flexible working hours, learning and development opportunities, child care contributions, and a retirement plan.
- Ownership and improvement of cloud infrastructure, including AWS and EKS, observability, disaster recovery, security, and compliance controls.
- Building standards, guardrails, and self-service tooling for engineering teams to safely run workloads on AWS.
- Driving organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and helping teams define meaningful SLOs.
- Strong systems foundation: You understand Linux internals and distributed systems well enough to debug complex production behavior
- Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-quality software in at least one of Python, Go, Rust, or Java
- Security and compliance awareness: You understand how infrastructure decisions affect access control, auditability, disaster recovery, logging, and standards such as SOC 2
- Production debugging depth: You can investigate complex failures, communicate clearly during incidents, and turn findings into durable improvements
- Hands-on cloud and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops-workflow (ArgoCD) and container orchestration (Kubernetes)
- 5+ years of professional experience in SRE, platform, or infrastructure engineering
- Clear communication: You can explain trade-offs to engineering teams and help others adopt better platform practices without unnecessary friction
- Observability and reliability experience: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and meaningful SLOs/SLIs
- An open source portfolio or contributions
- Prior technical leadership experience, especially in infrastructure, reliability, or platform engineering
This listing was posted by a verified recruiter at SimScale. Report this listing
JobSpring