Skip to content
← Back to job listings

Senior Site Reliability Engineer / Platform Engineer

SimScale · Munich, Germany

Quick applyfull-timeabout 2 months ago

About The Role

Join SimScale, a leading browser-based simulation platform, as a Senior Site Reliability Engineer / Platform Engineer. In this hands-on role, you will own and improve the cloud infrastructure, drive organization-wide adoption of observability tools, shape multi-region architecture, and own cloud cost and efficiency at scale. You will work closely with a small infrastructure team supporting 50+ engineers across the company. Enjoy benefits such as mobile working, competitive health benefits, discounted gym membership, flexible working hours, learning and development opportunities, child care contributions, and a retirement plan.

  • Ownership and improvement of cloud infrastructure, including AWS and EKS, observability, disaster recovery, security, and compliance controls.
  • Building standards, guardrails, and self-service tooling for engineering teams to safely run workloads on AWS.
  • Driving organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and helping teams define meaningful SLOs.
  • Strong systems foundation: You understand Linux internals and distributed systems well enough to debug complex production behavior
  • Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-quality software in at least one of Python, Go, Rust, or Java
  • Security and compliance awareness: You understand how infrastructure decisions affect access control, auditability, disaster recovery, logging, and standards such as SOC 2
  • Production debugging depth: You can investigate complex failures, communicate clearly during incidents, and turn findings into durable improvements
  • Hands-on cloud and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops-workflow (ArgoCD) and container orchestration (Kubernetes)
  • 5+ years of professional experience in SRE, platform, or infrastructure engineering
  • Clear communication: You can explain trade-offs to engineering teams and help others adopt better platform practices without unnecessary friction
  • Observability and reliability experience: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and meaningful SLOs/SLIs
  • An open source portfolio or contributions
  • Prior technical leadership experience, especially in infrastructure, reliability, or platform engineering

This listing was posted by a verified recruiter at SimScale. Report this listing