Skip to content
← Back to job listings

Senior Site Reliability Engineer

Carousell Group · Kuala Lumpur, Federal Territory of Kuala Lumpur, Malaysia

Software DevelopmentImported listingfull-timeabout 24 hours ago

About The Role

Responsibilities

-
Operate against our existing SLIs, SLOs and error budgets, hold services to them, review and adjust targets as systems evolve, and use budget burn to arbitrate between reliability work and feature velocity.
-
Own the incident lifecycle end to end detection, response, mitigation, blameless postmortems with tracked follow-through troubleshooting across the whole stack (OS, application, database, cache, network), and mature the on-call rotation around it: alerts tuned for signal, runbooks kept current, MTTD and MTTR trending down.
-
Extend and improve our observability stack instrument new services, close coverage gaps in metrics, logs and traces, and raise dashboard and alert quality so teams can diagnose their own services.
-
Design and maintain safe release processes: canary, progressive rollout, automated rollback across dev, staging and production.
-
Eliminate toil through automation; treat repetitive manual operations as bugs to be engineered away.
-
Own infrastructure as code provisioning, configuration, policy, and the documentation around it.
-
Capacity planning, performance tuning and cloud cost efficiency forecast growth, model headroom, upgrade before saturation.
-
Implement and maintain infrastructure security controls secrets and credential management in Vault, access control, monitoring and response.
-
Conduct production readiness reviews and resilience testing for new and existing services.
-
Operate our MCP gateway and internal AI tooling as production infrastructure availability, access control, rate limiting, cost and usage visibility and apply AI-assisted automation to operational work such as incident triage, log summarisation and runbook generation.
Reliability and infrastructure
-
5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.
-
Kubernetes in production deployment, upgrades, troubleshooting and ongoing maintenance.
-
Strong Linux fundamentals and performance tuning (RHEL / CentOS / Debian / Ubuntu), plus Docker and container networking.
-
Hands-on Google Cloud Platform experience, with infrastructure as code and declarative provisioning (Terraform or equivalent).
-
Secrets and credential management with HashiCorp Vault (or equivalent) policies, rotation and least-privilege access.
-
CI/CD delivery pipelines on GitHub Actions (or equivalent), including progressive delivery and automated rollback.
-
Bash plus working proficiency in a general-purpose language (Go, Python or similar) for building real tooling.
-
Monitoring and observability with Prometheus, Grafana and Google Cloud Operations instrumentation, dashboards, alerting and tracing.
-
Able to work independently on large, complex projects with minimal guidance.
AI-assisted operations
-
Practical experience applying AI and LLM-based tooling to engineering or operational workflows, with a clear view of where it helps and where it doesn't.
-
Working familiarity with MCP (Model Context Protocol) and agent harnesses tool-calling loops, context management, guardrails and failure handling or the appetite and fundamentals to pick them up quickly.
Why Join Us?
At Mudah, we are evolving towards an AI-first engineering organisation. This role is an opportunity to go beyond traditional technical leadership and help shape how engineering teams build software with AI and agentic workflows.
You will have the opportunity to influence both the technology we build and the way we build it.
By proceeding with your application, you are adhering to our PDPA policies. In case you are interested to know more, read about our Candidates Personal Data Privacy Statement.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing