← Back to job listings

Senior Site Reliability Engineer
Kody · Hong Kong
About The Role
Job Summary
Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen , you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.
Key Responsibilities
- Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
- Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
- SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
- Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
- Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
- Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.
Requirements
Qualifications & Requirements
- Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
- Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux , networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
- Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
- Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
- SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
- Location & Communication: Based in Hong Kong or Shenzhen . Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.
Leadership & Operational Excellence
- Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
- Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
- Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
- Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.
Benefits
- Competitive Package
- A dynamic and innovative team
- Collaborative, inclusive working environment
Similar roles you might like
See all →M
Full-stack Software Engineer
Manulife
Salary not disclosedPosted 1 day ago
PI
Assistant Manager, Software Engineering (1-Year Contract)
PUMA International Trading Services Limited
Salary not disclosedPosted 2 days ago
M
Senior Full-stack Software Engineer
Manulife
Salary not disclosedPosted 2 days ago
1B
VP Software Developer - APAC Flow Volatility Risk Pricing
1254 BC Asia Ltd
Salary not disclosedPosted 3 days ago
HK
AI Developer I, HK - Winter 2027
Hong Kong Data Lab
Salary not disclosedPosted 3 days ago
1C
Technical Lead, Network Services, IT
1899 CITIC Securities International Company Limited
Salary not disclosedPosted 3 days ago
W
Software Developer II - Mobile
WATI.io
Salary not disclosedPosted 3 days ago
1C
Principal DevOps, Algo KDB, IT
1899 CITIC Securities International Company Limited
Salary not disclosedPosted 4 days ago
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
