Skip to content
← Back to job listings

Senior Site Reliability Engineer

rapidai · Bangalore, India

Software DevelopmentSenior LevelImported listingfull-time23 days ago

About The Role

RapidAI is the trusted leader in deep clinical AI, helping hospitals deliver faster, more informed care through intelligent imaging and integrated workflows. The Rapid Enterprise™ Platform supports disease states across the care spectrum, but it’s our clinical depth that drives the most meaningful impact — improving decision-making, patient outcomes, and health-system performance. Used by more than 2,500 hospitals in over 100 countries and backed by 700+ clinical studies, including research that helped expand national stroke-treatment guidelines, RapidAI is the most clinically validated AI platform in healthcare.

What You Do

  • Own the availability, performance, and incident response for Rapid's production EKS clusters
  • Design and operate the full observability stack — metrics, logs, traces — withOpen Telemetry as the foundation
  • Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
  • Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
  • Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
  • Tune autoscaling, networking, and cost efficiency across AWS workloads
  • On-call rotation with the expectation you'll also fix the underlying cause, not just the alert

What We Looking For

10+ years in SRE, DevOps, or infrastructure engineering roles

Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and thesurrounding ecosystem

Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic

Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)

Strong foundation in Linux, networking, and distributed systems fundamentals

Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out

Startup mindset: you make decisions with incomplete information and iterate quickly

This is an external listing. JobSpring does not represent or verify the employer. Report this listing