Site Reliability Engineer
PagerDuty · Toronto, Canada
About The Role
Join PagerDuty as a Site Reliability Engineer on the Core Infrastructure team. You'll help build and operate the foundational infrastructure that powers our real-time digital operations platform. Your work will directly impact the reliability, scalability, and security of our services. You'll support and improve foundational infrastructure, contribute to the reliability and scalability of our core platform, and participate in agile rituals. You'll also monitor system health and participate in 24/7 on-call rotations.
- Contribuer à la fiabilité et à l'évolutivité de la plateforme de base de PagerDuty en renforçant les systèmes existants et en soutenant le déploiement de nouvelles capacités d'infrastructure.
- Participer aux rituels agiles (réunions debout, planification, rétrospectives) et communiquer les progrès/risques dès que possible.
- Surveiller la santé du système à l'aide de métriques, de journaux et d'alertes, et participer aux rotations d'astreinte 24/7 pour aider à détecter, répondre et résoudre les incidents.
- Working knowledge of networking fundamentals, such as load balancing, DNS,
TLS, and ingress traffic flow
- Hands-on experience operating Linux-based systems in production
environments
- Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
- 0 to 1+ years of experience in Site Reliability Engineering, DevOps, or Platform
Engineering roles
- Experience with container orchestration (e.g., EKS, Kubernetes)
- Experience working on cloud-native infrastructure (e.g., AWS, GCP, Azure),
including networking and compute concepts
- Proficiency in at least one programming language (e.g., Python, Ruby, Go, etc.)
- Experience with AWS cloud networking concepts such as VPCs, subnets,
routing, security groups, and load balancers
- Experience operating or contributing to production Kubernetes platforms (e.g.,
EKS), including cluster upgrades, networking, or ingress configuration
- Experience with monitoring, observability, and logging platforms (e.g., DataDog,
New Relic, SumoLogic, Splunk, Prometheus, Grafana)
- Familiarity with service meshes, ingress controllers, or API gateways (e.g.,
Envoy, Istio, NGINX)
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
