Skip to content
← Back to job listings

Site Reliability Engineer

PagerDuty · Toronto, Canada

RemoteImported listingfull-time10 days ago

About The Role

Join PagerDuty as a Site Reliability Engineer on the Core Infrastructure team. You'll help build and operate the foundational infrastructure that powers our real-time digital operations platform. Your work will directly impact the reliability, scalability, and security of our services. You'll support and improve foundational infrastructure, contribute to the reliability and scalability of our core platform, and participate in agile rituals. You'll also monitor system health and participate in 24/7 on-call rotations.

  • Contribuer à la fiabilité et à l'évolutivité de la plateforme de base de PagerDuty en renforçant les systèmes existants et en soutenant le déploiement de nouvelles capacités d'infrastructure.
  • Participer aux rituels agiles (réunions debout, planification, rétrospectives) et communiquer les progrès/risques dès que possible.
  • Surveiller la santé du système à l'aide de métriques, de journaux et d'alertes, et participer aux rotations d'astreinte 24/7 pour aider à détecter, répondre et résoudre les incidents.
  • Working knowledge of networking fundamentals, such as load balancing, DNS,

TLS, and ingress traffic flow

  • Hands-on experience operating Linux-based systems in production

environments

  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
  • 0 to 1+ years of experience in Site Reliability Engineering, DevOps, or Platform

Engineering roles

  • Experience with container orchestration (e.g., EKS, Kubernetes)
  • Experience working on cloud-native infrastructure (e.g., AWS, GCP, Azure),

including networking and compute concepts

  • Proficiency in at least one programming language (e.g., Python, Ruby, Go, etc.)
  • Experience with AWS cloud networking concepts such as VPCs, subnets,

routing, security groups, and load balancers

  • Experience operating or contributing to production Kubernetes platforms (e.g.,

EKS), including cluster upgrades, networking, or ingress configuration

  • Experience with monitoring, observability, and logging platforms (e.g., DataDog,

New Relic, SumoLogic, Splunk, Prometheus, Grafana)

  • Familiarity with service meshes, ingress controllers, or API gateways (e.g.,

Envoy, Istio, NGINX)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing