← Back to job listings
SE
Senior Manager of Reliability Engineering Management
ServiceNow · Dublin, Ireland
About The Role
Join ServiceNow as a Senior Manager of Reliability Engineering Management. Lead a global team of Site Reliability Engineers (SRE) focused on ensuring the reliability and availability of critical enterprise platforms and applications. Drive the SRE strategy and operating model, own observability capabilities, and lead cloud modernization initiatives. Collaborate with various teams to improve the reliability and availability of critical platforms and services.
- Lead and develop a global team of Site Reliability Engineering (SRE) leaders, managers, and engineers, with accountability for talent development, performance, prioritization, succession planning, and execution.
- Define and drive the SRE strategy and operating model across reliability, observability, automation, incident response, production readiness, and continuous improvement.
- Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, error budgets, and service health practices to improve detection, diagnosis, and overall production reliability.
- Know operating systems in various levels of troubleshooting and diagnostics?
- Have experience in leading a team of engineers and exposure to people management?
- Have experience with cloud technologies and hyperscalers such as AWS, GCP, or Azure, and a passion for driving cloud modernization and cloud-native transformation?
- Have a technical background in roles like systems engineering or devops or site reliability engineering?
- Have low tolerance to repetitive tasks and automate your way through work?
- If you Answered 'yes' to these questions, we want to hear from you. Hit the Apply button and let's have a chat about the role and your skills and experiences
- Excellent written and verbal communication skills, with the ability to translate complex technical topics into clear business outcomes and leadership decisions
- Experience applying automation, orchestration, and infrastructure-as-code to reduce operational toil and improve repeatability and reliability
- Demonstrated ability to influence organizational boundaries and drive alignment among Engineering, Product, Architecture, Security, Infrastructure, and Operations teams
- Strong technical foundation across Linux, distributed systems, databases, networking, systems troubleshooting, scripting, and software engineering fundamentals
- Experience operating high-scale software, platform, and infrastructure-as-a-service environments with demanding availability and reliability requirements
- 5+ years of people-management experience, including experience leading managers, senior technical leaders, and geographically distributed engineering teams
- Experience building or operating production-like staging/test environments and establishing validation strategies for reliability, resiliency, performance, integration, and release readiness
- Experience designing and operating observability platforms at scale, including metrics, logging, tracing, alerting, dashboards, golden signals, SLI/SLOs, and error budgets
- Strong understanding of incident management, problem management, operational readiness, and continuous improvement practices
- Deep working knowledge of one or more major hyperscalers - AWS, GCP, or Azure with experience designing and operating resilient, scalable, highly available production systems
- Experience leveraging or critically evaluating AI and AI-assisted operations to automate workflows, accelerate diagnosis and remediation, and improve engineering productivity
- Experience with CI/CD and release pipelines, including automated confidence gates, progressive/phased deployments, zero-downtime deployment strategies, rollback mechanisms, and production-readiness controls
- Strong experience with cloud modernization and migration, including modernizing legacy platforms and tooling into cloud-native architectures
- Significant experience leading Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or Cloud Infrastructure organizations in large-scale production environments
- Strong understanding of Kubernetes, containers, orchestration, service networking, autoscaling, and modern cloud-native architecture patterns
- Ability to lead effectively through ambiguity and change while maintaining a strong focus on execution, customer impact, and engineering excellence
- RHCE, CCNA, ITIL or other industry certifications
Similar roles you might like
See all →H
Lifecycle Marketing Manager
Hunter
$84,000 – 97,000/yrPosted today
0A
Engineering Associate Manager
0701 AGS - IP Company
Salary not disclosedPosted today
OI
Manager, Accounting
Omnissa International Unlimited Company
Salary not disclosedPosted today
3K
AWS Business Architect
3620 Kyndryl UK Limited
Salary not disclosedPosted today
5P
Senior Business Systems Analyst
501 PGIM Ireland Limited
Salary not disclosedPosted today
IM
Senior Data Scientist, AI Engineering
Ireland: Mastercard Ireland Limited
Salary not disclosedPosted today
A
Solutions Engineer
AMAX
Salary not disclosedPosted today
2H
Administration Executive - Kilkenny
2386IE00 Howden Insurance (Ireland) Limited
Salary not disclosedPosted today
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
