Skip to content
← Back to job listings

Director, Site Reliability Engineering - Incident Management

Yum! · Plano, TX, United States

Software DevelopmentExternal listingtemporaryabout 1 hour ago

About The Role

Purpose of The Position

The Director, Site Reliability Engineering – Incident Management owns the enterprise Reliability practice across Byte, KFC, and Taco Bell digital platforms, including its strategy, standards, governance, and delivery. Incident Management is the most visible part of that practice and sits with this role exclusively. This leader also owns the reliability platform, meaning the tooling, products, and capabilities that make reliability real for engineers and for markets, and serves as the accountable face to brands and markets for reliability process implementation, reporting, and the operational relationship. This leader establishes the strategy, governance, and operational excellence required to ensure highly available, resilient, and customer-centric technology services while developing high-performing teams and partnering across Engineering, Product, Infrastructure, and Brand leadership to continuously improve reliability and business continuity. This leader operates with exceptional diligence and care, bridging deep technical detail with human and business context so that executives, engineers, and restaurant teams experience clear, calm, and trustworthy communication in the moments that matter most.

Scope and Magnitude

  • ~30+ Teams – Byte, KFC and Taco Bell
  • Global Team - Vietnam, India, Colombia and US
  • Platform Team – Responsible for all products (Edge, Commerce, POS, KDS, Menu, Portal, etc.)

Position Functions

Strategic Leadership

  • Provide strategic leadership for Incident Management and Technical Operations across Byte, KFC, and Taco Bell digital platforms, ensuring high availability, operational excellence, and a consistent customer experience.
  • Establish the vision, governance, and operating model for enterprise Incident Management, driving standardized processes, tooling, and best practices across multiple brands and technology organizations.
  • Drive enterprise operational readiness for major product launches, restaurant initiatives, promotions, and seasonal events through effective change management, risk assessment, and cross-functional planning.
  • Own the enterprise framework and targets for service level objectives and indicators, set in partnership with the engineering teams that own the services, which remain accountable for achieving them. Define and monitor MTTR, incident trends, service health, operational maturity, and executive KPIs, using data to prioritize investments and improve platform resilience.
  • Own the enterprise Incident Management governance framework, ensuring consistent execution, accountability, and continuous improvement across all brands.
  • Own enterprise observability strategy, setting the standard and direction for how platform health is measured and seen, with Platform Engineering partnering on the underlying platform and instrumentation.
  • Own the strategy and roadmap for the reliability platform, including the tooling, products, and self-service capabilities that deliver reliability to engineering teams and to markets.
  • Provide executive-level communications and operational updates to Digital & Technology leadership, Brand CDTOs/CTOs, and executive stakeholders during major incidents and through regular operational business reviews.
  • Develop trusted partnerships with executive leaders across Product, Engineering, Infrastructure, Security, Restaurant Operations, and Brand Technology to align operational priorities with business objectives and customer experience goals.
  • Jointly own Business Continuity and Disaster Recovery with Platform Engineering, covering recovery strategies, resiliency testing, crisis management processes, and operational preparedness across Byte, KFC, and Taco Bell. Decisions are made in a standing joint review, with anything unresolved escalating to the Senior Director, Engineering within one cycle. During a declared continuity event or major incident, the Incident Commander decides in the moment.
  • Sponsor continuous improvement initiatives that enhance operational resilience, reduce enterprise risk, and strengthen the organization's ability to respond to large-scale operational events.

Technical Leadership

  • Serve as executive Incident Commander during critical enterprise incidents, providing leadership during major outages while ensuring effective cross-functional coordination across Engineering, Product, Infrastructure, Security, Restaurant Operations, Brand Leadership, and external partners.
  • Partner with Engineering, Platform, Security, and Architecture leaders to influence technology strategy, improve platform reliability, reduce operational toil, and accelerate automation.
  • Own the incident management and operational resilience strategy, setting the direction for detection, response, recovery, and preparedness across highly available digital platforms.
  • Influence architecture and platform design decisions to strengthen resiliency, reduce failure points, and improve the reliability of customer-facing services.
  • Set enterprise observability standards and direction, partnering with Platform Engineering on the platform and instrumentation that deliver against them.
  • Own detection, response, and remediation automation, partnering with Platform Engineering, which owns provisioning, deployment, and self-service automation.
  • Partner with Cloud, Platform, Engineering, Security, and Architecture leaders on cloud and platform engineering decisions that improve scalability, reliability, and operational resilience.
  • Champion a culture of operational excellence by driving post-incident reviews, root cause analysis, corrective actions, and continuous learning across engineering organizations.
  • Manage relationships with key technology vendors and managed service providers, ensuring operational performance, accountability, and adherence to service level commitments.
  • Lead cross-brand initiatives to mature observability, incident response automation, disaster recovery readiness, and operational resilience capabilities.

People Leadership

  • Lead, coach, and develop a high-performing organization of managers and technical leaders responsible for Incident Management and Technical Operations, building organizational capability, succession plans, and a culture of operational excellence and accountability.
  • Provide leadership for globally distributed Incident Management teams, ensuring consistent operational standards, clear ownership, and effective collaboration across regions and time zones.
  • Champion the team's culture of care, ensuring leaders support their people through high-pressure incidents, communicate with empathy, and build durable trust with every stakeholder and partner team.

Working Relationships

Internal

  • DTLT
  • Brand CDTOs/CTOs
  • Engineering and Product Teams
  • Reliability and Platform Engineering (GRE)
  • Security
  • Restaurant Operations

External

  • Vendors - managed service providers and technology partners (incident tooling, observability, cloud), with accountability for operational performance and SLA adherence
  • Brand market teams
  • Franchisees (potential)

Specialized or Technical Knowledge/Skills

  • 10+ years of experience in Site Reliability Engineering, Production Operations, Infrastructure Operations, or Technical Operations, including significant leadership experience managing managers and multiple operational teams.
  • Proven experience leading enterprise Incident Management programs supporting large-scale, customer-facing digital platforms.
  • Fluency with distributed systems, cloud infrastructure, CI/CD, telemetry, and modern SRE practices.
  • Demonstrated success developing operational strategy, governance, and standardized operating models across multiple organizations or business units.
  • Experience leading and developing leaders, building high-performing organizations, and driving employee engagement and organizational effectiveness.
  • Deep understanding of SRE principles, reliability engineering, observability, incident response, change management, operational resilience, Business Continuity, and Disaster Recovery.
  • Experience leading large cross-functional organizations through high-severity incidents while communicating effectively with executive leadership.
  • Strong business acumen with the ability to balance operational risk, customer experience, and business priorities.
  • Experience leading globally distributed teams and coordinating operations across multiple time zones.
  • Exceptional diligence and follow-through, with disciplined ownership of commitments, details, and outcomes across long-running operational programs.
  • Outstanding stakeholder communication, able to bridge deep technical detail with human and business context and translate incidents and risk into language executives, engineers, and restaurant teams can act on.
  • Leads with genuine care for people and partners, building the trust that turns incident response into a durable stakeholder relationship.

Preferred Requirements

  • Bachelor of Science in Computer Science
  • Experience supporting Quick Service Restaurant (QSR), retail, hospitality, or high-volume eCommerce platforms.
  • Experience leading enterprise operations across multiple brands or business units.
  • Expertise with cloud-native platforms, observability solutions, automation frameworks, and modern DevOps/SRE practices.
  • Experience driving organizational transformation, operational maturity, and large-scale process improvement initiatives.
  • Experience leading Business Continuity, Disaster Recovery, or enterprise resiliency programs.
  • Strong executive communication and stakeholder management skills, including presenting operational performance, risk assessments, and strategic recommendations to senior leadership.

Success Metrics

  • KPIs - MTTR, customer-impacting incident count and trend, repeat incident rate, BC/DR test coverage against recovery objectives, and stakeholder satisfaction with incident communications
  • A consistent, enterprise Incident Management program operates seamlessly across Byte, KFC, and Taco Bell, with standardized processes, governance, and measurable operational excellence.
  • Platform reliability improves year over year through reduced customer-impacting incidents, faster recovery times, and proactive risk management.
  • Business Continuity and Disaster Recovery capabilities are mature, regularly tested, and enable rapid recovery from significant operational events with minimal business disruption.
  • Engineering organizations consistently adopt operational best practices, resulting in improved service reliability, increased automation, and reduced operational toil.
  • Executive leadership has timely visibility into platform health, operational risk, resilience metrics, and strategic investment priorities.
  • Major launches, restaurant promotions, and seasonal events are executed with predictable operational readiness and minimal customer disruption.
  • Incident reviews consistently drive systemic improvements, resulting in fewer repeat incidents and stronger engineering ownership.
  • The Incident Management organization is recognized as a trusted strategic partner that enables business growth while protecting customer and restaurant experiences across all brands.
  • Leaders within the organization are developed, engaged, and prepared to scale operations through strong succession planning, coaching, and talent development.

Salary Range: $164,500 - $193,600 annually + bonus eligibility and stock-based compensation. This is the expected salary range for this position. Ultimately, in determining pay, we'll consider the successful candidate’s location, experience, and other job-related factors.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing