Talent.com
Yum Brands
Director, Site Reliability Engineering - Incident ManagementYum Brands • Plano, TX, United States
Director, Site Reliability Engineering - Incident Management

Director, Site Reliability Engineering - Incident Management

Yum Brands • Plano, TX, United States
23 hours ago
Job type
  • Full-time
Job description

Director, Site Reliability Engineering Incident Management

The Director, Site Reliability Engineering Incident Management owns the enterprise Reliability practice across Byte, KFC, and Taco Bell digital platforms, including its strategy, standards, governance, and delivery. Incident Management is the most visible part of that practice and sits with this role exclusively. This leader also owns the reliability platform, meaning the tooling, products, and capabilities that make reliability real for engineers and for markets, and serves as the accountable face to brands and markets for reliability process implementation, reporting, and the operational relationship. This leader establishes the strategy, governance, and operational excellence required to ensure highly available, resilient, and customer-centric technology services while developing high-performing teams and partnering across Engineering, Product, Infrastructure, and Brand leadership to continuously improve reliability and business continuity. This leader operates with exceptional diligence and care, bridging deep technical detail with human and business context so that executives, engineers, and restaurant teams experience clear, calm, and trustworthy communication in the moments that matter most.

~30+ Teams Byte, KFC and Taco Bell

Global Team - Vietnam, India, Colombia and US

Platform Team Responsible for all products (Edge, Commerce, POS, KDS, Menu, Portal, etc.)

Responsibilities

Position Functions:

Strategic Leadership

  • Provide strategic leadership for Incident Management and Technical Operations across Byte, KFC, and Taco Bell digital platforms, ensuring high availability, operational excellence, and a consistent customer experience.
  • Establish the vision, governance, and operating model for enterprise Incident Management, driving standardized processes, tooling, and best practices across multiple brands and technology organizations.
  • Drive enterprise operational readiness for major product launches, restaurant initiatives, promotions, and seasonal events through effective change management, risk assessment, and cross-functional planning.
  • Own the enterprise framework and targets for service level objectives and indicators, set in partnership with the engineering teams that own the services, which remain accountable for achieving them. Define and monitor MTTR, incident trends, service health, operational maturity, and executive KPIs, using data to prioritize investments and improve platform resilience.
  • Own the enterprise Incident Management governance framework, ensuring consistent execution, accountability, and continuous improvement across all brands.
  • Own the enterprise observability strategy, setting the standard and direction for how platform health is measured and seen, with Platform Engineering partnering on the underlying platform and instrumentation.
  • Own the strategy and roadmap for the reliability platform, including the tooling, products, and self-service capabilities that deliver reliability to engineering teams and to markets.
  • Provide executive-level communications and operational updates to Digital & Technology leadership, Brand CDTOs/CTOs, and executive stakeholders during major incidents and through regular operational business reviews.
  • Develop trusted partnerships with executive leaders across Product, Engineering, Infrastructure, Security, Restaurant Operations, and Brand Technology to align operational priorities with business objectives and customer experience goals.
  • Jointly own Business Continuity and Disaster Recovery with Platform Engineering, covering recovery strategies, resiliency testing, crisis management processes, and operational preparedness across Byte, KFC, and Taco Bell. Decisions are made in a standing joint review, with anything unresolved escalating to the Senior Director, Engineering within one cycle. During a declared continuity event or major incident, the Incident Commander decides in the moment.
  • Sponsor continuous improvement initiatives that enhance operational resilience, reduce enterprise risk, and strengthen the organization's ability to respond to large-scale operational events.

Technical Leadership

  • Serve as executive Incident Commander during critical enterprise incidents, providing leadership during major outages while ensuring effective cross-functional coordination across Engineering, Product, Infrastructure, Security, Restaurant Operations, Brand Leadership, and external partners.
  • Partner with Engineering, Platform, Security, and Architecture leaders to influence technology strategy, improve platform reliability, reduce operational toil, and accelerate automation.
  • Own the incident management and operational resilience strategy, setting the direction for detection, response, recovery, and preparedness across highly available digital platforms.
  • Influence architecture and platform design decisions to strengthen resiliency, reduce failure points, and improve the reliability of customer-facing services.
  • Set enterprise observability standards and direction, partnering with Platform Engineering on the platform and instrumentation that deliver against them.
  • Own detection, response, and remediation automation, partnering with Platform Engineering, which owns provisioning, deployment, and self-service automation.
  • Partner with Cloud, Platform, Engineering, Security, and Architecture leaders on cloud and platform engineering decisions that improve scalability, reliability, and operational resilience.
  • Champion a culture of operational excellence by driving post-incident reviews, root cause analysis, corrective actions, and continuous learning across engineering organizations.
  • Manage relationships with key technology vendors and managed service providers, ensuring operational performance, accountability, and adherence to service level commitments.
  • Lead cross-brand initiatives to mature observability, incident response automation, disaster recovery readiness, and operational resilience capabilities.

People Leadership

  • Lead, coach, and develop a high-performing organization of managers and technical leaders responsible for Incident Management and Technical Operations, building organizational capability, succession plans, and a culture of operational excellence and accountability.
  • Provide leadership for globally distributed Incident Management teams, ensuring consistent operational standards, clear ownership, and effective collaboration across regions and time zones.
  • Champion the team's culture of care, ensuring leaders support their people through high-pressure incidents, communicate with empathy, and build durable trust with every stakeholder and partner team.

Working Relationships:

Internal

  • DTLT
  • Brand CDTOs/CTOs
  • Engineering and Product Teams
  • Reliability and Platform Engineering (GRE)
  • Security
  • Restaurant Operations

External

  • Vendors - managed service providers and technology partners (incident tooling, observability, cloud), with accountability for operational performance and SLA adherence
  • Brand market teams
  • Franchisees (potential)

Qualifications

Specialized or Technical Knowledge/Skills

  • 10+ years of experience in Site Reliability Engineering, Production Operations, Infrastructure Operations, or Technical Operations, including significant leadership experience managing managers and multiple operational teams.
  • Proven experience leading enterprise Incident Management programs supporting large-scale, customer-facing digital platforms.
  • Fluency with distributed systems, cloud infrastructure, CI/CD, telemetry, and modern SRE practices.
  • Demonstrated success developing operational strategy, governance, and standardized operating models across multiple organizations or business units.
  • Experience leading and developing leaders, building high-performing organizations, and driving employee engagement and organizational effectiveness.
  • Deep understanding of SRE principles, reliability engineering, observability, incident response, change management, operational resilience, Business Continuity, and Disaster Recovery.
  • Experience leading large cross-functional organizations through high-severity incidents while communicating effectively with executive leadership.
  • Strong business acumen with the ability to balance operational risk, customer experience, and business priorities.
  • Experience leading globally distributed teams and coordinating operations across multiple time zones.
  • Exceptional diligence and follow-through, with disciplined ownership of commitments, details, and outcomes across long-running operational programs.
  • Outstanding stakeholder communication, able to bridge deep technical detail with human and business context and translate incidents and risk into language executives, engineers, and restaurant teams can act on.
  • Leads with genuine care for people and partners, building the trust that turns incident response into a durable stakeholder relationship.

Preferred Requirements:

  • Bachelor of Science in Computer Science
  • Experience supporting Quick Service Restaurant (QSR), retail, hospitality, or high-volume eCommerce platforms.
  • Experience leading enterprise operations across multiple brands or business units.
  • Expertise with cloud-native platforms, observability solutions, automation frameworks, and modern DevOps/SRE practices.
  • Experience driving organizational transformation, operational maturity, and large-scale process improvement initiatives.
  • Experience leading Business Continuity, Disaster Recovery, or enterprise resiliency programs.
  • Strong executive communication and stakeholder management skills, including presenting operational performance, risk assessments, and strategic recommendations to senior leadership.

Success Metrics:

  • KPIs - MTTR, customer-impacting incident count and trend, repeat incident rate, BC/DR test coverage against recovery objectives, and stakeholder satisfaction with incident communications
Create a job alert for this search

Director, Site Reliability Engineering - Incident Management • Plano, TX, United States

Similar jobs

Director Engineering - Sensing Solutions NPI

Honeywell InternationalRichardson, TX, United States
Full-time

Director Engineering - Sensing Solutions NPI.The Director Engineering - Sensing Solutions NPI will be responsible for technology roadmap, strategy development and new product introductions for the ... Show more

 • Promoted

Revenue Cycle Director, Critical Access Hospitals

Prime HealthcareDallas, TX, United States
Full-time

Prime Healthcare Revenue Cycle Director.Prime Healthcare is an award-winning health system headquartered in Ontario, California.Prime Healthcare operates 55 hospitals and has more than 360 outpatie... Show more

 • Promoted

Director, Risk Management

Crow HoldingsDallas, TX, United States
Full-time

Trammell Crow Residential (TCR) is a leading multifamily real estate developer with a local presence in 16 key U.Over 45 years, TCR has built more than 292,000 premier multifamily residences, deliv... Show more

 • Promoted

Director of Engineering

ECLARORichardson, TX, United States
Full-time

ECLARO is looking for a Director of Engineering for our client in Richardson, TX.ECLARO's client is a leading technology solutions provider, collaborating with customers to manage their needs and a... Show more

 • Promoted

AT&T SE Systems Engineering Director

Hewlett Packard EnterpriseDallas, TX, United States
Full-time

AT&T SE Systems Engineering Director.This role has been designated as 'Remote/Teleworker', which means you will primarily work from home.Hewlett Packard Enterprise is the global edge-to-cloud compa... Show more

 • Promoted

Director of Enterprise Resilience

Health Care Service CorporationRichardson, TX, United States
Full-time

Director Of Enterprise Resilience.The Director of Enterprise Resilience will be responsible for the strategic initiatives, implementation and governance of business resiliency and crisis management... Show more

 • Promoted

CRITICAL INCIDENT MANAGER

YochanaFrisco, TX, United States
Full-time

Incident Management Responsibilities.Understand the incident and the diagnostic/resolution actions attempted already by the service desk and any other technology tracks.Use the designated or allott... Show more

 • Promoted

Director of Project Deliverables - Critical Facilities Design

PkazaDallas, TX, United States
Full-time

Director Of Project Deliverables - Data Center Design.We are seeking an experienced Director who will lead the development, project deliverables, and oversight of critical power and mechanical infr... Show more

 • Promoted

Critical Incident Manager

Omni InclusivePlano, TX, United States
Full-time

Understand the incident and the diagnostic/resolution actions attempted already by the service desk and any other technology tracks.Use the designated or allotted communication bridge, monitoring f... Show more

 • Promoted

Director of Quality & Risk Management

DataOneDallas, TX, United States
Full-time

DataOne Systems is a leading provider of Engineering, Furnishing, and Installation (EF&I) solutions, delivering high-quality Layer 1 infrastructure that supports data centers, telecommunications en... Show more

 • Promoted

Executive Director Infrastructure Engineering

DTCCDallas, TX, United States
Full-time

Executive Director Infrastructure Engineering.Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive invest... Show more

 • Promoted

Director _ Engineering & IoT, Utilities

ClifyXDallas, TX, United States
Full-time

Digital / EIS: Internet of Things(IoT) for Energy Management~ Geographic Information System (GIS) Domain - Generic (E0~E1).Focus on the core content of the job post, removing extra metadata, requis... Show more

 • Promoted

Regional Safety Director, Mission Critical

Suffolk ConstructionDallas, TX, United States
Full-time

Suffolk is seeking people who are bold.Looking for the career opportunity of a lifetime.We'll challenge and inspire you to be your very best.We'll embrace what makes you unique and lift you up as y... Show more

 • Promoted

Director, Engineering Residences Americas

Rosewood HotelsDallas, TX, United States
Temporary

Director, Engineering, Residences Americas.As part of the Residential Operations team within the Americas region, this role supports the engineering oversight, operational standards, and long-term... Show more

 • Promoted

Director, Emerging Infrastructure Ecosystems

EquinixDallas, TX, United States
Full-time

Director, Emerging Infrastructure Ecosystems.Equinix is the world's digital infrastructure company, shortening the path to connectivity to enable the innovations that enrich our work, life and plan... Show more

 • Promoted

Security & Risk Consulting Client Engagement Director - 1898 & Co.

Burns & McDonnellDallas, TX, United States
Full-time

Security & Risk Consulting Client Engagement Director.Burns & McDonnell, as we lead the charge in securing critical infrastructure and shaping the future of industrial cybersecurity.Our team partne... Show more

 • Promoted

IT Release & Major Incident Manager: 25-06381

AkrayaDallas, TX, United States
Full-time

IT Release & Major Incident Manager.As the IT Release & Major Incident Manager, you will be at the forefront of ensuring seamless IT releases and effective major incident management, minimizing dis... Show more

 • Promoted

Director of Engineering

The Falcon GroupDallas, TX, United States
Full-time

We are seeking a licensed Director of Engineering to lead our Dallas - Ft.Worth Metro office, the office is located in Dallas, TX.This role will be instrumental in building and growing our Texas bu... Show more

 • Promoted

Director, eDiscovery

AnkuraDallas, TX, United States
Full-time

Ankura is a team of excellence founded on innovation and growth.Bright Labs, proudly a part of Ankura, is at the forefront of the eDiscovery industry, having contributed to some of the most signifi... Show more

 • Promoted

Perm - Leadership - Director of HRIS (Days) Dallas TX

Reliant Staffing SolutionsDallas, TX, United States
Full-time

As one of the largest public hospital systems in the country, Parkland Health & Hospital System is dedicated to providing valuable health and well-being services to those entrusted to our care.Perm... Show more