Resilience at Cloud Scale: Azure CTO on Outages, Hardware, and AI

Mark Russinovich discusses what resilience means at Azure’s scale, using a 2014 near-outage as a starting point to explain how Azure tests failure, defines service health, and responds to incidents. He also covers how AI and AI agents change reliability assumptions, and where Azure’s architecture frameworks fit.

Overview

Azure CTO Mark Russinovich talks with Adam Bogobowicz about how Azure approaches resilience at cloud scale, and why the definition of “resilient” keeps changing as the platform grows and workloads evolve.

Key themes covered in the conversation include:

Resources mentioned

Video chapters