Your architecture diagram is not your resilience

https://azure.microsoft.com/en-us/blog/wp-content/uploads/2026/09/Designing-for-Resilience-6.jpg

This first article in our resilience series draws on a conversation with Mark Russinovich about how resilience is changing in the AI era and what it takes to continuously validate it at scale.

What surprises me most about resilience failures is how ordinary the drift is. A workload is deployed across availability zones, but a health probe still points to a single dependency. A database supports failover, but the application’s connection string is pinned to one region. Nothing looks broken. The architecture diagram still shows a resilient design even as the operational reality underneath it changes.

For years, resilience was something you set up once: configure disaster recovery, write a runbook, run the occasional failover test. That kept the lights on, but it treated resilience as a project with an end date rather than a property you maintain.

So, when an availability zone or a region has a bad day,...

Copyright of this story solely belongs to azure.microsoft.com. To see the full text click HERE

Read more