How STPA Helps Find Unknown Unknowns Before They Cost You Millions
One reason complex systems are hard to reason about is that catastrophic failures are often not caused by a single broken component. Instead, they emerge from the way seemingly healthy components interact.
A retry mechanism can amplify load instead of helping recovery. An autoscaler can make a correct decision based on stale metrics. Two automated systems can optimize for different local goals and quietly fight each other in production.
In all of these cases, there may be no obviously broken component. No single service is necessarily down, and no engineer made a clearly wrong decision. The loss appears in the gaps between the parts.
Most engineering teams are pretty good at identifying known risks. We conduct architecture and incident reviews, run threat-modeling exercises, write postmortems, and build dashboards and alerts. All of that is useful.
But most of these practices have the same limitation: they usually start from risks we...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE