Your next incident won't define your system's fortunes. Your work before the storm will
One question I'm often asked: 'how do we create performant and resilient systems, at scale?' It is a sound question to ask.
But behind this query, there often sits an uncomfortable truth. Many large-scale systems morph and iterate over time, often without embedding the 'preventative and performative' elements that ensure high levels of service quality, availability and experience for the end user. This is a material risk to large-scale systems today, and systems at scale can literally live or die by it.
So important, then, that the 'real world readiness' of the system to react, adapt and bounce back at scale is built into its very core, by design.
This is why I developed the 'Three Ps' Framework, for performant systems. It's the set of guardrails and 'silent scaffolding' I apply to my products at scale, to not only ensure high levels of performance and 'non-functional' robustness when the system is live and running, but also to power its recovery and bounce-back when it goes down. All systems experience such moments, at some point or other. Best, then, to be prepared. Enter the Three Ps.
The Three Ps of Performant Systems
Plan: map out blind spots, the risk areas, the unhappy paths, and run through the key 'What If?' scenarios, to stress test your system in advance. This 'P' serves a cunning twin purpose: it helps prevent many problems from hitting the system in the first place, whilst also providing the fuel for a further 'P' to be built and readied in advance: Prepare.
Prepare: for the worst, whilst always hoping for the best (the worst will often happen anyway, so best to be ready). Your disaster management and recovery systems, your runbooks, incident and escalation paths that need to play out in real time, need to be stress tested whilst the sun is shining, not once the storm itself arrives. They need to be documented, and easily accessible too, ready for those moments of truth when they hit.
Perform: the final 10% of how you act in the moments of crisis, or when the system is tested. Even this can be dry run or rehearsed (part of 'Prepare'), to build the crisis management muscles before they're ever even needed. It's like how fire crews don't just sit idle waiting to attend the next fire or incident, they train, they rehearse, they develop new smarts to counter emerging threats. So too with resilient systems at scale. The 90% readiness comes from the first two Ps; the remaining 10% comes from how your team performs in those moments of truth.
Over to you: how do you engineer your systems at scale to be highly performant in the 'usual' times, but able to react and recover when the system is under real world stress?
More Product Truths in the book, 'Product Truths: Five Principles for Building Products That Work at Scale'. Available at all major digital bookstores (Amazon Kindle, Kobo, Apple Books): https://books2read.com/producttruths.
The Three Ps of Performant Systems: Copyright Ian Finn, 2026