The transformation update reaches the board as a single colour. Somewhere beneath it sit hundreds of judgements: a delivery leader’s honest worry, a programme director’s balanced view, a portfolio office’s weighting model, an executive committee’s discussion of tone. Each round of consolidation is defensible on its own terms, and the artefact that survives the journey is a composite: a maturity score of 3.4, an amber portfolio status, a heat map whose cells have been negotiated into softer shades across three drafts. Directors receive it in good faith and read it the way it invites being read, as a summary. It is not a summary. It is the residue left after the information a board most needs has been processed out.
Carillion showed how much can disappear on the way up. In November 2016, an internal peer review of the Royal Liverpool Hospital contract concluded it was making a loss; management overrode the assessment and booked a healthy margin instead, a difference of roughly £53 million that reappeared, almost to the pound, in the July 2017 profit warning. The joint parliamentary inquiry concluded the mystery was not that the company collapsed but that it lasted so long. Delivery professionals have a name for the milder, everyday version: the watermelon, green on the outside and red inside, produced wherever reporting a problem costs more than concealing one. Boards know this failure mode, and the better ones now ask hard questions about whether bad news survives the climb. Candour, though, is only half the problem. Even when every score in the chain is honest, the arithmetic alone can destroy the signal.
Averaging is not simplification. It is information destruction.
Picture the assessment on a single slide: six dimensions of execution health, each scored out of 30, with scores of 27, 24, 22, 18, 25, and 23. The average is a shade over 23, and on most reporting scales that presents as comfortable, a number to note and move past. Whether the comfort is deserved turns on a question the slide never answers: how do the six dimensions relate to one another? If they stand alone, averaging is fair, because strength in one place genuinely offsets weakness in another. A diversified investment portfolio works this way, which is partly why the habit feels so natural around a board table. If, instead, the six are links in a chain, with nothing reaching the customer except by passing through all of them, then the average is a fiction. The organisation will deliver at 18, whatever the other five numbers say. And the board’s real question was never how much strength the organisation holds in total; it was always what the organisation will actually deliver. For a portfolio, the average answers that question. For a chain, the weakest link answers it, and the average is where the answer goes to hide.
This is not a pathology confined to careless companies. The UK government’s own reporting on its major projects portfolio shows the habit at its most institutionalised: the annual report on major projects explains how average portfolio ratings are calculated by assigning numbers to red, amber, and green assessments and dividing by the number of projects. The National Audit Office has spent a decade documenting what grows in that soil, finding that incentives towards over-optimism are strong, disincentives weak, and the overall picture of performance opaque even to the centre. The deeper cost of the composite is what it does to spending. A board reading an average funds improvement wherever progress is cheapest to demonstrate, which is usually high in the organisation, close to the reporting. A board reading the weakest score funds the crack.
The weakest layer sets the verdict
Edition 5 set out the Coherence Stack in full: six layers, from Purpose and Vision at the foundation, moving through Strategy, Strategic Outcomes, the Operating Model, and Leadership and Culture, to the Management System through which the organisation steers and learns. Each layer rests on the one below, and weakness low in the stack cannot be compensated from above. That ordering claim has a consequence for measurement, and it is worth stating as a rule. Within a layer, scores can be summed, because six statements about one layer are six sightings of the same thing, and summing protects the reading from any single harsh or generous judgement. Across layers, the logic reverses. The layers are not six sightings of one thing; they are six different things arranged in series, and value must pass through all of them to become results. Measurements average; chains fail at the weakest link. And an organisation is a chain: the six readings on the slide above were layer scores from the Stack, which makes it an operating-model story, capped at 18, and everything the average added was noise.
The verdict rule is as follows: the weakest layer sets the verdict, regardless of what the other five say. It also sets the ceiling on spend. A leadership programme commissioned above a cracked operating model buys better behaviour inside a structure that defeats it. A new set of objectives cascaded above an unchosen strategy gives every function a sharper way to measure incompatible things. Worse than wasted, the money manufactures the appearance of action while the fault ages, which is why boards that fund from the average so often find themselves approving the same category of remedy three years in a row. The reporting demand this implies is short enough to minute. Stop asking management for the score. Ask which layer is weakest, and what evidence supports the answer.
Statements you can falsify, not aspirations you can admire
Finding the weakest layer requires an instrument built for the purpose, and the construction matters more than the length. The Coherence Stack diagnostic, published alongside this edition, puts 36 statements to the board, six per layer: four testing whether the layer is sound, and two testing whether it transmits into the layer built on it, because layers usually fail at the joint before they fail in the middle. Every statement is written to be falsifiable against evidence the organisation already holds. The claim that “Proposals are regularly declined because they do not fit the strategy” can be contradicted by the investment committee’s minutes. The claim that “We could tell whether the strategy is working before the financial results arrive” can be contradicted by the board pack itself. “Bad news travels upward quickly and without penalty” can be contradicted by the last surprise.
The register is deliberate. Maturity models fail as diagnostic instruments because their language is aspirational, and agreeing with an aspiration costs nothing: every executive is on a journey towards level four. A falsifiable statement exacts a different price, because a director who scores it generously is asserting something a named document or a named decision can disprove. The full instrument, with the model, the questions to put to management at each layer, and the scoring and reading guide, is free to download here. It may be shared in full with your board.
Disagreement is not noise. It is the second finding.
How the instrument is completed determines what it can find. The protocol is independent scoring before any discussion: each director, each attending executive, and two or three senior delivery leaders complete it alone, and the pictures are compared afterwards. Divergence between the pictures is the second finding. When the board’s picture and the delivery leaders’ picture disagree, the disagreement usually points at Layers 5 and 6, because those are the layers where what the board is told and what the organisation experiences part company. A consensus workshop that scores the statements collectively produces the average by social means, performing in a meeting room the same information destruction the composite performs in a spreadsheet, with the added defect that the most senior voice weights the mean.
What a board does with a verdict
A verdict earns its place on the agenda only if it changes what the board does next, and it changes three things. It changes the diagnostic sequence: the presenting symptom is traced downward to the layer that produced it before any remedy is approved, because cracks surface above their cause and a board inspecting only the visible damage will keep funding repairs to the wrong floor. It changes the investment sequence: the change portfolio is reordered so that spend below the crack precedes spend above it, however visible and however sponsored the upper-floor initiatives may be. And it changes the assurance scope: internal audit or an external reviewer is commissioned against the weakest layer specifically, testing the two lowest-scoring statements against artefacts rather than assurances. A reporting demand, an investment sequence, and an assurance scope are three instruments a board controls without asking anyone’s permission.
Questions for directors
Has this board ever been told which layer of the organisation is weakest, rather than the average or composite score?
Of the remedies approved in the last two years, how many sat above the fault they were meant to repair?
Which current red or amber status has been traced below the layer where it presents?
If directors, executives, and delivery leaders scored the organisation independently, where would their pictures diverge, and who around the table would predict the divergence accurately?
What evidence, rather than assurance, supports the most recent green rating the board accepted?
The composite score survives in board packs because it suits everyone who touches it. It is easy to produce, comfortable to present, and reassuring to receive, and reassurance is the one thing a diagnostic must never offer. A board that adopts the verdict rule gives up a tidy number and gains something better: a reading that tells it where the organisation will fail, before the failure files its own report.
If this is the kind of working discipline you want your board’s strategy oversight built on, subscribe to Strategy in the Boardroom. I write for directors, executives, and advisers who believe strategy is a discipline the board must own, and each edition aims to change a decision, a demand, or a control. Edition 7, The Board’s Two Dashboards, takes the measurement argument further: why monitoring the health of the enterprise and steering strategic change are different jobs, and why most board packs cannot do both.




