The record
Anyone can explain a programme after it has failed. The only honest test of a method is whether it says the right thing about a decision that has not yet been taken, in public, with a date on it, and with the conditions that would prove it wrong written down in advance.
This page holds every forward call MOSAIC has published on a live programme. Entries are added and never edited. Each one names the verdict, the condition it rests on, the evidence that would show us to be wrong, and the date on which it can be checked. When the check date passes the outcome is recorded here, whichever way it goes.
A call that cannot be shown to be wrong is not a call. Each entry states, in advance, the specific published evidence that would count as MOSAIC having been mistaken. No entry is written so that any outcome can be claimed as a success.
Entries are appended. Wording is never revised, verdicts are never softened, and no entry is removed once the check date has passed. Where we are wrong, the entry stays with the outcome recorded against it.
Each page is submitted to the Internet Archive on the day it is published, so the date and wording of every call can be verified by anyone, from a source that is not us.
The register
Both programmes are read from the published record only, using evidence available at the decision point. Figures describe the size of the decision as it stood. Nothing here is a claim about the competence of any individual, and nothing here is a recommendation to any party.
Prediction 01
Read live against a completion date that has already moved more than once, currently stated as March 2028. Outline business case £400m against £841m reported.
Prediction 02
Read live at the 2026 restart, after a three year pause. Department estimate $16.1bn against an independent cost estimate of $49.8bn.
Why only two. A forward call requires a live programme, a decision that has not yet been taken, and enough published evidence to reason from. Most programmes fail one of those three tests. We would rather hold two calls we can be judged on than twenty we cannot.
How this will be scored
Setting the rules before there is a number to report is the only moment at which they can be set honestly. These rules will not change once calls begin to resolve.
Eleven of fourteen will always be reported as eleven of fourteen. At that sample size the true accuracy sits somewhere between 52 and 92 per cent, and a single call moves the headline figure by seven points. Around forty resolved calls are needed before the range is tight enough to claim anything, and eighty before it is tight enough to defend.
Roughly eighty per cent of large transformations miss their targets. Predicting failure every single time therefore scores around eighty per cent with no skill whatsoever. Any accuracy figure published here will be shown alongside what indiscriminate pessimism would have scored on the same set of calls, and the difference between the two is the only number that means anything.
Predicting that a large programme will struggle is cheap. Predicting that one will succeed, against a base rate that says it should not, is where skill actually shows. Every call to move at full pace is a bet against the base rate, and those calls are reported on their own line, because they are the ones that cannot be produced by temperament.
The verdict is packaging. The condition is the product. So the harder test is not whether a programme struggled, but whether the specific condition named in advance turned out to be the thing that bound. A call can be right about the outcome and wrong about the reason, and where that happens it is recorded as a miss.
Every call, date, source and outcome is downloadable, so that anyone can recompute the score and arrive at the same answer. The most cited dataset in this field is also the most criticised, precisely because its method is closed and its definition of failure is contested. A number nobody can check is a number nobody should trust, including this one.
Why this page exists
Assurance ratings on individual engagements are confidential, so firms selling assurance do not publish whether their own calls were right. Government does publish some of this: the Infrastructure and Projects Authority reports a Delivery Confidence Assessment for every major project each year. What is missing is a supplier publishing its own record against dated calls it made in advance. That is what this page is. This page is the only part of Sense Shift that cannot be argued with, because in time it will simply be a score.
Every MOSAIC read asks an organisation to name, in advance, the one condition that must be true before more money moves, and to say what would prove the assumption wrong. This page applies that standard to Sense Shift.
It stays on the page. A register that only records successes is marketing. The value of this one depends entirely on the entries we would rather remove.
Not investment advice, not a recommendation to any party, and not a judgement on the people running these programmes. It is a call on organisational absorption capacity, made from published evidence, at a stated date.