The record

Calls made before the outcome was known.

Anyone can explain a programme after it has failed. The only honest test of a method is whether it says the right thing about a decision that has not yet been taken, in public, with a date on it, and with the conditions that would prove it wrong written down in advance.

This page holds every forward call MOSAIC has published on a live programme. Entries are added and never edited. Each one names the verdict, the condition it rests on, the evidence that would show us to be wrong, and the date on which it can be checked. When the check date passes the outcome is recorded here, whichever way it goes.

Every entry is falsifiable

A call that cannot be shown to be wrong is not a call. Each entry states, in advance, the specific published evidence that would count as MOSAIC having been mistaken. No entry is written so that any outcome can be claimed as a success.

Nothing is edited after publication

Entries are appended. Wording is never revised, verdicts are never softened, and no entry is removed once the check date has passed. Where we are wrong, the entry stays with the outcome recorded against it.

Every entry is independently timestamped

Each page is submitted to the Internet Archive on the day it is published, so the date and wording of every call can be verified by anyone, from a source that is not us.

The register

Two calls are open. Neither check date has passed.

Both programmes are read from the published record only, using evidence available at the decision point. Figures describe the size of the decision as it stood. Nothing here is a claim about the competence of any individual, and nothing here is a recommendation to any party.

Open
Published
17 August 2026
Check date
Within 90 days of the 2027–28 annual report

Prediction 01

NS&I Business Transformation

Read live against a completion date that has already moved more than once, currently stated as March 2028. Outline business case £400m against £841m reported.

The verdictPhased. The evidence supports continuing, but not committing the remaining migration tranches against the current date.
The conditionOne product held fully migrated for a complete reporting period, with volumes and error rates published, before the next tranche of migration is committed.
The callFull migration of all products will not be reported as complete on or before 31 March 2028 without a further re-baselining of the date, the scope, or both.
What would prove us wrongNS&I confirms in its published annual report and accounts that migration completed on or before 31 March 2028, against the scope as stated in August 2026, with no further re-baselining.
SourcesNational Audit Office reporting on NS&I Business Transformation; NS&I annual report and accounts.
OutcomeNot yet known. Recorded here when the check date passes.
Open
Published
17 August 2026
Check date
30 June 2028

Prediction 02

US Department of Veterans Affairs, electronic health record

Read live at the 2026 restart, after a three year pause. Department estimate $16.1bn against an independent cost estimate of $49.8bn.

The verdictPhased. The restart is sound. Scheduling further sites before adoption is evidenced is not.
The conditionAt the restarted sites, clinician adoption evidenced by measured order routing accuracy and a closed loop alert system, sustained across one full quarter, before any further site is scheduled.
The callFurther deployment sites will be scheduled before that condition is evidenced as closed, and a further pause, rebaseline or schedule slip will be reported within twenty four months of the restart.
What would prove us wrongThe department completes its next tranche of deployments on the timetable published at the restart, with no pause and no rebaseline, and Inspector General reporting over that period shows no recurrence of order routing defects.
SourcesDepartment of Veterans Affairs Office of Inspector General reporting; independent cost estimate; congressional oral evidence.
OutcomeNot yet known. Recorded here when the check date passes.

Why only two. A forward call requires a live programme, a decision that has not yet been taken, and enough published evidence to reason from. Most programmes fail one of those three tests. We would rather hold two calls we can be judged on than twenty we cannot.

How this will be scored

Published now, while the score is still zero.

Setting the rules before there is a number to report is the only moment at which they can be set honestly. These rules will not change once calls begin to resolve.

Rule 01

The count, never the percentage alone

Eleven of fourteen will always be reported as eleven of fourteen. At that sample size the true accuracy sits somewhere between 52 and 92 per cent, and a single call moves the headline figure by seven points. Around forty resolved calls are needed before the range is tight enough to claim anything, and eighty before it is tight enough to defend.

Rule 02

Accuracy is reported against a naive baseline

Roughly eighty per cent of large transformations miss their targets. Predicting failure every single time therefore scores around eighty per cent with no skill whatsoever. Any accuracy figure published here will be shown alongside what indiscriminate pessimism would have scored on the same set of calls, and the difference between the two is the only number that means anything.

Rule 03

Full pace calls are reported separately

Predicting that a large programme will struggle is cheap. Predicting that one will succeed, against a base rate that says it should not, is where skill actually shows. Every call to move at full pace is a bet against the base rate, and those calls are reported on their own line, because they are the ones that cannot be produced by temperament.

Rule 04

The condition is scored, not only the outcome

The verdict is packaging. The condition is the product. So the harder test is not whether a programme struggled, but whether the specific condition named in advance turned out to be the thing that bound. A call can be right about the outcome and wrong about the reason, and where that happens it is recorded as a miss.

Rule 05

The data is published, not only the method

Every call, date, source and outcome is downloadable, so that anyone can recompute the score and arrive at the same answer. The most cited dataset in this field is also the most criticised, precisely because its method is closed and its definition of failure is contested. A number nobody can check is a number nobody should trust, including this one.

Why this page exists

A method that will not be judged is not a method.

Assurance ratings on individual engagements are confidential, so firms selling assurance do not publish whether their own calls were right. Government does publish some of this: the Infrastructure and Projects Authority reports a Delivery Confidence Assessment for every major project each year. What is missing is a supplier publishing its own record against dated calls it made in advance. That is what this page is. This page is the only part of Sense Shift that cannot be argued with, because in time it will simply be a score.