Pre-registered study · n=73

We audited 73 public-sector courses. Nearly half cannot be moved.

Source files go missing. The developer who built it leaves. The plug-in it needed gets switched off. Then someone asks for a small change to a course nobody can edit any more, so it gets rebuilt from scratch, and paid for twice.

We wanted to know how common that actually is, so we measured it.

73

Packages audited across 7 publishers

46.6%

Shipped with no editable source

15.1%

Runtime already dead

70/73

Still SCORM 1.2

The sample

Seven publishers across four regions: US, EU, UK and Australia. We published our predictions, and what result would prove us wrong, before we opened a single package.

PublishernNo sourceDead runtime
US federal HR agency1989.5%52.6%
US federal health agency2157.1%0%
UK national security body728.6%0%
EU institutional academy2114.3%4.8%
Three others (UK, AU)50%0%

By the time we went looking, one publisher's index had been deleted. We got the packages back through the Internet Archive. That is the whole problem in one sentence.

Why the publishers are not named

We publish findings, not publishers, the same rule we apply to every course we audit. None of these organizations asked to be an example, none has been contacted, and naming them would attach a failure rate to an institution over work it has had no chance to answer for.

What this costs, stated plainly: you cannot independently re-run our sample from this page. The pre-registration, the sample frame, the coding rubric and the harness are all public, and any researcher who wants the package list can have it on request. We do not publish it as a league table.

Why this is a floor, not a ceiling

Most public-sector eLearning is streamed from a website, not handed out as a file. You can only download a package because somebody built it to be shared. That means it is likely to be newer, easier to move, and better looked after. So the courses we can reach are the ones most likely to prove us wrong.

If the best case is trapped at 46.6%, the streamed majority is almost certainly worse. We publish this as a floor, and we say so in the methodology.

One more number worth sitting with: every single one of these packages still tracks correctly. They were built right. They became unusable anyway.

Method, published in full

The pre-registration, the sample frame, the coding rubric and the harness are all public.

  • We published our predictions, and what would prove us wrong, before we started
  • The full scoring rules are public
  • We built the checks twice, in two different ways, and made sure they agree rule for rule
  • Every change we make is logged, and every one can be undone

What we have not published yet

A human coding layer. The current figures come from an automated harness. Coding a subset by hand would let us report inter-rater reliability, and would support the stronger claim the study was built to test: that a lot of these packages are still accurate and technically unusable, which is a more interesting problem than "old."

Until that exists, we report what the harness found and label it as such.


Methodology · Pre-registration · How Preflight works