The Nummero Verification Standard

Everyone says their data is clean. We publish how we check.

The standard is a named, versioned set of checks. Every dataset is measured against it before it reaches a buyer, and the raw tool output travels with the dataset — so every figure we quote is reproducible, not asserted.

A worked example

"A supplier portfolio was presented at 200 million lines. The standard established 1.5 million."

In one repository, 15,093 logical lines sat beneath 20,633,590 lines of generated and vendored code — a ratio of 1,367 to 1. In another, 6,147 against 3,032,810. None of it visible in the headline figure.

  1. 01

    Authored versus generated lines

    Separates hand-written code from data dumps, assets and build output. Line counts are reported on authored source only.

  2. 02

    Commit-history plausibility

    Flags volume that cannot be reconciled with authorship over time. A repository's size must be explainable by its history.

  3. 03

    Multi-root detection

    Finds repositories holding several projects measured as one. Each root is measured and reported separately.

  4. 04

    Licence and third-party contamination

    Identifies purchased themes, vendored libraries and copyleft exposure before anything is licensed.

  5. 05

    Provenance scrubbing

    Removes supplier identifiers, machine paths and internal names before delivery.

  6. 06

    Duplication on a matched file set

    Ensures duplication and line counts describe the same codebase, measured against the same files.

  7. 07

    Build and runnability verification

    Confirms the tree builds from a clean clone, with the build log retained as evidence.

  8. 08

    Rights and ownership confirmation

    Written authority from the owner before anything moves. No authority, no dataset.