Verified data for frontier AI

The data AI can't find on the open web — and proof that it's real.

Nummero sources proprietary datasets from the companies that own them, verifies every one against a published standard, and licenses them to AI labs.

Repositories
Expert annotations
Workflow records
✓

Verification Standard

8 checks · results travel with data

Delivered

Model-ready dataset

Licensed · provenance attached

The Data

Proprietary data for the next generation of AI.

The most valuable training data lives inside companies and specialised domains, not on the public internet.

main · history

Private code and engineering data

Production software and the history behind it. Repositories, commits, pull requests, issue histories, tests and review activity. Every repository measured for authored versus generated lines before it is offered.

A ✓B

Domain expert data

Credentialed professionals in medicine, law, finance, accountancy and engineering, for preference data, evaluation and red-teaming.

$ run task_042

step 1 … ok

step 2 … ok

reward: 1.0

Verifier

✓assert output
✓check state
✓score

RL environments and verifiers

Executable task environments and expert-written verifiers, built by domain specialists working alongside software engineers.

SOP
Ticket
Thread

Enterprise workflow data

Real work produced inside real organisations: documents, support interactions, project activity and standard operating procedures.

The Nummero Verification Standard

Everyone says their data is clean. We publish how we check.

Every dataset is measured against a versioned public standard before it reaches a buyer, and the results travel with the dataset.

A worked example

"A supplier portfolio was presented at 200 million lines. The standard established 1.5 million."

In one repository, 15,093 logical lines sat beneath 20,633,590 lines of generated and vendored code — a ratio of 1,367 to 1. In another, 6,147 against 3,032,810. None of it visible in the headline figure.

Authored versus generated lines

separates hand-written code from data dumps, assets and build output

Commit-history plausibility

flags volume that cannot be reconciled with authorship over time

Multi-root detection

finds repositories holding several projects measured as one

Licence and third-party contamination

identifies purchased themes, vendored libraries and copyleft exposure

Provenance scrubbing

removes supplier identifiers, machine paths and internal names before delivery

Duplication on a matched file set

ensures duplication and line counts describe the same codebase

Build and runnability verification

confirms the tree builds from a clean clone

Rights and ownership confirmation

written authority from the owner before anything moves

Process

From raw supply to model-ready datasets.

  1. 01

    Define

    We work with your team to understand the capability, use case and specification.

  2. 02

    Source

    We identify datasets across our network of companies and supplier organisations.

  3. 03

    Verify

    Every dataset is measured against the published standard. Results travel with the data.

  4. 04

    Prepare

    Anonymisation, normalisation, structuring and filtering for downstream use.

  5. 05

    License and deliver

    Commercial rights coordinated, data delivered to agreed technical and legal requirements.

Commitments

Provenance is verifiable, not asserted.

Licensed from the owner

Every dataset licensed directly from the company or people who own it, with rights and consent confirmed in writing before anything moves.

Nothing scraped

No crawled corpora, no aggregator inventory, no grey-market data.

Measured, not claimed

Every figure we quote is reproducible from the raw tool output, which we supply alongside it.

For Data Owners

You may be holding the data labs need.

Many companies have spent years building datasets with real value for AI development. Nummero helps establish that value, find appropriate buyers, structure the licensing and deliver securely — without selling the underlying business or its intellectual property.

New revenue

License historical or continuously generated datasets to qualified AI companies.

Keep control

Define exactly what may be used, by whom, and under what rights.

We handle the complexity

Measurement, buyer requirements, processing, negotiation and delivery.

License your data

We will measure your codebase free of charge and tell you what it is actually worth.

About

Built by operators, not brokers.

Nummero is based in Bengaluru. The team has spent years building and marketing software, and has supplied verified code datasets to AI training companies.

The verification method behind the Nummero Verification Standard was developed by auditing repositories against a frontier lab's own measurement harness — the same checks buyers run, applied before the data is sold, not after.

Abdul Salim

Chief Executive Officer

Tell us what your models need next.

We source against a specific requirement — capability, domain, scale and format.

or email sales@nummero.com