Verified data for frontier AI
The data AI can't find on the open web — and proof that it's real.
Nummero sources proprietary datasets from the companies that own them, verifies every one against a published standard, and licenses them to AI labs.
Verification Standard
8 checks · results travel with data
Delivered
Model-ready dataset
Licensed · provenance attached
The Data
Proprietary data for the next generation of AI.
The most valuable training data lives inside companies and specialised domains, not on the public internet.
main · history
Private code and engineering data
Production software and the history behind it. Repositories, commits, pull requests, issue histories, tests and review activity. Every repository measured for authored versus generated lines before it is offered.
Domain expert data
Credentialed professionals in medicine, law, finance, accountancy and engineering, for preference data, evaluation and red-teaming.
$ run task_042
step 1 … ok
step 2 … ok
reward: 1.0
Verifier
RL environments and verifiers
Executable task environments and expert-written verifiers, built by domain specialists working alongside software engineers.
Enterprise workflow data
Real work produced inside real organisations: documents, support interactions, project activity and standard operating procedures.
The Nummero Verification Standard
Everyone says their data is clean. We publish how we check.
Every dataset is measured against a versioned public standard before it reaches a buyer, and the results travel with the dataset.
A worked example
"A supplier portfolio was presented at 200 million lines. The standard established 1.5 million."
In one repository, 15,093 logical lines sat beneath 20,633,590 lines of generated and vendored code — a ratio of 1,367 to 1. In another, 6,147 against 3,032,810. None of it visible in the headline figure.
Authored versus generated lines
separates hand-written code from data dumps, assets and build output
Commit-history plausibility
flags volume that cannot be reconciled with authorship over time
Multi-root detection
finds repositories holding several projects measured as one
Licence and third-party contamination
identifies purchased themes, vendored libraries and copyleft exposure
Provenance scrubbing
removes supplier identifiers, machine paths and internal names before delivery
Duplication on a matched file set
ensures duplication and line counts describe the same codebase
Build and runnability verification
confirms the tree builds from a clean clone
Rights and ownership confirmation
written authority from the owner before anything moves
Process
From raw supply to model-ready datasets.
- 01
Define
We work with your team to understand the capability, use case and specification.
- 02
Source
We identify datasets across our network of companies and supplier organisations.
- 03
Verify
Every dataset is measured against the published standard. Results travel with the data.
- 04
Prepare
Anonymisation, normalisation, structuring and filtering for downstream use.
- 05
License and deliver
Commercial rights coordinated, data delivered to agreed technical and legal requirements.
Commitments
Provenance is verifiable, not asserted.
Licensed from the owner
Every dataset licensed directly from the company or people who own it, with rights and consent confirmed in writing before anything moves.
Nothing scraped
No crawled corpora, no aggregator inventory, no grey-market data.
Measured, not claimed
Every figure we quote is reproducible from the raw tool output, which we supply alongside it.
For Data Owners
You may be holding the data labs need.
Many companies have spent years building datasets with real value for AI development. Nummero helps establish that value, find appropriate buyers, structure the licensing and deliver securely — without selling the underlying business or its intellectual property.
New revenue
License historical or continuously generated datasets to qualified AI companies.
Keep control
Define exactly what may be used, by whom, and under what rights.
We handle the complexity
Measurement, buyer requirements, processing, negotiation and delivery.
We will measure your codebase free of charge and tell you what it is actually worth.
About
Built by operators, not brokers.
Nummero is based in Bengaluru. The team has spent years building and marketing software, and has supplied verified code datasets to AI training companies.
The verification method behind the Nummero Verification Standard was developed by auditing repositories against a frontier lab's own measurement harness — the same checks buyers run, applied before the data is sold, not after.
Abdul Salim
Chief Executive Officer
Tell us what your models need next.
We source against a specific requirement — capability, domain, scale and format.
or email sales@nummero.com
