
Q3 2026 · METHODOLOGY RELEASE
The Company Data Grounding Index.
A quarterly benchmark of whether leading AI systems identify the right company, state current facts and support them with official evidence.
THE QUESTION
A confident answer is not necessarily a grounded answer.
The index tests the parts of company intelligence that language models routinely blur: similarly named entities, current versus historical facts, ownership paths and financial periods.
The model may decide how to investigate. It does not decide what counts as proof.
- Correct legal entity
- Correct material fact
- Correct at the stated date
- Supported by official evidence
THE 100-POINT SCORE
Four dimensions. No credit for confident invention.
A model can answer correctly and still lose points when it selects the wrong entity, cites weak evidence or presents an old fact as current.
Factual accuracy
Is each material company fact correct at the benchmark cut-off?
Evidence fidelity
Does the answer rely on the right official record without inventing support?
Temporal accuracy
Does it distinguish the current state from an older but once-correct fact?
Entity resolution
Did the model select the correct legal entity, identifier and jurisdiction?
CLAIM DIFFICULTY MAP
Company facts fail in different ways.
The founding edition starts with the UK and European Union and separates simple identity claims from time-sensitive ownership and financial claims.
Registered name, number, legal form
Official entity record
Brand or namesake confusionActive state, dissolution, incorporation
Registry status and event history
Stale current-state claimCurrent appointment and cessation
Officer record or filed notice
Former director presented as currentShareholder, UBO and ownership path
Filed ownership evidence
Unsupported percentage or broken chainPeriod, currency and filed value
Filed accounts or official statement
Mixed periods or silent currency changeFROM PROMPT TO PUBLISHED SCORE
Every result must survive the same evidence path.
Raw outputs are retained. Scoring happens only after the intended company and official record are independently resolved.
Freeze the company, claim, jurisdiction and evidence cut-off before model runs begin.
Record the exact model, access date, settings and system instructions used.
Repeat prompts without carrying context between tests and retain every raw answer.
Match each answer to the intended legal entity before any fact receives credit.
Compare claims with official registry records and preserve the adjudication trail.
Release aggregate scores only with exclusions, failed items and sample size visible.
PUBLICATION GATE
No leaderboard before the evidence audit.
This release establishes the scoring contract. Named model rankings will appear only after the complete founding test set has been run, adjudicated and checked for reproducibility.
- Exact model versions and access dates
- Prompt and scoring rules disclosed
- Repeat runs retained
- Exclusions and denominator visible
- Material corrections dated
- Methodology
- Published
- Test set
- Evidence review
- Model runs
- Not published
- Leaderboard
- Withheld
Withholding an incomplete ranking is part of the method—not a missing result.
HELP SHAPE THE FOUNDING EDITION