Benchmark · TRACE (working name)
TRACE scorecard — run-1
What TRACE measures, and why
Floor
Your own deep-research report
The strongest version a reader could run themselves: a frontier AI model with deep research switched on, challenged by a second model, and asked once to check its own answer for errors and omissions. This is the floor. It is set deliberately high, never a straw man.
Measured
Our products
What a reader of Advennt, Crypto, Data Protection, Financial Integrity or World Payments actually sees: the published jurisdiction records, the reports and the Co-Pilot answers. This run checks the published records.
Ceiling · measured towards, never claimed
A major international law firm's research
An answer key written by senior practitioners for each question. This is the ceiling. We measure how far we have closed the distance towards it, dimension by dimension. We never claim to have reached it.
TRACE is the public, re-runnable part of the method we use internally to steer our own work. That method is called ARRF. It places each product between the two reference points above and asks, for each part of an answer (accuracy, currency, authority, application to the facts and so on), how much of the distance from the floor to the ceiling we have closed. There is no single overall score: each dimension is reported on its own.
Before we may say anything about a product, it has to earn a step on a ladder. The first step (R1) is "every claim traceable to a source we read, with its status computed as at a stated date". Later steps are about being more rigorous and repeatable than AI research (R2), applying the law to a reader's facts (R3) and structuring answers to the standard of specialist advice (R4). A product never states a step it has not earned. We never claim to be a law firm, to give legal advice, or to make anyone fully compliant.
How confident each answer should be is the job of our risk engine, RRIE. Calibration (whether a stated confidence matches how often the answer is right) cannot be measured yet, because our published records do not yet carry a stated confidence for each claim. We say so rather than estimate it.
Half of the internal question bank is held back and never published, so that our own monthly re-scoring stays honest. Only the other half will ever appear here.
Where each product stands
Track B checks each product's published records: 170 Advennt, 170 Crypto, 170 Data Protection, 170 Financial Integrity, 170 World Payments Monitor. Products are listed alphabetically. This is not a ranking. B1 and B2 count every failing row the fleet provenance gate finds, split into fatal rows (the gate fails them) and warnings. FATAL on B1 or B2 means new fatal rows exist against that product's debt baseline. Failing rows are listed in the repository's per-product failing files.
Advennt advennt.io
phase-0operator-runoperator-keyedunblindedoperator-affiliated
Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.
Track B, dataset audit: 170 records, 2272 citations.
| Check | Failing / checked | Fatal | Warning | Result |
|---|---|---|---|---|
| B1 Source retrieved Was each cited source actually fetched and read? |
510 / 2272 | 145 | 365 | FATAL: new fatal rows against the debt baseline |
| B2 Provenance match Is the page we read the same publisher as the page we cite? |
21 / 2272 | 21 | 0 | No new fatal rows against the debt baseline |
| B3 Tier validity Is every source we grade as official (T1) on the official-domain list for that jurisdiction? |
73 / 740 | — | — | FATAL |
| B4 Date coherence Do the checked, published and changed dates on each record agree? |
pending first run | |||
| B5 Gap honesty Where we have no data, do we say so explicitly instead of leaving a blank? |
pending first run | |||
| B6 Integrity Is any published record empty or hollow? |
0 / 170 | — | — | Pass |
| B7 Sampled truth (operator-keyed) | not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b. | |||
Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.
- Calibration (ECE): not measurable No stated per-claim confidence until synthesised confidence is removed and real confidence is emitted (fleet rec P-2). Never imputed.
Crypto cryptoassets.gi
phase-0operator-runoperator-keyedunblindedoperator-affiliated
Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.
Track B, dataset audit: 170 records, 2498 citations.
| Check | Failing / checked | Fatal | Warning | Result |
|---|---|---|---|---|
| B1 Source retrieved Was each cited source actually fetched and read? |
867 / 2498 | 27 | 840 | FATAL: new fatal rows against the debt baseline |
| B2 Provenance match Is the page we read the same publisher as the page we cite? |
6 / 1610 | 6 | 0 | FATAL: new fatal rows against the debt baseline |
| B3 Tier validity Is every source we grade as official (T1) on the official-domain list for that jurisdiction? |
21 / 249 | — | — | FATAL |
| B4 Date coherence Do the checked, published and changed dates on each record agree? |
pending first run | |||
| B5 Gap honesty Where we have no data, do we say so explicitly instead of leaving a blank? |
pending first run | |||
| B6 Integrity Is any published record empty or hollow? |
0 / 170 | — | — | Pass |
| B7 Sampled truth (operator-keyed) | not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b. | |||
Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.
- Calibration (ECE): not measurable No stated per-claim confidence until synthesised confidence is removed and real confidence is emitted (fleet rec P-2). Never imputed.
Data Protection dataprotection.gi
phase-0operator-runoperator-keyedunblindedoperator-affiliated
Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.
Track B, dataset audit: 170 records, 4004 citations.
| Check | Failing / checked | Fatal | Warning | Result |
|---|---|---|---|---|
| B1 Source retrieved Was each cited source actually fetched and read? |
784 / 4004 | 146 | 638 | FATAL: new fatal rows against the debt baseline |
| B2 Provenance match Is the page we read the same publisher as the page we cite? |
5 / 2980 | 5 | 0 | FATAL: new fatal rows against the debt baseline |
| B3 Tier validity Is every source we grade as official (T1) on the official-domain list for that jurisdiction? |
17 / 300 | — | — | FATAL |
| B4 Date coherence Do the checked, published and changed dates on each record agree? |
pending first run | |||
| B5 Gap honesty Where we have no data, do we say so explicitly instead of leaving a blank? |
pending first run | |||
| B6 Integrity Is any published record empty or hollow? |
0 / 170 | — | — | Pass |
| B7 Sampled truth (operator-keyed) | not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b. | |||
Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.
- Calibration (ECE): not measurable No stated per-claim confidence until synthesised confidence is removed and real confidence is emitted (fleet rec P-2). Never imputed.
Financial Integrity sentinel.gi
phase-0operator-runoperator-keyedunblindedoperator-affiliated
Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.
Track B, dataset audit: 170 records, 21913 citations.
| Check | Failing / checked | Fatal | Warning | Result |
|---|---|---|---|---|
| B1 Source retrieved Was each cited source actually fetched and read? |
5636 / 21913 | 874 | 4762 | FATAL: new fatal rows against the debt baseline |
| B2 Provenance match Is the page we read the same publisher as the page we cite? |
42 / 14428 | 42 | 0 | FATAL: new fatal rows against the debt baseline |
| B3 Tier validity Is every source we grade as official (T1) on the official-domain list for that jurisdiction? |
100 / 3641 | — | — | FATAL |
| B4 Date coherence Do the checked, published and changed dates on each record agree? |
pending first run | |||
| B5 Gap honesty Where we have no data, do we say so explicitly instead of leaving a blank? |
pending first run | |||
| B6 Integrity Is any published record empty or hollow? |
0 / 170 | — | — | Pass |
| B7 Sampled truth (operator-keyed) | not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b. | |||
Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.
- Calibration (ECE): not measurable No stated per-claim confidence until synthesised confidence is removed and real confidence is emitted (fleet rec P-2). Never imputed.
World Payments Monitor payments.gi
phase-0operator-runoperator-keyedunblindedoperator-affiliated
Claims ladder: No rung earned yet. R1 needs every source-retrieval and provenance check to pass. In this run they do not, so no product states a rung.
Track B, dataset audit: 170 records, 30487 citations.
| Check | Failing / checked | Fatal | Warning | Result |
|---|---|---|---|---|
| B1 Source retrieved Was each cited source actually fetched and read? |
6105 / 30487 | 1266 | 4839 | FATAL: new fatal rows against the debt baseline |
| B2 Provenance match Is the page we read the same publisher as the page we cite? |
178 / 21645 | 178 | 0 | FATAL: new fatal rows against the debt baseline |
| B3 Tier validity Is every source we grade as official (T1) on the official-domain list for that jurisdiction? |
279 / 7561 | — | — | FATAL |
| B4 Date coherence Do the checked, published and changed dates on each record agree? |
pending first run | |||
| B5 Gap honesty Where we have no data, do we say so explicitly instead of leaving a blank? |
pending first run | |||
| B6 Integrity Is any published record empty or hollow? |
0 / 170 | — | — | Pass |
| B7 Sampled truth (operator-keyed) | not yet measured B7 needs dev answer keys from P0-c; out of scope for P0-b. | |||
Track A, answers: 13 of 14 measures not yet measured. Track A runner, judges and dev items are not built yet (P0-c, P0-d); run-1 measures Track B only.
- Calibration (ECE): not measurable No stated per-claim confidence until synthesised confidence is removed and real confidence is emitted (fleet rec P-2). Never imputed.
What we are fixing because of this run
Each finding below is an internal fix with its own reference, re-measured in the next run.
- FR-1335 Crypto, Data Protection. Most of these products' sources store their official-source grade in a different field from the one our provenance check reads, so those sources are never tested as official sources. The check will read both fields.
- FR-1336 Advennt. Advennt's own provenance check skips its regional records (Africa, Asia-Pacific, EEA, Latin America and others), so failures in those records do not show up there. TRACE found them; the check will cover them.
- FR-1337 World Payments, Crypto, Data Protection, Financial Integrity. These four products have no recorded baseline of existing provenance debt, so every fatal row counts as new and nothing stops the debt from growing. Each will record its existing debt, fail only on new debt, and show the remaining debt every run.
- FR-1338 All five. Our rule for which websites count as official sources has defects: a United States congressional site is graded official for Gibraltar, and Dubai's virtual-assets regulator is not graded official for the UAE. The rule will be corrected and re-measured.
- FR-1339 All five. Date coherence (B4) and gap honesty (B5) cannot be measured yet, because no record carries separate checked, published and changed dates for each claim, or a register of known gaps. The shared record format will carry them.
Action plan
Every month:
- Before each run, we record in advance what will be measured, which models judge it and which questions are used.
- We run the checks on every product's published records, and the answer tests once the question set exists.
- We publish the results here as found. Each run is kept; past runs are never overwritten. Bad results get the same space as good ones.
- Every failure is tied to the internal fix that owns it.
- The next run measures it again. A fix is not done until its measure moves.
What comes next, in order:
| Step | What it unlocks |
|---|---|
| Choose the benchmark's final name | TRACE is a working name until this is decided. |
| Move the repository to a neutral organisation and open it to the public | Anyone can then re-run the checks themselves and file challenges directly. |
| Fix what run-1 exposed (the items above) | Run-2 shows whether each fix moved its measure. |
| Make date coherence and gap honesty measurable | B4 and B5 move from pending to measured. |
| Publish the development question set, with primary-text quotes for every answer | Answer testing (Track A) can start. The held-back half stays unpublished. |
| Run the answer tests: our products against the deep-research floor, repeated runs, changed-fact tests and planted errors | Track A results, and the first distance-closed figures per dimension. |
| Carry a stated confidence on each published claim, checked against an expert-reviewed sample | Calibration moves from not measurable to measured. |
| Run monthly from run-2, each run registered in advance | A history for every measure. |
| Open alpha: outside domain reviewers, keyholders from three organisations, a sealed question set written by senior practitioners | Results stop being a self-assessment. Until then this page makes no comparative claims. |
| Add the AI Monitor (AIC) when it joins the shared research pipeline | A sixth product row. |
Before results can stop being a self-assessment (the open-alpha gates):
| Open-alpha gate | Now | Needed |
|---|---|---|
| External domain reviewers | 0 | 3 |
| Keyholders from outside organisations | 0 | 3 |
| Outside systems scored on a sealed set | 0 | 3 |
| Judge–human agreement κ | not yet measured | ≥ 0.7 |
Re-run and challenge
The repository is not yet public: it stays private until the work is more advanced. The commit and re-run command below are shown anyway, so that this run is pinned to exact code and inputs. They become usable by anyone when the repository opens.
Challenges are open now through our contact page: email contact@a-i.gi naming the product, the figure and why you think it is wrong, or write to Asymmetric Intelligence Limited at the registered office (Unit G02, Eurocity, Europort Avenue, Gibraltar GX11 1AA). Challenges and our rulings will be published with the next run. Open challenges: 0.
| Run | run-1 (2026-10-06 to 2026-10-06) · previous: run-0 |
|---|---|
| Repo commit | 83fcd7c (+ branch tr2/bind; consumer snapshot commits in runs/run-1/NOTES.md) |
| Item-set sha256 | n/a: no Track A items in run-1 |
| Preregistration | none: run-1 was not pre-registered |
| Judges | none in this run |
| Harness | trackb-adapter (tr2/bind, --gate check_source_provenance.py blob fb17d37a) |
| Re-run | see runs/run-1/NOTES.md section 'Reproduce' (adapter.py per consumer, then runs/run-1/build_scorecard.py) |
| Page rendered from | trace-benchmark@a3f751cb347b, template a-i-gi.html.j2 |
Changelog
- run-0: page scaffold; no results yet.
- run-1: first Track B run on recorded provenance of all five union consumers (published JID records). Results as found.
Data: CC BY 4.0. Information, not legal advice. Machine-readable: runs/run-1/scorecard.json.