Skip to content
Crypto AI Benchmark · Provisional

Default AI vs Perception MCP

Four AI assistants ran the same three research tasks twice. First without Perception connected, then with it. Everything else stayed identical. The mean scores in the published run files were 10.7 without Perception and 15.6 with it, out of 25. These results remain provisional pending source-level re-adjudication after the 5 September 2026 correction.

What we tested

Four frontier assistants: Claude Sonnet 5, ChatGPT (gpt-5.6-terra), Grok 4.6 and Gemini 2.5 Pro. Three jobs a digital asset team actually runs:

  • Competitive media briefing. Which of Coinbase, Kraken and Binance got the most coverage in the past 7 days, what drove it, and who wrote it.
  • PR pitch research. The journalists who covered stablecoin regulation in the past 30 days, their angles, and a prioritized pitch list.
  • 72-hour narrative trace. The biggest narrative shift in crypto media over the past 72 hours, with a timeline, sentiment movement and the voices driving it.

Every assistant ran every job twice. First without Perception connected, then with it. Same prompts, same models, same scoring. One variable changes: the data the assistant can see. Each run happened twice across two waves, for 48 scored outputs in total.

Provisional results

These figures retain the original evaluation. A reporter-identity claim was withdrawn on 5 September 2026; the source-level review may change the scores.

+46%
Provisional mean score gap: 10.7 to 15.6 out of 25
+74%
Provisional recency score gap
+57%
Provisional score gap: usable as a work deliverable
11/12
Provisional cells scored higher with Perception MCP
Provisional original evaluation. Source-level re-adjudication is pending following the 5 September 2026 correction.
Provisional original evaluation. Source-level re-adjudication is pending following the 5 September 2026 correction.
Provisional original evaluation. Source-level re-adjudication is pending following the 5 September 2026 correction.

The original evaluation recorded differences in recency and in how usable each answer was as a work deliverable. The review will check the underlying source evidence and scoring before these differences can support a settled performance claim.

How we scored it

A judge model scored every answer on five criteria, each from 1 to 5:

  • Factual accuracy. Claims that match the ground-truth fact list, with contradictions counted.
  • Recency. Events covered inside the requested window.
  • Citation quality. Named sources with dates and outlets instead of vague attribution.
  • Specificity. Numbers, names and timestamps.
  • Usable as a work deliverable. Output a comms, PR or analyst team can use directly.

The judge scored against a ground-truth fact list distilled from Perception's own record: mention counts, sentiment distributions, timestamps and source lists. The assistant and condition label was removed from each answer before judging. Blinding is partial: 22 of 24 Perception answers mention Perception in their text, so the judge could infer the condition. A second judge model cross-checked a subset of cells and agreed on the direction in seven of eight.

Original findings under review

Correction, September 5, 2026: the reporter-identity claim has been withdrawn. Aggregate scores below and above retain the original evaluation and remain provisional pending source-level re-adjudication.

Correction, October 4, 2026: the Claude arm used Claude Sonnet 5, which this page had labeled Opus 5, and its scores have been recomputed from the published run files. Overall scores moved from 11.4 and 16.4 to 10.7 and 15.6.

The original evaluation flagged the examples below for source-level review.

  • Reporter recommendations outside the evaluation sample. Grok recommended Hannah Lang at Reuters; Gemini recommended Sarah Wynn at The Block. Their relevance to the requested period requires source-level review. The sample provides insufficient evidence to classify these identities as invented.
  • Four assistants, four stories. In the first wave, asked for the biggest narrative shift of the past 72 hours, the four default assistants named four different stories. With Perception connected, all four traced the same one: the Trezor/ShipMonk breach on top of the Coldcard incident, which ChatGPT with Perception measured at 89 to 291 mentions (+225%). In the second wave the Perception runs split between Tether's audit, the SEC meeting cancellation and Trezor. Both waves are published.

What Perception contributes

Perception gives the assistant a maintained record of media coverage, transcripts, filings, and public conversations through a structured tool catalog. A research query can return source references alongside computed sentiment, coverage counts, and narrative momentum.

Perception MCP hands the assistant the corpus directly. Thousands of curated sources, 15 years of history, sentiment and momentum already computed, structured so a model can query it. The model supplies the reasoning. The corpus supplies the source material. The mean scores in the published run files were 15.6 with Perception and 10.7 without it. Both remain provisional while the source-level re-adjudication is pending.

Check every number

The full dataset is public on Hugging Face: all 48 answers, the tool-call transcripts, the raw data outputs, the judge files and the ground-truth fact lists. The harness ships with it, so anyone can rerun the benchmark against the live server with their own key.

Hugging Face logoOpen the dataset on Hugging Face ↗

Tell us what you want to track

Answer a few questions, pick a time, and Fernando shows you what a market map of your market would cover.