Skip to content
The argument

The corpus is the bottleneck.

Generic AI keeps getting better at searching the web. The reach of that search is the limit, and a curated, structured, 15-year digital asset record sits outside it.

Example query

“What does Bitcoin mining sentiment look like this week?”

One question, asked in both tools. The figures below are an illustration of the shape of each answer.

Generic AI

web search

A paragraph citing a CoinDesk piece, a rewrite of that same CoinDesk piece, a contributor article paraphrasing both, and a listicle from a site built to rank.

What it never saw

  • Canaan's earnings call transcript
  • The 8-K Marathon filed Tuesday
  • The podcast interview with Riot's CEO
  • A dozen posts from serious mining analysts
  • The Bloomberg piece behind the paywall
  • A Substack essay that went viral off-platform

All of it sits far from page one, because none of it was written to rank.

Perception

curated corpus
Articles, last 7 days
847
Outlet-weighted sentiment
62% positive
Drivers
Canaan's beat, Riot's Texas expansion, the Marathon rumor
Most-cited analysts
Three, ranked
Narrative velocity
Bullish framing accelerating since Tuesday

Same question, and one answer you can put in a memo.

Where the signal lives

We read where Google’s crawl stops

Claiming better AI features ages badly. Frontier models will close that gap within a couple of releases. What stays true is what a model can reach, and these sources sit outside a web search.

  • SEC filings
  • Earnings call transcripts
  • Podcast interviews
  • Conference keynotes
  • 400+ active X accounts
  • Substack newsletters
  • Subscriber-only analyst notes
  • Paywalled Bloomberg and WSJ

What holds up

Four things a better model does not give you

Thousandsof curated sources

Coverage

Google surfaces what was written to rank: rewritten press releases, aggregator sites, contributor columns recycling other people’s reporting. The signal in digital assets sits outside that. We read where the crawl stops.

By handsource selection

Curation

Bloomberg counts. Reuters counts. The WSJ markets desk, CoinDesk’s regulatory beat, The Block’s research team, EDGAR, the BIS papers nobody reads. An SEO farm publishing 40 rewritten releases a day does not, even when Google ranks it. An unfiltered corpus hands you the loudest voices by default.

2M+enriched records

Structure

Every article carries sentiment, confidence, entities, narrative category, outlet, author and date, computed and stored in a queryable table. Ask about a quarter of Coinbase coverage and the answer is a distribution with confidence intervals, not a summary of the vibes.

15 yearsof history

History

Daily index values back to February 2018. You cannot build this retroactively or scrape it overnight, because the sources that mattered have moved their archives behind paywalls or let them rot. Our cutoff is three minutes ago.

ChatGPT guesses; Perception looks it up.

In practice

Briefing a PM on Circle

From a generic tool an analyst gets a plausible paragraph. Some names look right, one quoted figure is invented, and there is no outlet breakdown, no time series, and no way into a specific week.

From Perception she gets 90 days of per-entity sentiment split by outlet category, the ten stories driving the narrative ranked by citation, the analyst upgrade history on the same timeline, and the SEC insider filings annotated with that week’s coverage.

One of those survives the briefing.

Who this is for

People who need the receipts

Researching Bitcoin on a Sunday afternoon? ChatGPT is fine. It is free and it reads a lot of things.

Making a position decision, briefing a CEO, or writing 3,000 words that have to survive a fact-check is a different job. That one needs the outlets, the dates and the structure behind every claim.

The corpus is the bottleneck. Perception is the corpus.

Fernando Nikolic

Fernando Nikolic

Founder, Perception

Try the same question in both tools

14 days, no credit card. Ask anything and compare the answers yourself.