The corpus is the bottleneck.
Generic AI keeps getting better at searching the web. The reach of that search is the limit, and a curated, structured, 15-year digital asset record sits outside it.
Example query
“What does Bitcoin mining sentiment look like this week?”
One question, asked in both tools. The figures below are an illustration of the shape of each answer.
Generic AI
web searchA paragraph citing a CoinDesk piece, a rewrite of that same CoinDesk piece, a contributor article paraphrasing both, and a listicle from a site built to rank.
What it never saw
- Canaan's earnings call transcript
- The 8-K Marathon filed Tuesday
- The podcast interview with Riot's CEO
- A dozen posts from serious mining analysts
- The Bloomberg piece behind the paywall
- A Substack essay that went viral off-platform
All of it sits far from page one, because none of it was written to rank.
Perception
curated corpus- Articles, last 7 days
- 847
- Outlet-weighted sentiment
- 62% positive
- Drivers
- Canaan's beat, Riot's Texas expansion, the Marathon rumor
- Most-cited analysts
- Three, ranked
- Narrative velocity
- Bullish framing accelerating since Tuesday
Same question, and one answer you can put in a memo.
Where the signal lives
We read where Google’s crawl stops
Claiming better AI features ages badly. Frontier models will close that gap within a couple of releases. What stays true is what a model can reach, and these sources sit outside a web search.
- SEC filings
- Earnings call transcripts
- Podcast interviews
- Conference keynotes
- 400+ active X accounts
- Substack newsletters
- Subscriber-only analyst notes
- Paywalled Bloomberg and WSJ
What holds up
Four things a better model does not give you
Coverage
Google surfaces what was written to rank: rewritten press releases, aggregator sites, contributor columns recycling other people’s reporting. The signal in digital assets sits outside that. We read where the crawl stops.
Curation
Bloomberg counts. Reuters counts. The WSJ markets desk, CoinDesk’s regulatory beat, The Block’s research team, EDGAR, the BIS papers nobody reads. An SEO farm publishing 40 rewritten releases a day does not, even when Google ranks it. An unfiltered corpus hands you the loudest voices by default.
Structure
Every article carries sentiment, confidence, entities, narrative category, outlet, author and date, computed and stored in a queryable table. Ask about a quarter of Coinbase coverage and the answer is a distribution with confidence intervals, not a summary of the vibes.
History
Daily index values back to February 2018. You cannot build this retroactively or scrape it overnight, because the sources that mattered have moved their archives behind paywalls or let them rot. Our cutoff is three minutes ago.
ChatGPT guesses; Perception looks it up.
In practice
Briefing a PM on Circle
From a generic tool an analyst gets a plausible paragraph. Some names look right, one quoted figure is invented, and there is no outlet breakdown, no time series, and no way into a specific week.
From Perception she gets 90 days of per-entity sentiment split by outlet category, the ten stories driving the narrative ranked by citation, the analyst upgrade history on the same timeline, and the SEC insider filings annotated with that week’s coverage.
One of those survives the briefing.
Who this is for
People who need the receipts
Researching Bitcoin on a Sunday afternoon? ChatGPT is fine. It is free and it reads a lot of things.
Making a position decision, briefing a CEO, or writing 3,000 words that have to survive a fact-check is a different job. That one needs the outlets, the dates and the structure behind every claim.
The corpus is the bottleneck. Perception is the corpus.
Fernando Nikolic
Founder, Perception
Try the same question in both tools
14 days, no credit card. Ask anything and compare the answers yourself.