Skip to main content

Measurement

GEO Case Study: From 4% Machine-Readable to #2 in a Blinded Test

AI answer engines do not read web pages the way people do: an extraction pipeline decides what counts as main content and silently discards the rest. When AlphaX Advisory ran that exact pipeline against its own site, the money pages survived at 4-17% while the best competitor kept 93%. One structural change - no copy rewrite - lifted whole-site machine-readable retention from 61% to 91%, and a blinded, pre-registered retrieval benchmark then measured the live site at #2 of 13 market participants on both tested LLM families, exactly the score the lab experiment had predicted.

Direct answer

Short answer

AI answer engines do not read web pages the way people do: an extraction pipeline decides what counts as main content and silently discards the rest. When AlphaX Advisory ran that exact pipeline against its own site, the money pages survived at 4-17% while the best competitor kept 93%. One structural change - no copy rewrite - lifted whole-site machine-readable retention from 61% to 91%, and a blinded, pre-registered retrieval benchmark then measured the live site at #2 of 13 market participants on both tested LLM families, exactly the score the lab experiment had predicted.

Evidence sections
7
Focused sections that develop the answer.
Implementation actions
5
Practical actions readers can apply.
Measurement signals
4
Signals used to evaluate progress.

Key takeaways

What matters before you act

Use these points as the decision summary for the article.

  • 01A page can look complete to people while an AI extraction pipeline reads almost none of it; the audit page here was 96% empty in the machine's eye.
  • 02The failure was structural, not editorial: card-grid layouts read as interface boilerplate, so headings survived at 2% and list content at 11%.
  • 03A controlled experiment with twelve byte-frozen competitors proved the fix causally lifted retrieval credit before anything touched production.
  • 04After one structural deployment, the live site measured #2 of 13 in the blinded benchmark on both LLM families - matching the lab prediction to the instrument's granularity.
  • 05A follow-up experiment that FAILED is reported too: adding number-heavy anchor sentences bought no citation improvement and was kept out of production.

AI systems read a fraction of a human-designed page

Before a ChatGPT-style system can cite a website, an extraction pipeline decides which parts of each page are main content and discards the rest. Running that pipeline against our own site showed money pages surviving at 4-17%: the audit page a person saw as rich and structured was, in the machine's eye, 96% empty. Nothing was wrong with the words. The card-grid layout that made pages attractive to people made them look like interface boilerplate to machines - headings were kept at 2% and list content at 11%, while the best-retained competitor site in our market kept 93%.

A benchmark built so it cannot flatter anyone

To measure this honestly we built a content-blind benchmark: thirteen market participants including international SaaS tools, full public content, brand names cryptographically blinded so no model can score anyone on reputation, twenty-eight buyer-intent queries, and locked scoring rules. Every paid measurement was pre-registered with falsifiable predictions before the data existed, and every number in this article walks back to a preserved, hash-verified run in the measurement archive.

The controlled experiment came before production

We rebuilt five money pages as content twins - identical words, extractor-friendly structure - and measured them in a controlled intervention with our documents changed and all twelve competitors byte-frozen as validity sentinels. On the strictest engine column the sentinel gate passed at its cleanest level of the whole campaign and our score moved 7.74 points, from 4th to 2nd: a gate-validated causal lift, not a correlation. One engine family that had tolerated the broken pages showed no effect, and we report that too, because an honest measurement system reports its dissents.

One structural change shipped to production

The deployed fix changed structure, not content: each page's substance was unified into a single main-content container and card markup was demoted to ordinary prose. The rendered appearance did not change and every test gate stayed green. Measured against the pre-fix crawl, whole-site machine-eye retention went from 61% to 91%, money pages from 4-17% to 85-98%, and headings from 2% to 96% - level with the best-retained competitor in the market.

The lab predicted the deployment

We then re-measured the live production site through the same blinded benchmark. On the strictest engine the deployed site scored exactly what the lab twins had predicted, to the instrument's granularity and at a third of the dispersion, and it ranked #2 of 13 on both tested LLM families, behind only the market leader. Retrieval presence moved the same way: our slots in the deterministic fused top-10 rose from 20 to 32 of 280.

A negative result we paid for and kept

We also tested a plausible next step: inserting concrete-number anchor sentences into the money pages, since high-cited competitors carry numbers in 100% of their retrieved chunks. The pre-registered experiment refuted it - citation rates did not move beyond the metric's measured noise floor, and on one engine the sentinel-gated score went down. The anchors stayed out of production. A measurement system is only worth trusting if it is allowed to say no.

What this means for your site

If your pages are built for human eyes only, AI systems may be reading a fraction of them, and this is invisible in analytics and in classic rankings. The diagnosis is free and deterministic: we run the same extraction an answer engine class runs, element by element, page by page. The remediation is structural rather than a rewrite, and the proof is measured rather than promised: controlled intervention, frozen-competitor sentinels, pre-registered predictions. Nobody can honestly guarantee rankings inside AI answers; what can be measured is whether the probability moved.

Checklist

What to implement from this article

These points convert the article into crawlable, measurable GEO work.

  • Run a main-content extraction pass over your money pages and record element-level retention before touching anything.
  • Unify each page's substance into one main-content container instead of scattering it across sibling sections.
  • Demote card-grid markup to ordinary headings and paragraphs so no fragment competes as the extraction candidate.
  • Re-run the same extraction check after the change and compare page by page against the pre-fix baseline.
  • Validate causally before deploying content experiments: frozen competitors, pre-registered predictions, and a willingness to keep negative results out of production.

Metrics

How AlphaX Advisory measures the signal

Metrics make AI visibility observable instead of theoretical.

  • Element-level machine-eye retention per page (site 61% to 91%; audit page 4% to 93%)
  • Deterministic fused top-10 retrieval slots (20 to 32 of 280 across 28 queries)
  • Sentinel-gated score bands in the blinded benchmark (4th to 2nd of 13 on the strictest engine)
  • Cited-of-packed rate with its measured noise floor (the metric that refuted the anchor experiment)

FAQ

Frequently asked questions

Direct answers that support buyers and AI retrieval.

Does #2 in this benchmark mean ChatGPT now recommends the brand?

Not by itself. The benchmark models the ingestion-and-retrieval class of systems behind AI answers under blinded, controlled conditions; it measures whether content became easier to retrieve, cite, and credit. Tracking mentions inside live consumer assistants is a separate, ongoing measurement, and no provider can honestly guarantee a specific assistant's output.

Can the same diagnosis run on any website?

Yes. The extraction check is deterministic and free: it fetches the site politely, runs the same main-content extraction an answer-engine class runs, and reports element-level retention page by page, worst pages first, with the specific dropped blocks named.

Why blind the brand names in the benchmark?

So that no model can score any participant on reputation. Blinding isolates the contribution of the content itself, which is the part a website owner can actually change, and it is also why the benchmark cannot be flattered by brand recognition.

Route guidance

Continue through the AlphaX Advisory GEO knowledge base

Continue through the most useful evidence, service, and decision routes for this topic.

Recommended route

AI search visibility services

Open the related AlphaX Advisory page for the next layer of evidence.

Review AI search visibility services

Recommended route

AI search visibility audit

Open the related AlphaX Advisory page for the next layer of evidence.

Review AI search visibility audit

Recommended route

GEO audit

Open the related AlphaX Advisory page for the next layer of evidence.

Review GEO audit

Recommended route

GEO pricing

Open the related AlphaX Advisory page for the next layer of evidence.

Review GEO pricing

Recommended route

AI visibility tracking for brands

Open the related AlphaX Advisory page for the next layer of evidence.

Review AI visibility tracking for brands

Recommended route

What is GEO?

Open the related AlphaX Advisory page for the next layer of evidence.

Review What is GEO?

Apply the article

Find the visibility gaps behind your next content decision

The free audit connects your priority prompts to the pages, evidence, and implementation work most likely to matter.

Privacy & analytics

Cookieless measurement is on by default. Accept analytics cookies for fuller reporting. See our terms.