AI systems read a fraction of a human-designed page
Before a ChatGPT-style system can cite a website, an extraction pipeline decides which parts of each page are main content and discards the rest. Running that pipeline against our own site showed money pages surviving at 4-17%: the audit page a person saw as rich and structured was, in the machine's eye, 96% empty. Nothing was wrong with the words. The card-grid layout that made pages attractive to people made them look like interface boilerplate to machines - headings were kept at 2% and list content at 11%, while the best-retained competitor site in our market kept 93%.
A benchmark built so it cannot flatter anyone
To measure this honestly we built a content-blind benchmark: thirteen market participants including international SaaS tools, full public content, brand names cryptographically blinded so no model can score anyone on reputation, twenty-eight buyer-intent queries, and locked scoring rules. Every paid measurement was pre-registered with falsifiable predictions before the data existed, and every number in this article walks back to a preserved, hash-verified run in the measurement archive.
The controlled experiment came before production
We rebuilt five money pages as content twins - identical words, extractor-friendly structure - and measured them in a controlled intervention with our documents changed and all twelve competitors byte-frozen as validity sentinels. On the strictest engine column the sentinel gate passed at its cleanest level of the whole campaign and our score moved 7.74 points, from 4th to 2nd: a gate-validated causal lift, not a correlation. One engine family that had tolerated the broken pages showed no effect, and we report that too, because an honest measurement system reports its dissents.
One structural change shipped to production
The deployed fix changed structure, not content: each page's substance was unified into a single main-content container and card markup was demoted to ordinary prose. The rendered appearance did not change and every test gate stayed green. Measured against the pre-fix crawl, whole-site machine-eye retention went from 61% to 91%, money pages from 4-17% to 85-98%, and headings from 2% to 96% - level with the best-retained competitor in the market.
The lab predicted the deployment
We then re-measured the live production site through the same blinded benchmark. On the strictest engine the deployed site scored exactly what the lab twins had predicted, to the instrument's granularity and at a third of the dispersion, and it ranked #2 of 13 on both tested LLM families, behind only the market leader. Retrieval presence moved the same way: our slots in the deterministic fused top-10 rose from 20 to 32 of 280.
A negative result we paid for and kept
We also tested a plausible next step: inserting concrete-number anchor sentences into the money pages, since high-cited competitors carry numbers in 100% of their retrieved chunks. The pre-registered experiment refuted it - citation rates did not move beyond the metric's measured noise floor, and on one engine the sentinel-gated score went down. The anchors stayed out of production. A measurement system is only worth trusting if it is allowed to say no.
What this means for your site
If your pages are built for human eyes only, AI systems may be reading a fraction of them, and this is invisible in analytics and in classic rankings. The diagnosis is free and deterministic: we run the same extraction an answer engine class runs, element by element, page by page. The remediation is structural rather than a rewrite, and the proof is measured rather than promised: controlled intervention, frozen-competitor sentinels, pre-registered predictions. Nobody can honestly guarantee rankings inside AI answers; what can be measured is whether the probability moved.