Original insights by Metehan Yeşilyurt
30-second rundown
Key learnings:
- ChatGPT appears to use hybrid retrieval: it can use passages already stored in its search index or fetch a page and divide it into passages during live retrieval.
- These passages are called chunks: every important chunk must make sense alone because the model may receive only selected pieces of a page.
- Relevance scores influence which chunks move forward: keep the subject, answer, comparison, and evidence together so the passage survives without its surrounding text.
This week: Choose one page that helps customers decide whether to buy. Read each important section without the paragraphs around it. If the meaning collapses, rewrite it so the product, claim, comparison, and evidence travel together.
Coming 3 months: Review your most commercially important pages monthly. Track whether the improved passages appear more consistently in ChatGPT, Gemini, and other AI answers for your priority customer questions.
Metehan Yeşilyurt and the Peec AI research team pried open a ChatGPT data stream and found a retrieval operation with more chambers than a crooked revolver.
They asked ChatGPT to compare wireless noise-cancelling headphones under $300. Five clean product cards appeared. Behind that respectable shop window, ChatGPT ran four searches and fetched 96 pages.
Those pages were not simply handed to the model from headline to final full stop. They were divided into chunks, scored, sorted, and pushed forward while weaker pieces vanished through a trapdoor.
A chunk is a smaller passage taken from a webpage. Chunking divides a page into those passages so a retrieval system can find and rank specific information.
The findings came from the ChatGPT interface on September 9, 2026, and the machinery may change. But the evidence left one ugly fingerprint on the wall: your entire page may never reach the model writing the answer.
1. ChatGPT Appears to Use Hybrid Retrieval
Hybrid retrieval means information can enter through more than one route.
Of the 96 pages in the headphone test, 76 appeared to be pre-indexed, meaning information from them had already been discovered and stored before the prompt. Their most relevant chunks arrived pre-selected with relevance scores.
Other pages were fetched and chunked during live retrieval. Two entrances, several trapdoors, and no promise that every page receives the same treatment.
The researchers found that chunks from the pre-indexed pages later appeared in the material prepared for the model in the same order and with almost identical relative scores. They travelled through the system like marked cards between dealers in a rigged casino.
This suggests ChatGPT was carrying selected passages forward instead of rereading every page as one complete document. A strong page can still disappear if its useful facts are trapped inside chunks that depend on missing context.
Your page may read beautifully from top to bottom. The Machine may rip out three passages, leave the rest bleeding beside the highway, and judge the whole operation by what survived.
2. Every Important Chunk Must Survive Alone
Consider a running-shoe product page that says:
It lasts up to 30% longer.
A person who read the previous paragraphs may know that “it” means the outsole, what shoe it is being compared with, and where the 30% came from. A retrieved chunk may arrive without any of that. Once the context is hauled away in an unmarked truck, the claim becomes a corpse without identification.
Now compare it with this:
The Trailcrusher running-shoe outsole lasted 30% longer than our previous model in a 500-kilometre laboratory abrasion test.
The product, claim, comparison, and evidence travel together. The Machine can tear that passage from the page and throw it across the desert, but it still answers the question.
The same rule applies to service pages, pricing explanations, comparisons, case studies, and FAQs. If a section could influence a sale, name the subject and keep the proof beside the claim.
Do not force the retrieval system to crawl uphill through three paragraphs to discover what “it,” “this,” or “they” means. It may grab the first usable fragment and abandon the rest under the merciless sun.
3. Relevance Scores Decide Which Chunks Move Forward
The retrieved chunks carried neural relevance scores, estimates of how closely each passage matched the user’s query.
One SoundGuys page appeared as three chunks in the search record. The same chunks appeared in the same order in the information prepared for the model, covering battery life, alternative purchases, and competing headphones. Their exact scores changed slightly when converted to a standardised scale, but their relative order remained almost identical.
That does not prove every ChatGPT search follows the same pipeline. It does suggest that pre-scored chunks can travel forward towards the final answer instead of every page being judged again from scratch.
The decimal arithmetic can remain in the evidence locker. The commercial verdict is cleaner: the page entered the final stretch as selected and ranked passages, not one sacred document.
There is no proven magic chunk length. Anyone selling one is peddling miracle tonic from the boot of a stolen Cadillac. The safer strategy is to answer one recognisable customer question at a time, use descriptive headings, name the product or service, keep proof beside claims, and publish original information worth retrieving.
The whole page still matters. It provides the context, authority, and evidence the retrieval system can use. But the whole page may never attend the final hearing.
Write the document, then inspect its most valuable chunks like severed evidence on a steel table. Each one must identify the subject, answer the query, and carry enough proof to survive alone under the desert sun.
Because when The Machine starts chunking, it will not remember your page. It will remember the pieces.


