ChatGPT Built a Private Highway Through the Search Desert

​

Unleash this on:

Original insights by Tomek Rudzki

30-second rundown

Key learnings:

  • ChatGPT appears to operate its own search indexes: internal records revealed stored content for the web, news, PDFs, shopping, local results, images, and other categories.
  • Cached pages can delay updates: ChatGPT may answer from a stored copy instead of fetching your live page, so a correction published today may not reach its answers immediately.
  • Google and Bing visibility is no longer enough: ChatGPT combines its own indexes with external providers and specialist sources such as Yelp and TripAdvisor.

This week: Choose five pages containing facts that influence a sale, such as prices, availability, product specifications, opening hours, or delivery areas. Ask ChatGPT questions those pages answer, record whether its information is correct and current, and check which sources it cites. Fix the page or third-party profile carrying the most damaging error.

Coming 3 months: Repeat the test monthly across ChatGPT, Google, Bing, and the third-party platforms customers use to compare you. Track outdated facts, missing pages, and incorrect profiles in one document, then assign one person to correct persistent gaps.


Peec AI found strong observational evidence that ChatGPT operates its own family of search indexes, internally labelled Labrador, while still pulling information from Google, Microsoft, and specialist providers.

A search index is a stored collection of webpages and information that a system can retrieve without searching the live web from scratch every time.

This changes the old map.

Ranking in Bing may help ChatGPT find you, but it does not prove that your latest page entered ChatGPT’s own index, survived its cache, or reached the model assembling the answer.

You can rank in the old search engines and still be missing, outdated, or misrepresented inside ChatGPT.

The Machine has built a private highway across the search desert. It still uses the public roads, but now it controls several hidden exits, a warehouse full of copied pages, and a toll collector with no published schedule.

1. ChatGPT Appears to Have Its Own Search Estate

For two years, the easy story was that ChatGPT borrowed Bing’s map, occasionally rifled through Google’s glove compartment, and called the resulting pile a search product.

Tomek Rudzki found a more complicated operation.

Between May 21 and July 21, 2026, ChatGPT’s server-sent events, or SSEs, exposed a field showing where retrieved information came from. SSEs are streams of data sent to the browser while ChatGPT assembles an answer.

Three recurring values pointed toward external search providers.

The fourth, Labrador, pointed inward.

Labrador did not appear to be one enormous warehouse. The observed events named separate indexes for the general web, PDFs, YouTube, news, arXiv research papers, Wikipedia, local information, finance, legal information, medical information, shopping, and images.

These are sometimes called vertical indexes. Each one stores information for a particular category, such as products, local businesses, or news, instead of throwing the entire web into one boiling vat.

The records included page content, crawl dates, and publication dates. That is the anatomy of a search index, not a polite forwarding service.

The evidence is observational rather than an official description of ChatGPT’s architecture. The system can change, and active experiments may show different machinery to different users. But the blood trail does not end with one data stream.

OpenAI job listings have sought engineers to build indexing systems, retrieval pipelines, serving layers, and vector stores. A retrieval pipeline is the machinery that finds, filters, and ranks information before passing selected evidence to the model.

Nick Turley also testified during the Google antitrust proceedings that OpenAI began building a search index in 2023 after encountering quality problems with external data.

ChatGPT appears to use hybrid retrieval: it combines its own stored information with outside search engines and specialist providers instead of trusting one source for every question.

A Bing ranking is useful. It is not a stamped visa proving admission to Labrador.

2. Caching Creates a Dangerous Delay

The shopping experiments exposed the machinery more clearly.

During August and early September 2026, Peec AI observed experiment names including prefer-index-over-serp and shopping-index. Some searches appeared to favour OpenAI’s own product index over ordinary search results.

Rudzki and his colleagues also tested a locked-down ChatGPT mode that could not browse live webpages. It still returned stored versions of major publisher sites.

That points toward caching, which means saving a copy of a page so it can be retrieved later without fetching the live version again.

A separate billion-page experiment observed GPTBot crawling millions of pages at tens of thousands of requests per hour. The most reasonable inference is that OpenAI stores webpage content for later retrieval and filtering.

This creates a gap between changing your website and changing what ChatGPT knows.

A hotel may update its website to say the swimming pool is closed for renovation. A cached copy may still tell customers the pool is open. A retailer may reduce a warranty from five years to three while the old promise continues wandering through AI answers like a drunk carrying yesterday’s newspaper.

Important facts should appear in clean, server-rendered text. Server-rendered text is included directly in the page code when it loads, making it easier for crawlers to access than information inserted later by complicated scripts.

Use visible update dates where they help readers understand freshness. Keep important claims specific and consistent. Do not bury a changed price, opening time, or product specification inside an image, rotating banner, or digital trapdoor.

Updating the page is only the first step. The stored copy must also be replaced.

3. Outside Sources Still Control Part of the Story

ChatGPT’s internal index does not make Google, Bing, or third-party platforms irrelevant.

Peec AI found signs of Google-backed retrieval and Bing use within Deep Research. The observed system also pulled information from Yelp, TripAdvisor, Google Maps routes, and additional scraping providers.

This is a federated retrieval system. It gathers evidence from several indexes and providers, standardizes the results, and ranks the combined pile before the model writes its answer.

That makes off-site accuracy part of GEO, not merely a reputation chore.

A restaurant can publish the correct opening hours on its website while an outdated Yelp listing tells ChatGPT the doors are closed. A hotel can describe its refurbished rooms accurately while an old comparison page keeps feeding the Machine a version of the property that died two renovations ago.

Local businesses should keep major directory and review profiles correct. Product companies should monitor the comparison sites, retailers, and reviews that shape purchase decisions. Publishers should check whether their strongest facts appear accurately in both conventional search and AI answers.

Your website is not the only witness being questioned.

The old search map has not disappeared. ChatGPT built a private highway across it, connected it to several outside roads, and started moving evidence between them after midnight.

The winning strategy is not to guess which provider will rule forever.

Keep important facts crawlable. Date information when freshness matters. Correct the third-party profiles customers trust. Test the questions that influence a purchase and record what the Machine actually retrieves.

Because the customer will not care which index supplied the wrong answer.

They will only remember that your business appeared to be wrong.

Unleash this on:
Avatar photo
WH.

All the paranoia of a field correspondent. None of the plane tickets.

WH. has spent 14 years inside the SEO machine and started The Vector Gazette, because he got tired of watching entrepreneurs make catastrophic decisions based on advice from people who discovered GEO last Tuesday.