Teneo
Back to Blog
How Does AI Search Fetch Web Pages? Why Real Connections Matter

How Does AI Search Fetch Web Pages? Why Real Connections Matter

Use CasesBeaconSeptember 2026ยท11 min read

ChatGPT, Perplexity and Google AI Mode rewrite your question, then fetch live pages. Where that fetch comes from decides which page comes back. A robots.txt survey of the top 300 sites shows why AI search is buying real connections.

By Teneo Protocol

Share

ChatGPT, Perplexity and Google AI Mode rewrite your question, then fetch live pages. Where that fetch comes from decides which page comes back. A robots.txt survey of the top 300 sites shows why AI search is buying real connections.

AI search fetches web pages in two steps. First it turns your question into one or more search queries and runs them against a web index to find candidate pages. Then, when the answer needs current or detailed information, it requests the actual page at the moment you ask, reads it, and cites it. ChatGPT, Perplexity and Google AI Mode all work this way, and all three say so in their own documentation. The second step is the one this post is about, because where that fetch comes from decides which page comes back.

What happens between the question and the answer?

The question is rewritten before anything is fetched. OpenAI's help centre says ChatGPT search "typically rewrites your query into one or more targeted queries" and sends them to partner search providers, then may send "additional, more specific queries" after reviewing the first results. Google describes the same pattern for AI Mode as query fan-out: "breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf". Perplexity's Deep Research "performs dozens of searches automatically" and "reads hundreds of sources", according to its help centre.

Those queries return candidate pages from an index. An index is a copy, and for anything that changes, a copy is not enough. So the product fetches the page itself. OpenAI runs a fetcher called ChatGPT-User "for certain user actions in ChatGPT", per its crawler documentation. Perplexity runs Perplexity-User, which "visits pages when users ask questions", per its bot page. Google lists Google-Agent among its user-triggered fetchers, "used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request". OpenAI and Perplexity both publish the IP ranges their fetchers use, so a website can tell exactly who is asking.

OpenAI's crawler documentation listing OAI-SearchBot, ChatGPT-User and GPTBot. Source: OpenAI.
OpenAI's crawler documentation listing OAI-SearchBot, ChatGPT-User and GPTBot. Source: OpenAI. View source

Does AI search read the live web or an index?

Both, and the split is documented. ChatGPT "may search the web automatically when your question would benefit from current information", according to OpenAI's help centre. Perplexity's help centre says it "uses advanced AI to search the internet in real-time". Google said at I/O in May 2026 that AI Mode's agents draw on "real-time info on finance, shopping and sports".

The index answers stable questions. The live fetch answers everything with a price, a stock level, a timetable, a review count or a news cycle attached. Those are the questions where a stale or partial page produces a confident wrong answer, because the model has no way to know what it was not shown.

Why does location change what an AI search fetch sees?

Because websites serve different pages to different places, and a fetch inherits the location of the connection it comes from. Google's own crawler documentation is the clearest statement of the problem. It says "the default IP addresses of the Googlebot crawler appear to be based in the USA", that for sites which adapt content to the visitor's location Google "might not crawl, index, or rank all your content for different locales", and that Google therefore also crawls "with IP addresses based outside the USA". Its advice to site owners about those crawls is to "treat it like you would treat any other user from that country".

AI search products meet the same fact from the user's side. OpenAI says ChatGPT "may use an approximate location based on your IP address to provide relevant local results", and that "a VPN or network location may affect the approximate location inferred from your IP address". The location of the person asking picks the query. The location of the fetch decides what the page contains. When the two differ, a person in Lagos gets a page priced and stocked for Virginia.

This is the first reason a data centre is the wrong place to fetch from. A data centre is in one place, and that place is usually not where the user is.

What happens when a website blocks the fetch?

The answer is built from whatever came back, and the model cannot see what was withheld. Websites have three ways to say no. They can ask, through robots.txt. They can enforce at the network edge, by refusing known bot addresses. And they can serve a different page to a request they distrust. All three are becoming more common.

We measured the first one. On 25 September 2026 we fetched robots.txt from the 300 most-visited domains on the Tranco list, from a home connection in Ireland. 152 of them served a robots.txt file. Of those 152, 24 fully disallow GPTBot, OpenAI's training crawler, and 25 fully disallow ClaudeBot. The search fetchers get a different reception: 9 fully disallow OAI-SearchBot and 12 fully disallow ChatGPT-User. 15 domains block GPTBot but leave OAI-SearchBot open, which is the clearest sign that site owners now distinguish training from search. PerplexityBot, which Perplexity describes as a search crawler, is the exception: 22 refuse it, about as many as refuse the training crawlers.

Bar chart of the 152 top-300 Tranco domains that served a robots.txt on 25 September 2026, showing how many fully disallow each AI fetcher. OpenAI's and Anthropic's training crawlers are refused two to three times as often as their search and user fetchers; PerplexityBot is the exception. Source: Teneo survey of Tranco list domains.
Bar chart of the 152 top-300 Tranco domains that served a robots.txt on 25 September 2026, showing how many fully disallow each AI fetcher. OpenAI's and Anthropic's training crawlers are refused two to three times as often as their search and user fetchers; PerplexityBot is the exception. Source: Teneo survey of Tranco list domains. View source

Sites want to be in the answer, then, and a share of them do not want to be in the model. But robots.txt is a request. OpenAI says of ChatGPT-User that because "these actions are initiated by a user, robots.txt rules may not apply". Perplexity says Perplexity-User "generally ignores robots.txt rules". Google says the same of its user-triggered fetchers: "because the fetch was requested by a user, these fetchers generally ignore robots.txt rules". So enforcement moves to the network layer, where the origin of the request is what gets judged.

That is where the rules are changing now. From 15 September 2026, new domains on Cloudflare block AI training bots and AI agent bots by default on pages that display ads. Search crawlers remain allowed. Cloudflare's own definition of the agent category is the fetch this post is about: it "covers automated activity acting in real time on a person's behalf, such as chat fetch bots and browser-use agents". In June 2026 Google added a Search Console toggle that lets a site opt out of appearing in, and grounding, its generative AI Search features. A fetch that arrives from a published bot address range, or from an address block a site has decided to distrust, is the fetch most likely to be refused or given less.

Cloudflare changelog entry setting new AI traffic defaults from 15 September 2026. Source: Cloudflare.
Cloudflare changelog entry setting new AI traffic defaults from 15 September 2026. Source: Cloudflare. View source

Why is demand for real connections rising?

AI search is the first of the four buyer groups in our guide to what bandwidth sharing is used for, and it is the one whose growth is easiest to measure from public numbers. OpenAI said on 31 August 2026 that ChatGPT has "more than 1 billion weekly active users". Google said on 3 June 2026 that AI Overviews "has over 2.5 billion monthly active users" and that AI Mode "has surpassed one billion monthly users", with AI Mode queries "more than doubling every quarter since launch".

Every one of those questions can fan out into several queries, and every query that needs a current page triggers a fetch. Cloudflare's measurement of its own network for July 2025 found OpenAI's crawlers requested about 1,091 pages for every visit they referred back to a site, and Perplexity's about 195. The number of pages machines read is already far larger than the number of people they send anywhere.

Put the two trends together. More people ask AI search for things that live on current pages. More sites refuse, challenge or thin out fetches that arrive from data-centre address space. The only supply that satisfies both is a real connection in the right place, requesting public pages the way a person in that place would. That is what bandwidth sharing provides, and it is why demand for it rises every year.

How does Teneo supply AI search with real connections?

On the supply side, Teneo Beacon is the app you install to share unused bandwidth with the Teneo network. It runs on Windows, macOS, Linux, Android and iOS. Sharing is opt-in, and you stop it by closing the app. The network only ever carries requests for public pages through your connection. It cannot read your files, accounts or browsing history. Operators are rewarded for keeping Beacon online, and the mechanics are on the Beacon page and in the setup guides.

On the demand side, buyers reach that supply through the Teneo Protocol, where hundreds of AI agents are deployed. The data platforms in scope are X, Reddit, TikTok, LinkedIn, Instagram, YouTube, Google Maps, Google Search and Yelp, all of which vary by region. Each request is paid per call in USDC through x402 and settled on-chain on peaq. The Agent SDK documents the order of events: users see the price before executing a task, and payment is verified and settled on-chain before the agent processes it. Builders set a price per unit for each command in their agent config, and a new agent starts private and moves to public after a review of up to 72 hours. The Amazon agent alone has processed more than 33,000 requests for product, search and review data, and every one of them needed the page a shopper in that market would see.

Frequently asked questions

How does AI search get its information?

It rewrites your question into search queries, runs them against a web index to find candidate pages, then fetches the pages it needs at the moment you ask and builds the answer from what came back, with citations. ChatGPT, Perplexity and Google AI Mode all describe this process in their own documentation.

Does ChatGPT search the live web in real time?

Yes, when the question benefits from current information or when you turn search on. OpenAI's help centre says ChatGPT may search automatically, uses partner search providers, and fetches pages with its ChatGPT-User agent for user-initiated actions. Stable questions may still be answered from training data.

Why can a data centre not just fetch the page?

Because websites serve different pages by location and treat traffic from known bot and cloud addresses differently. Google's own documentation says its US-based crawler may not see content a site adapts by locale, and Cloudflare now blocks AI agent bots by default on ad-supported pages for new domains. A data-centre fetch is in the wrong place, and it is identifiable.

Does AI search see the same page I would see?

Only if the fetch comes from a connection like yours, in your location. A fetch from another country or from data-centre address space can return different prices, availability, language or a blocked page, and the model cannot tell the difference.

What happens when a website blocks an AI search fetcher?

The fetch returns an error, a challenge page or a reduced page, and the answer is built without that source. User-triggered fetchers from OpenAI, Perplexity and Google generally ignore robots.txt, so blocking increasingly happens at the network edge, based on where the request comes from.

How does Teneo Beacon fit in?

Teneo Beacon is the opt-in app that shares unused bandwidth with the Teneo network on Windows, macOS, Linux, Android and iOS. Requests for public pages route through real connections in real places, which is what AI search needs to see the page a person there would see. Buyers pay per request in USDC through x402.

What to do next

If you want to supply the connections AI search needs, install Teneo Beacon and follow the setup guide for your platform. If you are building a search or answer product and need to see public pages the way a person in a given city sees them, start at the Teneo Protocol and the Agent Console. For the wider picture of who buys shared bandwidth and why, read what bandwidth sharing is used for.

Key takeaways

  • -Bandwidth sharing
  • -AI search
  • -Teneo Beacon
  • -DePIN