What AI assistants actually send when they fetch your page
- Published: 2026-09-10
- Measured: 2026-09-04 to 2026-09-08
- By: verifyagents.org
- Source data: an instrumented endpoint at https://verifyagents.org/echo, and edge logs from a fleet of 174 Cloudflare zones
- Canonical URL: https://verifyagents.org/study
Two questions sit behind every report of AI traffic. When an AI assistant reads your page, can you tell which one it was? And when a person arrives because an assistant recommended you, can you tell that is what happened?
We measured both. The answers are "sometimes, less often than the documentation suggests" and "almost never."
This document is the measurement. It takes no position on what anyone should do about it. A separate proposal at https://verifyagents.org/proposals/referral-binding sketches one possible fix, and the findings here stand whether or not that proposal is any good.
What is measured and what is not. Every claim below rests on a server-side log: either our own endpoint's record of a request, or Cloudflare edge analytics for zones we operate. Nothing rests on what an assistant said it did, which Section 2 shows is unreliable in both directions. Sample sizes are stated inline and several are small. The fetch tests are one run per assistant, except Perplexity's agent surface, which is six.
1. The three populations
Three populations of AI traffic reach a website, and only the first two are countable today.
- Fetches. Crawler and user-triggered requests from assistant infrastructure, in principle identifiable by published IP ranges and, increasingly, by a Web Bot Auth signature. In practice the identifiable share is smaller than the documentation suggests. Of six assistant surfaces that fetched a URL on request on 2026-09-08, none signed, and three could not be identified by the method their own operator publishes (Section 2).
- Tagged clicks. Human visits where the assistant added a tracking parameter to the link. Only ChatGPT does this at any scale. In a logged-in session on 2026-09-08 its search citations and its inline brand links both carried
utm_source=chatgpt.com, while a URL it had been asked to fetch directly and then echoed back carried nothing. Google stamps a per-query token on some AI Mode advertising clicks and nothing on organic ones. - Dark clicks. Human visits from an assistant where nothing identifies the origin. The assistant's native app strips the
Refererheader, the link carries no parameter, and the visit lands in analytics as direct traffic.
2. What eight assistant surfaces did when asked to fetch one URL
On 2026-09-08 we asked eight assistant surfaces, across seven vendors, to fetch a header mirror we run at https://verifyagents.org/echo and paste back what it returned. The endpoint blocks nobody and records every request, so an assistant's account of what happened can be checked against what arrived.
Six surfaces sent a request. None of the six signed it. Three of the six arrived in a state where a site following the operator's own published guidance could not identify them:
| Assistant | Identifier sent | In a published IP list | Verifiable by the published method |
|---|---|---|---|
| Le Chat (Mistral) | MistralAI-User/1.0 with a docs link | exact match, one of four published addresses | yes |
| ChatGPT | ChatGPT-User/1.0 | operator publishes lists | by user agent and list |
| Claude | Claude-User/1.0 | operator publishes a list | by user agent and list |
| Gemini | Google, a token Google does not document | no, none of Google's five lists | no |
| Perplexity (agent) | none | no, rotating consumer and hosting addresses | no |
| Grok | none | no list is published at all | no |
Two surfaces sent nothing. Perplexity's search mode displayed a completed fetch step for a request that never left its network, and Copilot reported a failure that was real. A ninth, DeepSeek, could not be tested.
Two of the eight described their own retrieval incorrectly, in opposite directions: one showed a fetch that did not happen, and one reported a block that did not happen while being served a normal response. The only dependable account of an AI fetch is the origin's own log, which is what everything in this section rests on.
The rest of this section follows one of those assistants in detail, because repetition shows what a single request cannot. We asked its agent surface the same question six times in half an hour. Every request was unsigned and carried no Perplexity identifier of any kind. Nothing else about them was stable.
| Time (UTC) | Country | Network | Claimed browser / OS | Referer claimed |
|---|---|---|---|---|
17:48:55 | MX | AS8151 Telmex | Chrome 116, Windows | none |
17:50:42 | HN | AS271971 Cable Satelite | Edge 140, Windows | verifyagents.org |
17:54:29 | US | AS16628 DedFiberCo | Vivaldi 6, Linux | www.yahoo.org |
17:58:07 | US | AS62874 Web2Objects | Chrome 140, macOS | verifyagents.org |
18:19:54 | NL | AS132817 DZCRD Networks | Vivaldi 5.1, macOS | verifyagents.org |
18:21:21 | MX | AS13999 Mega Cable | Edge 115, Windows | none |
Six exits, six networks, four countries across two continents, six operating-system-and-browser identities spanning Windows, Linux and macOS, mixed between consumer broadband and hosting. None of the addresses appears in either of the operator's two published IP lists. Five of the six runs used one selected model and the sixth used the vendor's default orchestrator, with no difference in behavior, so what assembles these requests sits below model selection.
Four of the six asserted a Referer that was false: the URL was fetched directly in response to a prompt, with no navigation from any page. An origin counting referral sources by Referer would have credited one of these visits to yahoo.org.
Each of those four also sent sec-fetch-site: none, which contradicts the Referer it sent. Per Fetch metadata, none describes a navigation with no initiator, such as a typed address, and a browser sends no Referer in that case; where a Referer exists the value is same-origin or cross-site. So the headers were assembled rather than emitted by a browser, and an origin can detect that particular inconsistency for free. It identifies no operator, and it stops working the day a client fixes it.
We tested two of that vendor's surfaces and make no claim about the others.
The point of the table is not this operator. It is that every identifier an origin can inspect was either absent or untrue, while the request itself was ordinary. Counting by user agent files these four as desktop humans. Checking published IP ranges rejects them. Reading Referer misattributes them. That is the environment any referral mechanism has to work in, and it is why the token in this document is derived from a signature rather than asserted by the caller.
Cloudflare has stated the third case in its own words. Its published analysis of the crawl-to-refer ratio notes that "traffic referred by Claude's native app does not include a Referer: header," and that because of this "these calculations may overstate the respective ratios, but it is unclear by how much."
3. Two attempts to recover the missing clicks, both failed
The gap is not recoverable from the site side with the data an ordinary analytics setup has. Two independent attempts, run across the same fleet in 2026, both failed:
- Timing. Across 22 sites, 128 UTM-tagged assistant clicks and seven days of edge logs, only 15% of tagged assistant clicks followed a verified assistant fetch of the same site within two minutes, against a 4% base rate for an organic control group. That is a real signal, a likelihood ratio of about 3.6, and it is nowhere near enough to classify an individual session. 26% of tagged clicks had no fetch of that site in the preceding seven days at all. The likeliest explanation is that assistants often answer from an index crawled earlier rather than fetching live, but the study measured the absence of a fetch and not the reason for it.
- Behavior. A behavioral classifier trained on 54 sites and 60 days of sessions reached a cross-site AUC of 0.65 on dirty labels and 0.615 on clean ones. It cannot size the dark share, let alone attribute a visit.
Counting by user agent is worse than useless. Over four days on the same fleet, 54,580 HTML requests carried one of thirteen AI bot user agent strings that Cloudflare's verified bot categories did not confirm, against 23,749 requests in the same window that Cloudflare did verify as AI crawler, AI search or AI assistant traffic. Seven in ten requests bearing an AI user agent could not be verified as coming from the operator they named. On the one day the traffic was taken apart in detail, a single credential scanner accounted for most of it: individual addresses rotating through all thirteen tokens against paths like /.env and /root/.aws/credentials, with the top five addresses alone making up 42%. The other three days were counted rather than dissected, and the scanner's daily volume varies by an order of magnitude. Within that, the verified assistant fetches, the ones where a person is waiting for an answer, ran at 180 to 325 a day.
The conclusion we draw, and the only one these measurements support: the information needed to attribute an assistant referral exists on the assistant's side, at the moment it renders the link, and nowhere else afterwards. Recovering it at the origin is not an engineering problem that better analytics will solve.
4. Limits
The fetch tests are single runs. One request per assistant, on one day, from one prompt, except Perplexity's agent surface at six. An assistant that behaved differently an hour later would not show up here. The Perplexity repetition exists because a single request could not have established that the exits rotate.
Two assistants could not be tested at all. DeepSeek's signup was blocked to us, and Copilot's request never left Microsoft, so it cannot be observed to sign or not sign. Neither is recorded as a negative result.
The fleet is not a random sample of the web. It is 174 Cloudflare zones belonging to one agency's clients, weighted toward small and mid-sized United States service businesses. The traffic mix on a large publisher or an ecommerce site will differ.
The credential-scanner forensics are one day. The four-day totals are counts. Only the first day's traffic was taken apart to identify what the unverified requests were.
"Unverified" is not "spoofed." Cloudflare's verification failing means the request could not be confirmed as coming from the operator it named. Gemini's fetch in Section 2 came from genuine Google infrastructure and would also fail that check, because the user agent is undocumented and the address is on none of Google's published lists.
The timing and behavioral studies used GA4-grade data. A site with server-side session records, or one that could see the full query string at the edge, might do better than we did. We could not test that.
5. Reproducing this
The endpoint at https://verifyagents.org/echo mirrors the headers of any request it receives and blocks nobody. Ask an assistant to fetch it and paste back what it returns, then compare that against what your own logs record. The interesting part is not usually the answer the assistant gives you.
Daily key-directory snapshots for every operator named here are at https://verifyagents.org/observatory, and the per-assistant scoreboard, with the date and source for every cell, is at https://verifyagents.org/transparency.