# What AI assistants actually send when they fetch your page

- **Published:** 2026-09-10
- **Measured:** 2026-09-04 to 2026-09-08
- **By:** verifyagents.org
- **Source data:** an instrumented endpoint at https://verifyagents.org/echo, and edge logs from a
  fleet of 174 Cloudflare zones
- **Canonical URL:** https://verifyagents.org/study

Two questions sit behind every report of AI traffic. When an AI assistant reads your page, can you
tell which one it was? And when a person arrives because an assistant recommended you, can you tell
that is what happened?

We measured both. The answers are "sometimes, less often than the documentation suggests" and
"almost never."

This document is the measurement. It takes no position on what anyone should do about it. A separate
proposal at https://verifyagents.org/proposals/referral-binding sketches one possible fix, and the
findings here stand whether or not that proposal is any good.

**What is measured and what is not.** Every claim below rests on a server-side log: either our own
endpoint's record of a request, or Cloudflare edge analytics for zones we operate. Nothing rests on
what an assistant said it did, which Section 2 shows is unreliable in both directions. Sample sizes
are stated inline and several are small. The fetch tests are one run per assistant, except
Perplexity's agent surface, which is six.

## 1. The three populations

Three populations of AI traffic reach a website, and only the first two are countable today.

1. **Fetches.** Crawler and user-triggered requests from assistant infrastructure, in principle
   identifiable by published IP ranges and, increasingly, by a Web Bot Auth signature. In practice
   the identifiable share is smaller than the documentation suggests. Of six assistant surfaces that
   fetched a URL on request on 2026-09-08, none signed, and three could not be identified by the
   method their own operator publishes (Section 2).
2. **Tagged clicks.** Human visits where the assistant added a tracking parameter to the link.
   Only ChatGPT does this at any scale. In a logged-in session on 2026-09-08 its search citations
   and its inline brand links both carried `utm_source=chatgpt.com`, while a URL it had been asked
   to fetch directly and then echoed back carried nothing. Google stamps a per-query token on some
   AI Mode advertising clicks and nothing on organic ones.
3. **Dark clicks.** Human visits from an assistant where nothing identifies the origin. The
   assistant's native app strips the `Referer` header, the link carries no parameter, and the
   visit lands in analytics as direct traffic.

## 2. What eight assistant surfaces did when asked to fetch one URL

On 2026-09-08 we asked eight assistant surfaces, across seven vendors, to fetch a header mirror we run at
`https://verifyagents.org/echo` and paste back what it returned. The endpoint blocks nobody and
records every request, so an assistant's account of what happened can be checked against what
arrived.

Six surfaces sent a request. **None of the six signed it.** Three of the six arrived in a state
where a site following the operator's own published guidance could not identify them:

| Assistant | Identifier sent | In a published IP list | Verifiable by the published method |
|---|---|---|---|
| Le Chat (Mistral) | `MistralAI-User/1.0` with a docs link | exact match, one of four published addresses | **yes** |
| ChatGPT | `ChatGPT-User/1.0` | operator publishes lists | by user agent and list |
| Claude | `Claude-User/1.0` | operator publishes a list | by user agent and list |
| Gemini | `Google`, a token Google does not document | **no**, none of Google's five lists | **no** |
| Perplexity (agent) | none | **no**, rotating consumer and hosting addresses | **no** |
| Grok | none | **no list is published at all** | **no** |

Two surfaces sent nothing. Perplexity's search mode displayed a completed fetch step for a request
that never left its network, and Copilot reported a failure that was real. A ninth, DeepSeek, could
not be tested.

Two of the eight described their own retrieval incorrectly, in opposite directions: one showed a
fetch that did not happen, and one reported a block that did not happen while being served a normal
response. The only dependable account of an AI fetch is the origin's own log, which is what
everything in this section rests on.

The rest of this section follows one of those assistants in detail, because repetition shows what a
single request cannot. We asked its agent surface the same question six times in half an hour. Every request was unsigned
and carried no Perplexity identifier of any kind. Nothing else about them was stable.

| Time (UTC) | Country | Network | Claimed browser / OS | `Referer` claimed |
|---|---|---|---|---|
| `17:48:55` | MX | AS8151 Telmex | Chrome 116, Windows | none |
| `17:50:42` | HN | AS271971 Cable Satelite | Edge 140, Windows | `verifyagents.org` |
| `17:54:29` | US | AS16628 DedFiberCo | Vivaldi 6, Linux | `www.yahoo.org` |
| `17:58:07` | US | AS62874 Web2Objects | Chrome 140, macOS | `verifyagents.org` |
| `18:19:54` | NL | AS132817 DZCRD Networks | Vivaldi 5.1, macOS | `verifyagents.org` |
| `18:21:21` | MX | AS13999 Mega Cable | Edge 115, Windows | none |

Six exits, six networks, four countries across two continents, six operating-system-and-browser
identities spanning Windows, Linux and macOS, mixed between consumer broadband and hosting. None of
the addresses appears in either of the operator's two published IP lists. Five of the six runs used
one selected model and the sixth used the vendor's default orchestrator, with no difference in
behavior, so what assembles these requests sits below model selection.

Four of the six asserted a `Referer` that was false: the URL was fetched directly in response to a
prompt, with no navigation from any page. An origin counting referral sources by `Referer` would
have credited one of these visits to `yahoo.org`.

Each of those four also sent `sec-fetch-site: none`, which contradicts the `Referer` it sent. Per
Fetch metadata, `none` describes a navigation with no initiator, such as a typed address, and a
browser sends no `Referer` in that case; where a `Referer` exists the value is `same-origin` or
`cross-site`. So the headers were assembled rather than emitted by a browser, and an origin can
detect that particular inconsistency for free. It identifies no operator, and it stops working the
day a client fixes it.

We tested two of that vendor's surfaces and make no claim about the others.

The point of the table is not this operator. It is that every identifier an origin can inspect was
either absent or untrue, while the request itself was ordinary. Counting by user agent files these
four as desktop humans. Checking published IP ranges rejects them. Reading `Referer` misattributes
them. That is the environment any referral mechanism has to work in, and it is why the token in this
document is derived from a signature rather than asserted by the caller.

Cloudflare has stated the third case in its own words. Its published analysis of the crawl-to-refer
ratio notes that "traffic referred by Claude's native app does not include a `Referer:` header," and that
because of this "these calculations may overstate the respective ratios, but it is unclear by how
much."

## 3. Two attempts to recover the missing clicks, both failed

The gap is not recoverable from the site side with the data an ordinary analytics setup has. Two
independent attempts, run across the same fleet in 2026, both failed:

- **Timing.** Across 22 sites, 128 UTM-tagged assistant clicks and seven days of edge logs, only
  15% of tagged assistant clicks followed a verified assistant fetch of the same site within two
  minutes, against a 4% base rate for an organic control group. That is a real signal, a likelihood
  ratio of about 3.6, and it is nowhere near enough to classify an individual session. 26% of tagged
  clicks had no fetch of that site in the preceding seven days at all. The likeliest explanation is
  that assistants often answer from an index crawled earlier rather than fetching live, but the
  study measured the absence of a fetch and not the reason for it.
- **Behavior.** A behavioral classifier trained on 54 sites and 60 days of sessions reached a
  cross-site AUC of 0.65 on dirty labels and 0.615 on clean ones. It cannot size the dark share,
  let alone attribute a visit.

Counting by user agent is worse than useless. Over four days on the same fleet, 54,580 HTML requests
carried one of thirteen AI bot user agent strings that Cloudflare's verified bot categories did not
confirm, against 23,749 requests in the same window that Cloudflare did verify as AI crawler, AI
search or AI assistant traffic. Seven in ten requests bearing an AI user agent could not be verified as coming from the
operator they named. On the one day the traffic was taken apart in detail, a single credential scanner
accounted for most of it: individual addresses rotating through all thirteen tokens against paths
like `/.env` and `/root/.aws/credentials`, with the top five addresses alone making up 42%. The
other three days were counted rather than dissected, and the scanner's daily volume varies by an
order of magnitude. Within that, the verified assistant fetches, the ones where a person is
waiting for an answer, ran at 180 to 325 a day.

The conclusion we draw, and the only one these measurements support: the information needed to
attribute an assistant referral exists on the assistant's side, at the moment it renders the link,
and nowhere else afterwards. Recovering it at the origin is not an engineering problem that better
analytics will solve.

## 4. Limits

**The fetch tests are single runs.** One request per assistant, on one day, from one prompt, except
Perplexity's agent surface at six. An assistant that behaved differently an hour later would not
show up here. The Perplexity repetition exists because a single request could not have established
that the exits rotate.

**Two assistants could not be tested at all.** DeepSeek's signup was blocked to us, and Copilot's
request never left Microsoft, so it cannot be observed to sign or not sign. Neither is recorded as
a negative result.

**The fleet is not a random sample of the web.** It is 174 Cloudflare zones belonging to one
agency's clients, weighted toward small and mid-sized United States service businesses. The traffic
mix on a large publisher or an ecommerce site will differ.

**The credential-scanner forensics are one day.** The four-day totals are counts. Only the first
day's traffic was taken apart to identify what the unverified requests were.

**"Unverified" is not "spoofed."** Cloudflare's verification failing means the request could not be
confirmed as coming from the operator it named. Gemini's fetch in Section 2 came from genuine Google
infrastructure and would also fail that check, because the user agent is undocumented and the
address is on none of Google's published lists.

**The timing and behavioral studies used GA4-grade data.** A site with server-side session records,
or one that could see the full query string at the edge, might do better than we did. We could not
test that.

## 5. Reproducing this

The endpoint at https://verifyagents.org/echo mirrors the headers of any request it receives and
blocks nobody. Ask an assistant to fetch it and paste back what it returns, then compare that
against what your own logs record. The interesting part is not usually the answer the assistant
gives you.

Daily key-directory snapshots for every operator named here are at
https://verifyagents.org/observatory, and the per-assistant scoreboard, with the date and source
for every cell, is at https://verifyagents.org/transparency.
