# Referral Binding for Web Bot Auth

- **Version:** draft-00, 2026-09-08
- **Status:** a sketch, put out for comment. Not adopted by any standards body, not submitted
  anywhere, and not implemented by any assistant. Section 13 is what we are actually asking.
- **Editor:** verifyagents.org
- **Normative anchors:** [RFC 9421](https://www.rfc-editor.org/rfc/rfc9421.html) (HTTP Message Signatures), and [I-D.ietf-webbotauth-httpsig-protocol](https://datatracker.ietf.org/doc/draft-ietf-webbotauth-httpsig-protocol/) (Web Bot Auth), adopted as a working group document on 2026-09-01
- **Canonical URL:** https://verifyagents.org/proposals/referral-binding

## Abstract

An AI assistant fetches a page while composing an answer, cites the page, and a person clicks the
citation. The fetch is the identifiable half, at least in principle: Web Bot Auth gives an operator
a way to sign it with an Ed25519 key so that any origin can confirm who made it. The click that
follows is not identifiable at all. Native assistant apps send no `Referer`, one vendor tags some
links with a UTM parameter and the rest tag nothing, and a site with ordinary analytics cannot tell
an assistant referral from someone typing the address by hand.

This document sketches Referral Binding: a short token, derived from the signed fetch itself, that
the assistant appends to every link it renders from that fetch. It requires no new key, no new
signature, no per-user state, and no cooperation from the receiving site. A site that does nothing
still sees a marker that the visit came from an assistant. A site that keeps its own log of verified
Web Bot Auth requests can recompute the token and join the click to the exact fetch that produced
it.

Two things ought to be said before the mechanism. The measurements behind it, published at
https://verifyagents.org/study, found that of eight assistant surfaces asked to fetch one
instrumented URL, six sent a request and **none of the six signed it**. Signing is therefore not a
precondition this document may assume; it is the thing to argue for first, and the mechanism here is
one answer to what a signature is worth to the operator that produces it.

And a plain UTM parameter would solve much of the same problem more cheaply. Section 2.1 says so
plainly and sets out the narrower case for doing anything more, because that is the objection this
document most deserves and we would rather raise it than have it raised for us.

## 1. Problem

Three populations of AI traffic reach a website, and only the first two are countable today.

1. **Fetches.** Crawler and user-triggered requests from assistant infrastructure, in principle
   identifiable by published IP ranges and, increasingly, by a Web Bot Auth signature.
2. **Tagged clicks.** Human visits where the assistant added a tracking parameter to the link.
3. **Dark clicks.** Human visits from an assistant where nothing identifies the origin. The
   assistant's native app strips the `Referer` header, the link carries no parameter, and the visit
   lands in analytics as direct traffic.

We measured all three, and the measurement is published separately at
https://verifyagents.org/study. The four findings this document rests on:

- **Of eight assistant surfaces asked to fetch one instrumented URL, six sent a request and none
  signed it.** Three of the six could not be identified by the method their own operator publishes.
  Signing the fetch is not a solved precondition this document may assume.
- **Only ChatGPT tags its organic citations at any scale.** Everyone else renders a bare link, so
  the tag is not being stripped in transit, it is never applied.
- **The gap cannot be closed from the site side.** A timing join recovers 15% of known assistant
  clicks at a likelihood ratio of about 3.6, and a behavioral classifier reaches an AUC of 0.65. Both
  attempts ran across 174 zones and both failed.
- **Asserted values are not worth much.** Of six requests from one assistant's agent surface, four
  asserted a `Referer` that was false, one of them naming `yahoo.org`, and every self-declared field
  we could check was either absent or untrue.

The information needed to attribute an assistant referral exists on the assistant's side, at the
moment it renders the link, and nowhere else afterwards.

## 2. What this proposes

At render time the assistant already knows which fetch produced the citation it is about to link.
It appends one query parameter derived from that fetch:

```
https://example.com/pricing?wbab=19be2817c5c441df35d7561bd&utm_source=assistant.example&utm_medium=ai-assistant
```

The `wbab` value is a truncated hash of the signed fetch. The assistant stores nothing. The site
recognizes the parameter's presence as an assistant referral without any setup at all, and a site
that logs its verified Web Bot Auth requests can recompute the same value and learn which fetch,
which operator, and when.

Everything else in this document is the detail needed to make two implementations agree.

### 2.1 Why not simply a UTM parameter

This is the first question the proposal deserves, and the honest answer is that a UTM gets you most
of the way.

The token rides in the query string exactly as `utm_source` does. It survives what a UTM survives
and is lost where a UTM is lost, and it is copied into other people's links the same way. There is
no transport advantage here, and the problem being solved was never a tag going missing in transit.
Only ChatGPT applies a tag at all; every other assistant renders a bare link, so there is nothing to
strip. If all eight surfaces sent `utm_source=<their host>&utm_medium=ai-assistant` tomorrow, most
of what Section 1 describes would be fixed, at a cost to each operator of one string constant.

**We think that should happen, with or without this document.** It is cheaper than anything proposed
here and it needs no signature, no key and no specification.

What the derived token adds, once an operator has decided to touch its rendering path at all, is
three things a declared parameter cannot carry:

- **Which fetch, not which company.** `utm_source=chatgpt.com` names an operator. The token names one
  request: this page, at this second, under this key. The page an assistant read is frequently not
  the page it linked to, and only the token can tell an origin which of its pages earned the
  citation.
- **A value that can be checked rather than believed.** A declared parameter is correct when the
  operator is honest and its client is working, and worthless otherwise, with no way to tell the two
  apart at the origin. The token can be recomputed against a fetch the origin logged itself. Section
  1's fourth finding is the argument: every self-declared field we were able to check on one
  assistant's agent surface was absent or untrue.
- **One thing to look for.** UTM conventions are per-vendor and unregistered. OpenAI already runs at
  least two, tagging Yelp links `utm_source=openai&utm_medium=feed_v2&adjust_creative=openai` while
  its organic citations carry `utm_source=chatgpt.com`.

What it does not add is proof to anyone but the origin itself. Section 8 is explicit about that, and
it is the limitation we would most like to be argued out of.

## 3. Terminology

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT",
"RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as
described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown
here.

Those keywords are used in three places only: on the assistant, which implements this document; on
an origin that chooses to implement the join, for which they define what a conforming implementation
does; and on any future specification that extends this one. No origin is required to implement the join at all, since the parameter
arrives whether or not the site has read this document, and browsers, CDNs and intermediaries implement
nothing here. Advice to a party with no conformance obligation is written as guidance or as a
security consideration, never as a keyword.

- **Assistant.** Software that fetches web content while composing an answer for a person and then
  renders links to that content. It is the signer of the fetch and the minter of the token. The
  Web Bot Auth protocol draft calls this party the **agent**; the terms are used here for the same
  role, and this document prefers "assistant" because the mechanism applies only to the
  answer-composing case and not to crawling.
- **Bound fetch.** An HTTP request carrying an RFC 9421 signature with `tag="web-bot-auth"` that
  also carries `created` and `keyid`. That is a statement about the request's form, not about its
  cryptographic verification: a token derives from a bound fetch whether or not anyone has checked
  the signature. Whether the signature verified is a separate property, which the origin records
  alongside the fetch and which Section 8 says is the one that carries any weight.
- **Referral token.** The value of the `wbab` query parameter, derived from a bound fetch as
  described in Section 4.
- **Bound referral.** A click whose URL carries a referral token.
- **Unbound referral.** A click that carries `utm_source` naming an assistant but no referral
  token. It says where the visit came from; it does not prove it.
- **Origin.** Used in the sense of RFC 9110 Section 4.3.1, the scheme, host and port that served
  the bound fetch, and by extension the party operating it, which is also the party that receives
  the later click. This document does not redefine the term.

## 4. Token derivation

### 4.1 Inputs

The token is derived from six fields. Five are present in every conformant Web Bot Auth request;
the sixth is a constant this document defines. Nothing needs to be generated, stored, or
negotiated.

| Field | Source | Notes |
|---|---|---|
| Context | The fixed string `web-bot-auth-referral-binding/1` | Domain separation. Bumped with the version character. |
| Signature | The raw signature bytes for the chosen label in the `Signature` header, re-encoded as unpadded base64url | Signers differ on padding and on base64 versus base64url. Normalizing here is what stops two correct implementations from disagreeing. |
| Created | The `created` signature parameter | REQUIRED by the Web Bot Auth protocol draft. |
| Key id | The `keyid` signature parameter | REQUIRED by the protocol draft, where it MUST be the JWK SHA-256 thumbprint. Some deployed directories use short labels instead. The derivation copies whatever the signature carries and takes no position. |
| Authority | The RFC 9421 `@authority` of the fetched request | |
| Path | The RFC 9421 `@path` of the fetched request | Never the query string. |

Section 4.2 says exactly how each of those six becomes bytes. The descriptions in this table are
not sufficient on their own: a hash has no notion of two values being equivalent, so every place
two implementers could serialize the same field differently is a place they derive different
tokens.

The signature label is not an input. Google labels its signature `g`, Cloudflare's examples use
`sig1`, and a token that depended on the label would break when a signer renamed it. Which
signature on the request the token is derived *from* is a different question, and Section 4.4
answers it.

The `nonce` is not an input either. RFC 9421 ends every signature base with the
`@signature-params` line, so the signature bytes already commit to the nonce, the expiry and the
tag when those are present. Naming the nonce a second time would add no entropy and one more thing
for two implementations to get wrong. This matters because the nonce cannot be relied on: the
Web Bot Auth working group document states that it "defines no additional nonce requirements," and
Cloudflare's verifier documentation says plainly that "there is currently no `nonce` validation,
nor does Cloudflare guard against replay attacks using a database of seen `nonces`."

Authority and path are named explicitly, even though a signature that covers them already commits
to them, because a signature that covers `@authority` alone verifies against any path on that host
until it expires. One such signature can legitimately accompany several fetches of one site. Naming
the fetched path makes the token per fetch rather than per signature. Assistants that intend to
emit tokens SHOULD cover `@path` in the signature as well, which makes the two agree.

### 4.2 Canonicalization

Two implementations agree only if they serialize the six fields identically, and a hash has no
notion of two values being equivalent.

RFC 9421 Section 2.2.6 scopes its normalization narrowly, and it is worth reading the three
sentences together because the middle one does the work. The `@path` value "is normalized according
to the rules provided in [HTTP], Section 4.2.3"; then, "Namely, an empty path string is normalized
as a single slash (/) character"; then, path components are "represented by their values before
decoding any percent-encoded octets, as described in the simple string comparison rules provided in
Section 6.2.1 of [URI]". The "Namely" limits the RFC 9110 reference to the empty-path rule alone,
which is the same construction Section 2.2.3 uses for `@authority`. So RFC 9421 is not ambiguous
here, and this document simply follows it.

It is worth stating explicitly anyway, because RFC 9110 Section 4.2.3 read whole would make
`/%7Esmith` and `/~smith` the same resource, and they hash differently. An implementer who follows
the first sentence past the second will build something that disagrees with a conformant peer.

The governing rule is that a URL is canonicalized **once**, at the point where it becomes a
request, and never again:

> **Canonicalize once.** The `authority` and `path` fields are the `@authority` and `@path`
> component values of the fetched request, computed as RFC 9421 Sections 2.2.3 and 2.2.6 define
> them, whether or not the signature covers those components. Where the signature does cover them,
> these are byte for byte the values in the signature base. An implementation that starts from a
> URL string rather than from a captured request MUST derive those values the way an HTTP client
> derives a request target from a URL, and MUST NOT apply any further normalization afterwards. An
> origin MUST log the request target as received and MUST NOT place a logging pipeline that
> rewrites URLs between verification and the token index.

The "whether or not the signature covers them" is not a detail. The protocol draft requires a
signer to cover only one of `@authority` or `@target-uri`, and treats `@path` as a narrowing
component a conformant signer may omit, so the signature base of a perfectly conformant fetch often
contains no `@path` line at all. Section 4.1 explains why the derivation names both regardless.

**Authority.** The host, ASCII-lowercased, with the port present if and only if it is not the
default port for the request's scheme. A host that is an internationalized domain name MUST be
represented in its A-label form (RFC 5890 Section 2.3.2.1); implementations MUST NOT hash a
U-label. Stating the A-label rule also disposes of what "lowercase" means for a Unicode string,
which is locale- and code-point-dependent in most runtimes. The `userinfo` subcomponent,
deprecated by RFC 9110 Section 4.2.4, MUST NOT appear.

The scheme is not an input, because `@authority` does not carry it. A fetch of `http://example.com/p`
and one of `https://example.com/p` derive the same token, so an origin serving both MUST index them
together. Note that this makes the token slightly less specific than the RFC 9110 Section 4.3.1
sense of "origin" that Section 3 uses, where the scheme is part of the triple.

**Path.** The absolute path of the request target as sent on the wire, or `/` if that is empty.
Dot segments are removed, because an HTTP client removes them before it sends the request; this is
the one transformation the rule above permits, and it happens where the URL becomes a request, not
afterwards. Beyond that, implementations MUST NOT decode percent-encoded octets, MUST NOT encode
octets that arrived undecoded, MUST NOT change the case of percent-encoding hexadecimal digits,
and MUST NOT add or remove a trailing slash.

Not decoding is a hygiene requirement and not only an interoperability one. The derivation string is
newline-delimited, so a field carrying a line feed would read as two fields. Path is the last field,
so a decoded `%0A` there cannot presently collide with any conformant six-tuple, but that safety
rests on field order rather than on anything structural, and an implementation that decoded what
this section says not to decode should fail loudly rather than depend on it.

**Signature.** The Byte Sequence value of the `Signature` member under the chosen label, encoded
with the base64url alphabet of RFC 4648 Section 5, **without** padding and without line breaks.
Implementations MUST NOT hash the header's own base64 serialization, which uses `+`, `/` and `=`.
This is the single place an implementation is most likely to go wrong.

**Created.** The Integer value of the `created` signature parameter, rendered in base 10 with no
sign, no leading zeros, no decimal point and no exponent. If `created` is absent, or is not an
Integer, no token is derivable. A `created` held as a floating-point number is the trap here:
`str(1788264000.0)` is `1788264000.0` in Python and `1788264000` in JavaScript, which is one fetch
and two tokens.

**Key id.** The String value of the `keyid` signature parameter after structured-field parsing:
without the surrounding double quotes and with backslash escapes resolved. If `keyid` is absent,
no token is derivable.

**Context.** The ASCII string `web-bot-auth-referral-binding/1`, unchanged.

```abnf
; The derivation string, hashed to produce the token.
derivation-string = context LF signature LF created LF keyid LF authority LF path
context           = %s"web-bot-auth-referral-binding/1"  ; case sensitive, RFC 7405
signature         = 1*( ALPHA / DIGIT / "-" / "_" )      ; base64url, unpadded, RFC 4648 S5
created           = "0" / ( %x31-39 *DIGIT )
keyid             = 1*%x20-7E                            ; the parsed String value, RFC 8941 S3.3.3
authority         = host [ ":" port ]                    ; host and port per RFC 3986 S3.2,
                                                         ; host in A-label form
path              = "/" *( pchar / "/" )                 ; pchar per RFC 3986 S3.3; the hex digits
                                                         ; of a pct-encoded octet are case
                                                         ; preserving here, so lowercase a-f is
                                                         ; permitted as well as uppercase

; The token itself. A separate grammar: it is the output, not part of the input above.
token             = version 24HEXDIG-LC
version           = DIGIT / %x61-7A                      ; one character
HEXDIG-LC         = DIGIT / %x61-66                      ; 0-9 a-f
```

No field may contain LF (0x0A). An implementation that finds one MUST NOT derive a token. For
conformant input it cannot: a structured-field String is drawn from `%x20-7E` (RFC 8941
Section 3.3.3), a request target is percent-encoded on the wire, and the remaining fields are
decimal digits or base64url. The rule exists so that an implementation which decoded something
this section says not to decode fails loudly rather than deriving a forgeable token.

#### Canonicalization vectors

Each row below is Vector A's signature, `created` and `keyid` with only the fetched URL varying,
so each is one canonicalization rule stated as a value rather than as a sentence. These are
computed by the implementation that serves the validator and asserted in its test suite.

| Fetched URL | `authority` | `path` | Token | The rule |
|---|---|---|---|---|
| `https://example.com/pricing` | `example.com` | `/pricing` | `19be2817c5c441df35d7561bd` | The baseline: Vector A. |
| `https://example.com/a/./b/../pricing` | `example.com` | `/a/pricing` | `1e10a18b4f827fd2542a91b7a` | Dot segments are removed. |
| `https://example.com/%7Esmith` | `example.com` | `/%7Esmith` | `1f8a61a92b2845edcfc69eeb5` | Percent-encoded octets are never decoded. |
| `https://example.com/~smith` | `example.com` | `/~smith` | `1f8bfe1ca5eee14643d9780a9` | And what arrived undecoded is never encoded. |
| `https://example.com/PRI%2fcing` | `example.com` | `/PRI%2fcing` | `1b44f222159db92028a7b832b` | Percent-encoding hex case is preserved, not normalized. |
| `https://example.com/PRI%2Fcing` | `example.com` | `/PRI%2Fcing` | `140fc2db41a51488eb38bf2e3` | Which makes this a different token from the row above. |
| `https://example.com/pricing/` | `example.com` | `/pricing/` | `1707fc5160f38c7d1282e0f08` | A trailing slash is neither added nor removed. |
| `https://exämple.com/pricing` | `xn--exmple-cua.com` | `/pricing` | `1b3fe4af94008f9bf38be42dd` | An internationalized host is hashed as its A-label. |
| `https://xn--exmple-cua.com/pricing` | `xn--exmple-cua.com` | `/pricing` | `1b3fe4af94008f9bf38be42dd` | So the two spellings of one host agree. |
| `https://EXAMPLE.com:443/pricing` | `example.com` | `/pricing` | `19be2817c5c441df35d7561bd` | ASCII-lowercased, default port dropped. |
| `https://example.com:8443/pricing` | `example.com:8443` | `/pricing` | `1c57f0a705fd4fb326f358afc` | A non-default port stays. |
| `https://example.com` | `example.com` | `/` | `14ff4d760bd28904123aff813` | An empty path is `/`. |
| `http://example.com/pricing` | `example.com` | `/pricing` | `19be2817c5c441df35d7561bd` | The scheme is not an input, so this is the first row's token again. |

### 4.3 Construction

Concatenate the six fields in the order above, separated by a single line feed (0x0A), with no
trailing line feed. Encode as UTF-8. Take SHA-256 of that byte string. Take the first 12 bytes of
the digest, encode them as lowercase hexadecimal, and prefix the version character `1`.

```
derivation-string = context LF signature LF created LF keyid LF authority LF path
token             = "1" || lowercase-hex( SHA-256( UTF-8( derivation-string ) )[0..11] )
```

The result is 25 characters: one version character and 24 hexadecimal characters, 96 bits of
digest. It is case insensitive, contains only RFC 3986 unreserved characters, never needs percent
encoding, and is greppable in a log file.

### 4.4 Which fetch mints the token

A conversation can produce several signed requests before a link is rendered. Three cases decide
which one the token comes from, and getting them wrong makes the join fail on requests that were
otherwise perfectly formed.

**Redirects.** The token is derived from the request whose response supplied the content being
cited, which is the final request in a redirect chain. An assistant MUST NOT derive a token from a
request that returned a 3xx. Where a redirect crosses authorities, say `example.com/x` to
`www.example.com/x`, a token derived from the first request names an authority that never served
the content and an origin that will never look it up. An assistant MUST NOT emit a token whose
`authority` differs from the authority of the link it is rendering.

**Methods.** This version is defined for safe methods only. An assistant MUST NOT emit a token
derived from a request whose method is other than `GET` or `HEAD`. `@method` is an optional
narrowing component in the protocol draft, so a signature covering `@authority` alone signs one base
for every method. Where a signature covers neither `@method` nor a `nonce`, a `GET` and a `HEAD` of
the same resource in the same `created` second derive the same token and the origin cannot tell them
apart. An assistant that needs them distinguished MUST cover `@method` or include a `nonce`, which
is the same remedy, for the same reason, as Section 4.6.

**More than one signature.** The protocol draft permits a request to carry more than one Web Bot
Auth signature, each under its own label. The assistant MUST derive from the signature it
produced. An origin recomputing tokens cannot reliably tell which that was, so it MUST derive one candidate
token per `web-bot-auth` signature on the request and index all of them. An implementation MUST
NOT combine signature parameters from one label with signature bytes from another: if the
`Signature` field carries no member under the label being derived from, no token is derivable for
that label, and pairing across labels mints a token for a request that never existed.

### 4.5 What the token deliberately does not encode

The visible token carries no timestamp, no key id, no operator name, and no structure of any kind.
Every one of those would invite a reader to parse it and would freeze a format that has to change.
The `gclid` parameter is the cautionary example: two decades of production use with no versioning
layer and no defined structure, which is why it cannot safely be changed now. Here the version
character is the entire forward-compatibility story. A verifier that does not recognize the version
character MUST treat the token as opaque and MUST NOT attempt a match.

### 4.6 Scope

One token per distinct signed fetch. Not one per person, not one per conversation, not one per
click.

Ed25519 is deterministic, so two fetches that produce the same signature base also produce the same
signature and therefore the same token. Without a `nonce`, the base is the same whenever the key,
the covered components and the `created` second are the same, so an assistant fetching one URL twice
within a second mints one token for both. An assistant that wants a token per fetch in that case
MUST include a `nonce`, which is the one thing a nonce does for this mechanism that nothing else
does. This is the reverse of the usual advice: Section 4.1 explains why the nonce is not a
derivation input, and this section explains why a signer may still want one.

When a cached answer is served to many people, they all carry the same token, and the origin sees a
count of clicks against one fetch with no way to tell those people apart. That is a real property
where it holds, and Section 7.2 sets out where it does not.

## 5. URL convention

An assistant that mints a token for a link:

- MUST place it in the query string as `wbab=<token>`. Not the fragment, which never reaches the
  server, and not a path segment, which changes the resource.
- SHOULD also append `utm_source=<the assistant's signature-agent host>`, so that analytics which
  know nothing about this document still record the referral. Where the assistant publishes a
  Web Bot Auth directory, this SHOULD be the host that serves it.
- SHOULD also append `utm_medium=ai-assistant`. Google Analytics 4's default AI Assistant channel
  matches on `medium` exactly equal to `ai-assistant`, evaluated after Organic Search and before
  Referral. GA4 also sets that medium itself when the referrer matches its own list of AI
  assistants, so a tagged link and a recognized referrer land in the same channel. Without a medium,
  GA4 reports the campaign and medium as `(not set)`.
- MUST NOT append it to a link the assistant did not produce from a bound fetch.
- MUST preserve any query parameters already on the target URL.
- MUST leave at most one `wbab` on the URL. Where the target URL already carries one, which is
  Section 6's own warned failure mode of a tagged URL copied into a link somewhere else, the
  assistant MUST replace its value rather than append a second.

An origin that receives more than one `wbab` MUST treat the token as absent. `URLSearchParams.get`
returns the first value and PHP's `$_GET` returns the last, so an origin that picks one is picking
at random. An origin MUST ASCII-lowercase the received value before comparing it.

If an assistant serves a link from an index rather than a live fetch, there is no bound fetch and
therefore no token. It SHOULD still send `utm_source`, which makes the referral unbound but visible.
Section 9.1 describes an extension for this case.

### 5.1 Error handling

Four cases. In two of them the answer is to recognize the referral and decline the join; in the
other two there is no referral to recognize.

**A version character this site does not implement.** An origin MUST recognize any value that,
after the ASCII-lowercasing above, matches `version 24HEXDIG-LC` as an assistant referral, whatever
the version character is, and MUST NOT attempt a match for a version it does not implement. This is
the whole of the forward-compatibility mechanism Section 4.5 describes: if a version-1 site treats a
version-2 token as malformed, then the first assistant to ship version 2 disappears from every
version-1 site's reporting, which is a worse outcome than the one the version character exists to
prevent.

**A malformed value.** Anything not matching that shape is not a referral token, and an origin MUST
NOT count it as a referral or attempt a match against it. It is more likely a copied URL fragment or
a scanner than a new version.

**More than one `wbab`.** Treated as absent, per Section 5.

**A well-formed token that matches no logged fetch.** This is the ordinary case, not an error. The
fetch may have been answered from cache, the headers may never have been logged, the fetch may be
older than the join window, or the token may be fabricated. None of those is distinguishable from
the others at the origin, and Section 8 says what an unmatched token is worth.

## 6. Site behavior

This section is guidance. A site implements nothing in this document and receives the parameter
anyway, so a site that ignores all of it still sees `wbab` in its logs and can count assistant
referrals it could not count before. The one keyword below binds only a site that chooses to build
the join, which is the case Section 3 admits. Where a site's choice has consequences for someone
other than the site, the consequence is stated in Section 7 or Section 8 rather than as a keyword
here. What an origin does with a token it cannot use is in Section 5.1.

**Recognize.** The presence of a well-formed `wbab` value is sufficient to classify the visit as an
assistant referral, whatever the `Referer` header says or fails to say.

**Absence means nothing.** A visit without a token is not evidence that it did not come from an
assistant. Section 5 says an index-served citation carries no token at all, and the measurement in
The study found that 26% of tagged clicks had no fetch of that site in the preceding seven days.
A site that reads a missing `wbab` as "not an assistant" will under-count by more than it
previously over-counted. The protocol draft has the sentence to borrow, in its Section 6.11:
absence of a signal is not evidence about the party that did not send it.

**Join, if you want the detail.** To learn which operator, which fetch and when, keep a log of
verified Web Bot Auth requests with the six derivation fields, recompute the token for each, and
look the click's token up. Derive at ingest and index on the token, rather than scanning the
window and deriving per row at read time: at a site taking 10 million signed fetches a day, seven
days is 7×10⁷ rows and the read-time reading is a full scan on every click. Retain, per token, the
`created`, `keyid`, `signature-agent`, authority and path. The signature is the only field the
origin does not already have in some form, and it is 64 bytes, exactly 86 characters as unpadded
base64url. That is the honest storage cost of the join.

At that index size the truncation holds comfortably. The birthday bound on 96 bits of digest at
7×10⁷ entries is roughly n²/2⁹⁷, about 3×10⁻¹⁴, which is why 12 bytes is enough and 16 would be
waste.

An origin implementing the join SHOULD bound it to 7 days from `created`. That is
among the shortest windows in comparable systems (Meta and TikTok default click-through windows, and
Safari's cap on script-set cookies; Google's `_gl` linker parameter is shorter still, at two
minutes) and long enough to cover a conversation someone returns to. A longer window makes the index larger without making the token more meaningful.

**The join is an opt-in logging change, not a matter of keeping a log.** No standard access-log
format captures arbitrary request headers. nginx `combined`, Apache `combined` and Cloudflare's
default Logpush field set carry `Referer` and `User-Agent` and nothing else, so `Signature` and
`Signature-Input` need a custom log format, a header allowlist, or a worker at the edge. Two further cases leave an origin with no row to join
against, and neither is the site's fault: a signed fetch of a cacheable page may be answered at the
CDN edge and never reach the origin at all, and the protocol draft's Section 6.6 says only that a
proxy SHOULD NOT strip these headers.

**Do not vary the response.** The token is an observation, not an input, and a site that varies its
response on it builds a match oracle: anyone can then test a guessed token against the site and
learn whether it corresponds to a real fetch. Section 8 sets out why that matters.

**Keep it out of the cache key, without keeping it out of your logs.** Most CDNs include the full
query string in the cache key by default, which means an unfiltered token fragments the cache one
entry per fetch. The order of operations decides whether the remedy also destroys the data, so
prefer the remedy that does not depend on it.

Stripping the parameter is the fragile choice. On Cloudflare a URL Rewrite Rule using
`remove_query_args()` runs in the `http_request_transform` phase, and Cloudflare's published phase
list puts that phase ahead of both Request Header Transform Rules (`http_request_late_transform`)
and Cache Rules (`http_request_cache_settings`). Everything downstream of the rewrite sees a URL
with no token on it. A deployment that strips the parameter therefore has to establish that its
own reader runs before that phase, and "strip it at the edge, read it somewhere else at the edge"
is not a safe assumption to make without checking.

Excluding it from the cache key instead leaves the parameter on the request all the way to the
origin, so nothing has to run in a particular order. Cache Rules offer two forms. "Ignore query
string" is available below Enterprise and is safe for a site whose pages do not vary by query
string, though it is blunt: it drops every parameter from the key, not just this one.
Per-parameter cache key exclusion is exactly the right instrument and is an Enterprise feature.

**Keep it out of the canonical URL.** Serve a self-referential `<link rel="canonical">` without the
token. Never write the token into internal links, sitemaps, `og:url`, or anything else that gets
crawled or shared. Google's own guidance warns against referral parameters and session identifiers
in indexed URLs, and the recurring failure mode with `gclid` is people copying a tagged URL into a
link on another site.

**Take it out of the address bar.** After reading the parameter, call
`history.replaceState` to remove it, before any analytics tag reads the URL. This keeps the token
out of copied and pasted links, and it is why the parameter must never be needed on a second page
view.

**Do not put it in a cookie or join it to a user.** See Section 7.

## 7. Privacy considerations

### 7.1 What the token reveals

To the origin: that this click came from this fetch. The origin already saw the fetch, and already
saw the click. The token joins two events it was party to, and adds no third fact.

To the assistant: nothing new. The assistant learns nothing by minting the token, because it never
sees whether the token comes back. There is no callback, no pixel, and no reporting endpoint in this
design. An assistant that wants referral numbers has to ask the origin for them, exactly as it does
now.

To anyone else: nothing about the person. Like any query parameter, the token is visible to
third-party scripts and subresources on the landing page until the site removes it, which is why
Section 6 says to take it out of the address bar early.

### 7.2 Cross-site linkability

None from the token values. The token is bound to one `@authority`, and a fetch of a different site
produces a different signature, a different authority and a different token. Two origins comparing
their tokens cannot tell that two of them came from the same conversation, the same person, or the
same session, because the token derives from nothing that describes the person. There is no user
identifier in the derivation because there is no user identifier available to it.

Two residual channels are worth stating plainly, because the argument above is weaker than it
first looks.

The first is metadata neither origin learns from the token itself. Each origin recovers its own
`created` and `keyid` from its own fetch log, so two colluding origins, or one party that fronts
both, can observe that they each saw a bound click from the same key within the same second. The
fetch-to-fetch half of that correlation exists in their raw request logs without this document. The
click-to-fetch half is what this document adds, so on this point the mechanism does make an existing
correlation easier to complete, and a CDN or bot management vendor fronting both origins is the
party best placed to do it.

The second is that the anonymity set is sometimes one. The per-fetch property protects a person only
when a fetch serves more than one of them. When an assistant fetches a page live while composing an
answer for a single person, the set is a single person, and for that fetch the token is a per-person
identifier for as long as the join window lasts. The measurements in the study at https://verifyagents.org/study suggest this is not a
corner case: 15% of tagged assistant clicks followed a fetch of the same site within two minutes,
which is the signature of a live per-query fetch. Implementations that care about this should prefer
serving cached answers, and origins should hold to the join window in Section 6 rather than
retaining tokens indefinitely. It is the weakest part of this proposal and the place where review is
most wanted.

### 7.3 Why a derived token does not leak the signature

The derivation is a one-way function truncated to 96 bits. It reveals no signature bytes, no key
material, and nothing about the private key. An origin that receives a token learns only whether it
matches a fetch it already logged.

### 7.4 The stripping criteria that browsers apply

Browsers that remove tracking parameters from navigation apply a consistent test. Brave strips
parameters "known to be specific to: a user, an email address, or an individual click," while
explicitly not interfering with campaign-level tracking. Mozilla targets identifiers used "for the
purpose of building a user profile." Apple targets "unique identifiers embedded into URLs."

A referral token as specified here is per fetch, not per user and not per click. Many people can
share one token, though Section 7.2 sets out the case where only one person does. It cannot be joined to an account, a session, a conversation, or an advertising
profile, and it is usable only by the site the fetch was made against. Those properties are what
keep it outside the criteria above, and they are properties of this mechanism rather than
observations about it. An implementation of this document:

- MUST NOT mint a distinct token per recipient of a cached answer.
- MUST NOT derive the token from anything describing the person, including a conversation
  identifier, an account identifier, or a session identifier.
- MUST NOT report token-to-outcome data back to the assistant.

An implementation that violates any of these has built a click identifier, and should expect to be
treated as one.

### 7.5 Relationship to the working group's privacy text

The Web Bot Auth protocol draft carries three passages that a fetch-to-click token has to answer
directly.

- Section 7.3 says clients "SHOULD take care to avoid signing information that could be used to
  correlate activity across contexts, especially where sensitive user data is involved." Referral
  Binding correlates two contexts by design, so this is the passage that matters most. The answer
  is that the correlation involves no user data, is scoped to a single authority, rotates per
  signature base and therefore per fetch for a signer that includes a `nonce` (Section 4.6), and is
  computable only by a party that already observed both events.
- Section 7.2 warns that a key tied to a specific human individual "makes the key usable for user
  tracking or profiling." It carries no normative keyword, and the passage is under active
  revision: the MUST that once sat there, in the architecture draft's Section 5.9, came out after
  issues #114 and #115 were opened against it, and both remain open. Referral Binding is compatible
  with the stricter reading as well as the current one. The token derives from a signature over a request, and from nothing the
  assistant knows about the person. Two people who receive the same cached answer receive the same
  token.
- Section 5.6 sets a bound for a session credential an origin issues after verifying a
  signature: "no longer-lived, and no wider in scope than the components the signature covered."
  Session establishment is out of that document's scope and a referral token is not a credential, so
  the rule does not reach this document. It is still the right bound, and this document adopts it
  voluntarily. The token is scoped to the fetched authority and to the fetched path, both of which
  the derivation names whether or not the signature covers them; a signer that covers `@authority`
  alone is conformant and has covered no path to speak of. Its useful life is the join window, which is RECOMMENDED at 7 days. It is an identifier, not
  a credential: possessing it authorizes nothing.

The working group charter puts "Authenticating the end user of a participating client or agent" and
"Tracking or assigning reputation to particular bots" out of scope. Referral Binding does neither. It is defined only for the
identifying path, signatures tagged `web-bot-auth`, and has nothing to say about the anonymous
credential work.

## 8. Security considerations

**A token proves nothing on its own.** Anyone can put any 25 characters in a URL. A token is
meaningful only when it matches a fetch the origin itself logged and verified. An unmatched token
is not evidence of anything, and a site MUST NOT grant access, pricing, or content on the strength
of one.

**Forgery.** To forge a matching token an attacker needs the signature bytes of a real bound fetch
of the origin, plus its `created`, `keyid`, authority and path. Those are visible to the assistant, to the
origin, and to any party that terminates TLS or reads the origin's request logs on its behalf, such
as a CDN or a bot management vendor. They are not visible to a passive network observer, because
the fetch travels over TLS. An attacker who does have them has observed a fetch that already
happened, and can at most claim that a click came from a fetch that genuinely occurred. There is
nothing to steal by doing so.

**Guessing.** 96 bits. An attacker with no knowledge of any fetch cannot produce a matching token.

**Enumeration.** An origin MUST NOT expose a lookup endpoint that answers whether an arbitrary
token matches a fetch. The join belongs in the origin's own analytics, not in a public API. The
validator in Section 14 is not such an endpoint: it derives a token from signature headers the
caller supplies, and it holds no fetch log to look anything up in.

**Replay.** A token can be replayed as many times as someone likes, in as many URLs as they like.
This is not an attack on the origin's accounting so much as noise in it, and it is bounded the same
way every click parameter is: by the join window, and by the fact that a replayed token points at a
fetch that really happened. Origins that care SHOULD count distinct tokens rather than total
tagged clicks.

**Not a secret, not a credential.** The token is neither. It is safe in logs, in referrer headers
and in analytics exports, and it confers nothing on the bearer.

**Why the token is derived rather than asserted.** An assistant could simply state its identity in a
parameter, and that would be far easier to adopt. The study's fetch tests are the argument against
it. In four
observed requests from one assistant, the user agent, the operating system, the network and the
`Referer` were all asserted, and every asserted value we could check was either absent or false. A
self-declared referral marker would be one more field of the same kind: correct when the operator is
honest and its client is working, worthless otherwise, and indistinguishable from either case at the
origin. Deriving the token from the signature means a claim about a fetch can be checked against
that fetch, which is the only property in this design that survives an uncooperative caller.

**Heuristics are not a substitute.** The `Referer` and `sec-fetch-site` contradiction the study records
is a real detector and it costs nothing to run, but it names no operator, it produces no join, and
it disappears the moment a client emits consistent headers. Origins SHOULD treat such checks as
noise filters, never as attribution.

**An origin can mint its own tokens, so this is not settlement data.** Every derivation input is
known to the origin from its own fetch log, which is what lets it recompute the token in the first
place. The same property lets it compute tokens for fetches that no human ever clicked and report
them as referrals. Nothing in this document detects that, and Section 7.1 deliberately gives the
assistant no callback that would. So a token is evidence to the origin about its own traffic, and it
is not evidence to a counterparty. Any arrangement that pays on referral counts needs a mechanism
this document does not provide: a token only the assistant can produce and anyone can verify, which
means a signature over the click rather than a hash of the fetch, at a cost of roughly 86 characters
instead of 25. The version character in Section 4.5 is the hook for specifying that, and Section 13
asks whether it should be.

**Expiry.** The `expires` parameter of the original signature is a bound on the fetch, not on the
token. An origin verifies the signature at fetch time. By the time the click arrives the signature
may be long expired, and that is expected. The join window in Section 6 is what bounds the token.

## 9. Extensions

### 9.1 Citations served from an index

Assistants cite pages they did not fetch during the conversation. In the measurements behind Section 1,
26% of tagged assistant clicks had no fetch of the site in the previous seven days.

There is no bound fetch to derive from in that case, and this version defines no token for it. The
natural extension is to derive the token from the signed crawl request that populated the index,
which would let an origin join a click to a crawl that may be weeks old. That trades a much longer
join window, and a much larger index on the origin's side, for coverage of the common case. It is
left to a future version rather than guessed at here.

### 9.2 Pre-committing the token at fetch time

An origin that wants to know the token before the click can ask for it. The `@query-param` component
lets an assistant include a parameter in the signed request, and the protocol draft's
`Accept-Signature` mechanism lets an origin state which components it wants covered. An assistant
could then fetch with the token already in the URL. This is a heavier arrangement than the
recomputation described in Section 6, it requires the origin to opt in, and it is documented here
only to note that it is possible.

Two things make it awkward beyond its weight. Cloudflare's verifier does not support the component,
spells it `@query-params` in its documentation, and recommends signing `@query` instead, so
anything built on this path should be tested against real verifiers first. And RFC 9421 Section
2.2.8 does not take the parameter's value from the URL as written: it parses the query per the HTML
form-encoding rules into name and value pairs and then re-encodes the value, so what gets signed is
a normalized form of the value rather than the bytes on the wire. For a token drawn only from
lowercase alphanumerics that normalization is the identity, so it costs nothing today; it means the
component's behavior for this parameter is incidental rather than guaranteed, and a longer or
differently shaped future token would not get that for free. The same section RECOMMENDS signing the
whole query string with `@query` where an application commonly covers several parameters.

### 9.3 A signed intent header

The protocol draft observes in Section 5.2.4 that "an agent might include an HTTP header expressing
its intent and sign it." A `Referral-Binding` request header stating that the assistant intends to
emit a token for this fetch would let an origin log selectively rather than logging every verified
request. If defined, it MUST be an RFC 9651 structured field and MUST be covered by the signature.
This document does not define one. The query parameter convention needs no protocol change at all,
and that is its main advantage.

## 10. Relationship to existing work

| Existing | What it does | What it does not do |
|---|---|---|
| Cloudflare Attribution Business Insights (2026-07-01) | Reports a crawl-to-referral ratio per operator, where a referral is a visit "tracked through UTM parameters" | Any per-click join, and nothing at all for operators that do not tag their links. Referral Binding is the assistant-side input this product is missing. |
| Cloudflare's Radar crawl-to-refer analysis (blog, 2025-07-01) | Ratios based on the `Referer` host | Native app traffic, by Cloudflare's own admission |
| `Forwarded: for="openai";use="reference"` (experimental) | An operator-level statement of intent on the inbound fetch, where `use=reference` means "index, excerpt, and link back" | Any per-request identifier, and anything on the click. Referral Binding is how an origin would measure whether the linking back actually happened. |
| `utm_source=chatgpt.com` | The only vendor-documented AI referral parameter, broader in practice than its documentation claims, and unaffected by the native apps: we logged tagged arrivals carrying no `Referer` at all | Any `utm_medium` or `utm_campaign` value, any identification of which fetch produced the link, and any way for the origin to check the claim against its own records. Section 2.1 sets out what this document adds over it, and what it does not |
| `oppref` (ChatGPT Ads) | An opaque per-click token minted server side | Organic coverage, and any way for a site to verify it |
| `adview_query_id` (Google AI Mode ads) | A per-query token on some ad clicks | Documentation, and any confirmed presence on organic links |
| RSL 1.0 | Requires "visible credit and a functional link" under its attribution payment type | Any way to measure whether the link was honored |
| IAB Tech Lab CoMP 1.0 | License tokens for content use | Attribution. The specification does not mention it. |

As of 2026-09-07 we found nothing in the working group's drafts, its 333 archived list messages,
its two GitHub repositories or its four sets of minutes that addresses referrals, click attribution,
or carrying a fetch identifier across to human traffic. The protocol draft does use the word
"attribution" throughout, in a different sense: the lookup that attributes a request to a
`Signature-Agent` URL. The nearest neighbor to this document is an unanswered 2025 request in the
key directory repository for a pseudonymous session identifier for abuse reporting, which the same
token would serve.

## 11. Test vectors

Both vectors are generated by the reference implementation that serves the validator at
https://verifyagents.org/proposals/referral-binding, and are asserted against it in its test
suite: the suite parses the headers, the derivation string, both tokens, the stated key id, the
published JWK and the rendered link out of this section and compares each against what the code
produces, so this section cannot drift away from the implementation without turning the build red.

The key is a throwaway generated for this document and signs nothing real. Both vectors are fixed
in time, so their signatures are expired, which is the normal state of a fetch signature by the time
a click arrives. Expiry has no part in the derivation. The protocol draft recommends an expiry of no
more than 24 hours, so no fixed vector can stay verifiable; a verifier that checks `expires` against
the wall clock will decline to re-confirm these two, and the tokens still derive.

```json
{
  "public": { "kty": "OKP", "crv": "Ed25519", "x": "QJ1Xn3k6gL7dEzCeM95oAweTRmja1VzNHeZfRnQCnAA" },
  "private": { "kty": "OKP", "crv": "Ed25519", "x": "QJ1Xn3k6gL7dEzCeM95oAweTRmja1VzNHeZfRnQCnAA", "d": "-skkN7o8nrSL57bU7u6BWJQGUgeTAZh3rcv2N0O5aXI" }
}
```

Its RFC 7638 thumbprint, and the `keyid` in both vectors, is
`S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM`.

### 11.1 Vector A: no nonce

Request: `GET https://example.com/pricing`, with `Signature-Agent: "https://assistant.example"`.

```
Signature-Agent: sig1="https://assistant.example"
Signature-Input: sig1=("@authority" "@path" "signature-agent";key="sig1");created=1788264000;keyid="S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM";alg="ed25519";expires=1788264300;tag="web-bot-auth"
Signature: sig1=:gJZNIMngohl6J4SCJJPCCjWW6B+HRc2QtZ4u5mD4IAWhis1t9jM+puyFUQPlJBYoz560fqGalKZGmPOJ2Xj1Bw==:
```

Derivation string, with `\n` shown as a line break:

```
web-bot-auth-referral-binding/1
gJZNIMngohl6J4SCJJPCCjWW6B-HRc2QtZ4u5mD4IAWhis1t9jM-puyFUQPlJBYoz560fqGalKZGmPOJ2Xj1Bw
1788264000
S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM
example.com
/pricing
```

Token: `19be2817c5c441df35d7561bd`

Rendered link:

```
https://example.com/pricing?wbab=19be2817c5c441df35d7561bd&utm_source=assistant.example&utm_medium=ai-assistant
```

Note the signature field in the derivation string: the `Signature` header carries standard base64
with padding, and the derivation uses unpadded base64url of the same bytes. This is the one place
an implementation is most likely to go wrong.

### 11.2 Vector B: the same request with a nonce

```
Signature-Agent: sig1="https://assistant.example"
Signature-Input: sig1=("@authority" "@path" "signature-agent";key="sig1");created=1788264000;keyid="S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM";alg="ed25519";expires=1788264300;nonce="AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8gISIjJCUmJygpKissLS4vMDEyMzQ1Njc4OTo7PD0+Pw==";tag="web-bot-auth"
Signature: sig1=:nCj+XQ3P+4E+gEpjgRXYuBA07LE2rE18PjusmQRLUnChPT8IkDPywnqLEEzX9eJHp3C3RBZHV2FobdlFjYgDCw==:
```

Token: `1b0a762ac565e10d47c5fa449`

The token differs from Vector A even though the derivation never names the nonce, because the nonce
is part of the `@signature-params` line that the signature covers. This is the demonstration behind
Section 4.1.

## 12. Adoption notes

What each operator would have to do, given what it publishes today. Directory status checked
2026-09-08 and tracked at https://verifyagents.org/observatory.

**OpenAI.** Furthest along on the click side, and the operator with the clearest missing piece on
the fetch side. It publishes a live key directory at `chatgpt.com` with one Ed25519 key and
`purpose: "ai"`, and it appends `utm_source=chatgpt.com` broadly: in a logged-in session on
2026-09-08, both search citations and inline brand links carried the parameter, with
`rel="noopener"` rather than `rel="noreferrer"`. It also runs at least one negotiated per-partner
scheme, tagging Yelp links with `utm_source=openai&utm_medium=feed_v2&adjust_creative=openai`,
which shows the rendering path already accommodates a counterparty's parameters.

The fetch side is the gap. Asked to fetch a URL that echoes the request headers back, ChatGPT
returned a request carrying `User-Agent: ...ChatGPT-User/1.0...` and no `Signature`,
`Signature-Input` or `Signature-Agent` header at all. The real-time fetcher behind a citation is
unsigned; only agent mode signs. Adoption is therefore two steps, not one: sign `ChatGPT-User`
requests, which five verifiers already validate and which needs no new key, and then derive and
append the token on the rendering path that already appends parameters. Section 9.1 covers the
remaining case where a citation comes from the index with no live fetch behind it.

**Google.** Signs a subset of requests at `agent.bot.goog`, with a self-signed directory, and
already stamps a per-query token (`adview_query_id`) on some AI Mode ad clicks. The fetch side is
compatible today. The click side is the harder ask, since AI Mode traffic is deliberately not
separated from Web search. The narrow version of the pitch is to standardize the token Google
already mints, and to apply it in the Gemini apps, whose iOS surface is reported to identify itself in
the user agent.

**Anthropic.** Publishes IP ranges for its fetchers but no key directory, so it would need to stand
one up first. It has taken no public position on link tagging. A request for UTM tagging was filed
in a Claude Code repository in April 2026 and closed automatically by a bot as off topic for that
repository, with no human reply, so nothing has actually been declined.

**Perplexity.** The furthest from adopting this document as written, and the operator with the most
to gain from solving the problem. Comet Plus pays publishers on a mix of "human visits, search
citations, and agent actions" with an undisclosed attribution method, so a shared, computable
definition of a referral is worth more to Perplexity than to anyone else. Three obstacles sit in
front of it, and two of them are structural rather than a matter of will.

First, it publishes IP ranges but no key directory, so it cannot derive a token from a signature it
does not produce. Second, its citation surface does not fetch: asked three times to retrieve a URL,
against three different targets including one on a zone we control and had built for the purpose,
its search mode reported a failed or completed fetch and sent no request in any of the three, which
we confirmed against our own edge logs. The third of those is the run recorded in the study. A per-fetch token is undefined for an assistant that answers without
fetching, so Section 9.1, the index-derived extension this version defers, is not an optional extra
for Perplexity; it is the only form of this proposal that could apply to its citations. Third, its
agent surface fetches without identifying itself at all (see the study).

The honest pitch is therefore not "adopt this parameter." It is that Perplexity currently cannot
prove a referral to the publishers it pays, that no third party can verify one on its behalf, and
that both facts follow from choices it could reverse: publish a directory, sign the agent fetch, and
work with us on the index-derived variant.

**Microsoft, xAI, Meta, DeepSeek, Mistral.** No key directory as of 2026-09-08. Web Bot Auth
adoption is the prerequisite, and Referral Binding is a reason to do it.

**Verifiers.** As of 2026-09-07, Cloudflare, Vercel, Akamai, AWS WAF on CloudFront and DataDome all
verify Web Bot Auth signatures, and none of them exposes anything that joins a verified fetch to a
later click.
Nothing in this document requires a verifier to change: the derivation runs on inputs an origin can
log for itself, and a site behind any of these verifiers can implement the join alone, subject to
the logging and caching caveats in Section 6. A verifier that exposed the six derivation fields, or
the token itself, to the origin it fronts would save every site behind it that work, and would be
the single cheapest thing any party could do for this proposal.

## 13. What we are asking

*These are real questions, not rhetorical ones. We have spent more time on the measurements than on
the mechanism, and we are a good deal more confident in the first than the second.*

**The three that matter most, before any of the detail below:**

1. **Has this been done?** We searched the working group's drafts, its archived list messages, its
   two repositories and its minutes, and found nothing addressing referrals or carrying a fetch
   identifier across to human traffic. The nearest neighbor was an unanswered request for a
   pseudonymous session identifier for abuse reporting. We would rather be told we missed something
   than duplicate it.
2. **Is the problem worth solving at this layer?** Section 2.1 concedes that a UTM convention gets
   most of the way for a fraction of the effort. If the right answer is "campaign for link tagging
   and stop there," that is a useful answer.
3. **Does the origin-can-mint limit in Section 8 sink it?** Every derivation input is known to the
   origin, so the token is evidence to a site about its own traffic and not to a counterparty. The
   fix is a signature over the click instead of a hash of the fetch, at roughly 86 characters
   instead of 25, and it puts a signing key on the rendering path. We do not know whether that
   trade is worth making, and it is the single most consequential thing we cannot decide alone.

**The narrower ones:**

1. Should the join window be 7 days or 30? Seven matches the shortest comparable windows and keeps
   the origin's index small. Thirty would match ad platform convention and cover conversations
   people return to weeks later.
2. Is a `Referral-Binding` request header worth defining, so origins can log selectively instead of
   logging every verified request? Section 9.3 argues not yet.
3. Should Section 9.1 be folded into this version rather than deferred? Index-served citations are
   the larger share of the problem, and a specification that covers only live fetches solves the
   smaller half.
4. Should this version define the signed, origin-unforgeable token described in Section 8 rather
   than defer it? A hash is short and needs no verification step; a signature is the only thing that
   makes the token evidence to anyone but the origin. The answer probably decides whether this is an
   analytics convention or an accounting one.
5. Is `wbab` the right parameter name, or should it be longer? It is unused in the wild, on no
   published strip list, and names the specification family, but it is opaque to a person reading a
   URL and, as Section 15 sets out, there is no registry for query parameter names and therefore
   nothing that prevents a collision. A longer name buys collision resistance at the cost of URL
   length, and four characters against a 25-character value is not much of a saving.

## 14. Try it

The validator at https://verifyagents.org/proposals/referral-binding takes a landing URL, the
fetch's `Signature-Input` and `Signature` headers, and the URL that was fetched. It derives the
token, reports whether it matches, and runs the full RFC 9421 signature debug on the fetch. Both
vectors above load into it as examples.

The API is one POST:

```
POST https://verifyagents.org/api/referral-binding/verify
Content-Type: application/json

{
  "click": "https://example.com/pricing?wbab=19be2817c5c441df35d7561bd&utm_source=assistant.example&utm_medium=ai-assistant",
  "signatureInput": "sig1=(\"@authority\" \"@path\" \"signature-agent\";key=\"sig1\");created=1788264000;keyid=\"S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM\";alg=\"ed25519\";expires=1788264300;tag=\"web-bot-auth\"",
  "signature": "sig1=:gJZNIMngohl6J4SCJJPCCjWW6B+HRc2QtZ4u5mD4IAWhis1t9jM+puyFUQPlJBYoz560fqGalKZGmPOJ2Xj1Bw==:",
  "targetUrl": "https://example.com/pricing",
  "signatureAgent": "sig1=\"https://assistant.example\""
}
```

## 15. IANA considerations

**A registry for the version character.** Section 4.5 makes one character carry the whole
forward-compatibility story, and a character with no registry behind it is not a mechanism. This
document requests a registry, "Web Bot Auth Referral Binding Versions", with:

- **Value space:** one character, `DIGIT / %x61-7A` (`0` to `9`, `a` to `z`), which is 36 values.
  The token's remaining 24 characters are lowercase hexadecimal, so a version character in `a`
  to `f` is distinguishable from a hex digit only by position, never by inspection. That is
  deliberate: the token is opaque and nothing should try to parse it.
- **Registration policy:** Specification Required.
- **Columns:** Version Character; Description; Reference.
- **Initial contents:** `1`, this document, defining SHA-256 over the six-field derivation string
  truncated to 12 bytes. Also `z`, Reserved, held to introduce a longer version prefix should the
  space approach exhaustion. Under Specification Required nothing holds a value back on its own, so
  the escape hatch has to be reserved now rather than described as a future option.
- **Exhaustion:** thirty-six versions of a mechanism this narrow is not a foreseeable problem, which
  is why one reserved value is the whole of the plan.

**The parameter name has no registry to go in, and that is a real weakness.** There is no IANA
registry for URI query parameter names. `wbab` is claimed by convention and nothing prevents a
collision; the mitigations available are that the name is currently unused in the wild, appears on
no published parameter-stripping list, and names the specification family. Section 13 asks whether
a longer and more collision-resistant name would be a better trade against the cost of a longer
URL. A reviewer should read the absence of a registry as an argument for the longer name, not as
an argument that the question does not matter.

**No other actions.** This document defines no new HTTP header field, no media type, no URI scheme
and no signature parameter, and it registers nothing in the RFC 9421 or Web Bot Auth registries.

## 16. Implementation status

*This section records the state of implementation at the time of writing, in the form RFC 7942 asks
for, and is to be removed before any publication as an RFC. Information current as of 2026-09-08.
RFC 7942 places this section before Security Considerations; it sits here because this document is
read as a web page first, and it should move before Section 8 on conversion to an Internet-Draft.*

One implementation exists, by the editor. It is the code that serves this document. It implements
draft-00, this document, and no other version.

- **Coverage.** Derivation, the canonicalization rules in Section 4.2, click parsing, the join
  check, and the RFC 9421 signature verification the check runs alongside it. It reconstructs the
  signature base itself, including keyed dictionary components, because the library it previously
  delegated to cannot express one.
- **Maturity.** Production, in the sense that it is deployed and serves the validator; it has not
  been used to derive tokens on an assistant's rendering path, because no assistant implements this
  document.
- **Licensing and contact.** The source is not publicly licensed. The implementation is reachable as
  a free, unauthenticated service at https://verifyagents.org/api/referral-binding/verify, and
  contact runs through the canonical URL at the head of this document.
- **What has been interoperability tested, and what has not.** The vectors in Section 11 and the
  canonicalization table in Section 4.2 are generated by this implementation and asserted against
  it. **That is a self-consistency check and not an interoperability report.** No second,
  independent implementation exists, so the interoperability claims in Section 4.2 rest on the
  divergences being demonstrated rather than on two implementations having been made to agree.
  A second implementation is the most useful contribution anyone could make to this document.

## 17. References

**Normative.** RFC 9421 (HTTP Message Signatures); BCP 14 (RFC 2119, RFC 8174); RFC 3986 (URI
Generic Syntax); RFC 4648 (Base64 encodings); RFC 5234 (ABNF) and RFC 7405 (case-sensitive ABNF
strings); RFC 5890 (Internationalized Domain Names); RFC 7638 (JWK Thumbprint); RFC 8941
(Structured Field Values); RFC 9110 (HTTP Semantics); RFC 9651 (Structured Field Values, which
obsoletes RFC 8941 and which Section 9.3 refers to);
`draft-ietf-webbotauth-httpsig-protocol` (Web Bot Auth).

This document cites RFC 8941 for structured fields, following RFC 9421, and RFC 9651 only where
Section 9.3 describes a future header field. The two agree on everything this document relies on.

**Informative.** RFC 7942 (Improving Awareness of Running Code); RFC 8032 (EdDSA); Cloudflare's bot
verification and Attribution Business Insights documentation; Google Analytics 4 channel-grouping
documentation; Brave's and Mozilla's query-parameter stripping policies; Apple's Web AdAttributionKit
documentation; RSL 1.0; IAB Tech Lab CoMP 1.0. Measurements cited in Section 1 are
the editor's own and are described in Section 18.

## 18. Acknowledgments and provenance

The measurements summarized in Section 1, and set out in full at https://verifyagents.org/study,
come from a fleet of 174 Cloudflare zones instrumented at the edge
through 2026, reported at https://completeseo.com. The Web Bot Auth facts are drawn from
RFC 9421, the working group's protocol draft, Cloudflare's bot verification documentation, and live
checks of operator key directories run on 2026-09-08 and repeated daily at
https://verifyagents.org/observatory. Claims in this document that rest on third party reporting
rather than a primary source, in particular the details of `oppref` and `adview_query_id`, are
identified as such in the research notes published alongside it.

This is a proposal from an independent party. It has not been submitted to the IETF, and no
operator named in it has endorsed it.
