Referral Binding for Web Bot Auth

draft-002026-09-08not adopted by any standards bodymarkdown source

Abstract

An AI assistant fetches a page while composing an answer, cites the page, and a person clicks the citation. The fetch is the identifiable half, at least in principle: Web Bot Auth gives an operator a way to sign it with an Ed25519 key so that any origin can confirm who made it. The click that follows is not identifiable at all. Native assistant apps send no Referer, one vendor tags some links with a UTM parameter and the rest tag nothing, and a site with ordinary analytics cannot tell an assistant referral from someone typing the address by hand.

This document sketches Referral Binding: a short token, derived from the signed fetch itself, that the assistant appends to every link it renders from that fetch. It requires no new key, no new signature, no per-user state, and no cooperation from the receiving site. A site that does nothing still sees a marker that the visit came from an assistant. A site that keeps its own log of verified Web Bot Auth requests can recompute the token and join the click to the exact fetch that produced it.

Two things ought to be said before the mechanism. The measurements behind it, published at https://verifyagents.org/study, found that of eight assistant surfaces asked to fetch one instrumented URL, six sent a request and none of the six signed it. Signing is therefore not a precondition this document may assume; it is the thing to argue for first, and the mechanism here is one answer to what a signature is worth to the operator that produces it.

And a plain UTM parameter would solve much of the same problem more cheaply. Section 2.1 says so plainly and sets out the narrower case for doing anything more, because that is the objection this document most deserves and we would rather raise it than have it raised for us.

1. Problem

Three populations of AI traffic reach a website, and only the first two are countable today.

  1. Fetches. Crawler and user-triggered requests from assistant infrastructure, in principle identifiable by published IP ranges and, increasingly, by a Web Bot Auth signature.
  2. Tagged clicks. Human visits where the assistant added a tracking parameter to the link.
  3. Dark clicks. Human visits from an assistant where nothing identifies the origin. The assistant's native app strips the Referer header, the link carries no parameter, and the visit lands in analytics as direct traffic.

We measured all three, and the measurement is published separately at https://verifyagents.org/study. The four findings this document rests on:

The information needed to attribute an assistant referral exists on the assistant's side, at the moment it renders the link, and nowhere else afterwards.

2. What this proposes

At render time the assistant already knows which fetch produced the citation it is about to link. It appends one query parameter derived from that fetch:

https://example.com/pricing?wbab=19be2817c5c441df35d7561bd&utm_source=assistant.example&utm_medium=ai-assistant

The wbab value is a truncated hash of the signed fetch. The assistant stores nothing. The site recognizes the parameter's presence as an assistant referral without any setup at all, and a site that logs its verified Web Bot Auth requests can recompute the same value and learn which fetch, which operator, and when.

Everything else in this document is the detail needed to make two implementations agree.

2.1 Why not simply a UTM parameter

This is the first question the proposal deserves, and the honest answer is that a UTM gets you most of the way.

The token rides in the query string exactly as utm_source does. It survives what a UTM survives and is lost where a UTM is lost, and it is copied into other people's links the same way. There is no transport advantage here, and the problem being solved was never a tag going missing in transit. Only ChatGPT applies a tag at all; every other assistant renders a bare link, so there is nothing to strip. If all eight surfaces sent utm_source=<their host>&utm_medium=ai-assistant tomorrow, most of what Section 1 describes would be fixed, at a cost to each operator of one string constant.

We think that should happen, with or without this document. It is cheaper than anything proposed here and it needs no signature, no key and no specification.

What the derived token adds, once an operator has decided to touch its rendering path at all, is three things a declared parameter cannot carry:

What it does not add is proof to anyone but the origin itself. Section 8 is explicit about that, and it is the limitation we would most like to be argued out of.

3. Terminology

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.

Those keywords are used in three places only: on the assistant, which implements this document; on an origin that chooses to implement the join, for which they define what a conforming implementation does; and on any future specification that extends this one. No origin is required to implement the join at all, since the parameter arrives whether or not the site has read this document, and browsers, CDNs and intermediaries implement nothing here. Advice to a party with no conformance obligation is written as guidance or as a security consideration, never as a keyword.

4. Token derivation

4.1 Inputs

The token is derived from six fields. Five are present in every conformant Web Bot Auth request; the sixth is a constant this document defines. Nothing needs to be generated, stored, or negotiated.

FieldSourceNotes
ContextThe fixed string web-bot-auth-referral-binding/1Domain separation. Bumped with the version character.
SignatureThe raw signature bytes for the chosen label in the Signature header, re-encoded as unpadded base64urlSigners differ on padding and on base64 versus base64url. Normalizing here is what stops two correct implementations from disagreeing.
CreatedThe created signature parameterREQUIRED by the Web Bot Auth protocol draft.
Key idThe keyid signature parameterREQUIRED by the protocol draft, where it MUST be the JWK SHA-256 thumbprint. Some deployed directories use short labels instead. The derivation copies whatever the signature carries and takes no position.
AuthorityThe RFC 9421 @authority of the fetched request
PathThe RFC 9421 @path of the fetched requestNever the query string.

Section 4.2 says exactly how each of those six becomes bytes. The descriptions in this table are not sufficient on their own: a hash has no notion of two values being equivalent, so every place two implementers could serialize the same field differently is a place they derive different tokens.

The signature label is not an input. Google labels its signature g, Cloudflare's examples use sig1, and a token that depended on the label would break when a signer renamed it. Which signature on the request the token is derived from is a different question, and Section 4.4 answers it.

The nonce is not an input either. RFC 9421 ends every signature base with the @signature-params line, so the signature bytes already commit to the nonce, the expiry and the tag when those are present. Naming the nonce a second time would add no entropy and one more thing for two implementations to get wrong. This matters because the nonce cannot be relied on: the Web Bot Auth working group document states that it "defines no additional nonce requirements," and Cloudflare's verifier documentation says plainly that "there is currently no nonce validation, nor does Cloudflare guard against replay attacks using a database of seen nonces."

Authority and path are named explicitly, even though a signature that covers them already commits to them, because a signature that covers @authority alone verifies against any path on that host until it expires. One such signature can legitimately accompany several fetches of one site. Naming the fetched path makes the token per fetch rather than per signature. Assistants that intend to emit tokens SHOULD cover @path in the signature as well, which makes the two agree.

4.2 Canonicalization

Two implementations agree only if they serialize the six fields identically, and a hash has no notion of two values being equivalent.

RFC 9421 Section 2.2.6 scopes its normalization narrowly, and it is worth reading the three sentences together because the middle one does the work. The @path value "is normalized according to the rules provided in [HTTP], Section 4.2.3"; then, "Namely, an empty path string is normalized as a single slash (/) character"; then, path components are "represented by their values before decoding any percent-encoded octets, as described in the simple string comparison rules provided in Section 6.2.1 of [URI]". The "Namely" limits the RFC 9110 reference to the empty-path rule alone, which is the same construction Section 2.2.3 uses for @authority. So RFC 9421 is not ambiguous here, and this document simply follows it.

It is worth stating explicitly anyway, because RFC 9110 Section 4.2.3 read whole would make /%7Esmith and /~smith the same resource, and they hash differently. An implementer who follows the first sentence past the second will build something that disagrees with a conformant peer.

The governing rule is that a URL is canonicalized once, at the point where it becomes a request, and never again:

> Canonicalize once. The authority and path fields are the @authority and @path > component values of the fetched request, computed as RFC 9421 Sections 2.2.3 and 2.2.6 define > them, whether or not the signature covers those components. Where the signature does cover them, > these are byte for byte the values in the signature base. An implementation that starts from a > URL string rather than from a captured request MUST derive those values the way an HTTP client > derives a request target from a URL, and MUST NOT apply any further normalization afterwards. An > origin MUST log the request target as received and MUST NOT place a logging pipeline that > rewrites URLs between verification and the token index.

The "whether or not the signature covers them" is not a detail. The protocol draft requires a signer to cover only one of @authority or @target-uri, and treats @path as a narrowing component a conformant signer may omit, so the signature base of a perfectly conformant fetch often contains no @path line at all. Section 4.1 explains why the derivation names both regardless.

Authority. The host, ASCII-lowercased, with the port present if and only if it is not the default port for the request's scheme. A host that is an internationalized domain name MUST be represented in its A-label form (RFC 5890 Section 2.3.2.1); implementations MUST NOT hash a U-label. Stating the A-label rule also disposes of what "lowercase" means for a Unicode string, which is locale- and code-point-dependent in most runtimes. The userinfo subcomponent, deprecated by RFC 9110 Section 4.2.4, MUST NOT appear.

The scheme is not an input, because @authority does not carry it. A fetch of http://example.com/p and one of https://example.com/p derive the same token, so an origin serving both MUST index them together. Note that this makes the token slightly less specific than the RFC 9110 Section 4.3.1 sense of "origin" that Section 3 uses, where the scheme is part of the triple.

Path. The absolute path of the request target as sent on the wire, or / if that is empty. Dot segments are removed, because an HTTP client removes them before it sends the request; this is the one transformation the rule above permits, and it happens where the URL becomes a request, not afterwards. Beyond that, implementations MUST NOT decode percent-encoded octets, MUST NOT encode octets that arrived undecoded, MUST NOT change the case of percent-encoding hexadecimal digits, and MUST NOT add or remove a trailing slash.

Not decoding is a hygiene requirement and not only an interoperability one. The derivation string is newline-delimited, so a field carrying a line feed would read as two fields. Path is the last field, so a decoded %0A there cannot presently collide with any conformant six-tuple, but that safety rests on field order rather than on anything structural, and an implementation that decoded what this section says not to decode should fail loudly rather than depend on it.

Signature. The Byte Sequence value of the Signature member under the chosen label, encoded with the base64url alphabet of RFC 4648 Section 5, without padding and without line breaks. Implementations MUST NOT hash the header's own base64 serialization, which uses +, / and =. This is the single place an implementation is most likely to go wrong.

Created. The Integer value of the created signature parameter, rendered in base 10 with no sign, no leading zeros, no decimal point and no exponent. If created is absent, or is not an Integer, no token is derivable. A created held as a floating-point number is the trap here: str(1788264000.0) is 1788264000.0 in Python and 1788264000 in JavaScript, which is one fetch and two tokens.

Key id. The String value of the keyid signature parameter after structured-field parsing: without the surrounding double quotes and with backslash escapes resolved. If keyid is absent, no token is derivable.

Context. The ASCII string web-bot-auth-referral-binding/1, unchanged.

; The derivation string, hashed to produce the token.
derivation-string = context LF signature LF created LF keyid LF authority LF path
context           = %s"web-bot-auth-referral-binding/1"  ; case sensitive, RFC 7405
signature         = 1*( ALPHA / DIGIT / "-" / "_" )      ; base64url, unpadded, RFC 4648 S5
created           = "0" / ( %x31-39 *DIGIT )
keyid             = 1*%x20-7E                            ; the parsed String value, RFC 8941 S3.3.3
authority         = host [ ":" port ]                    ; host and port per RFC 3986 S3.2,
                                                         ; host in A-label form
path              = "/" *( pchar / "/" )                 ; pchar per RFC 3986 S3.3; the hex digits
                                                         ; of a pct-encoded octet are case
                                                         ; preserving here, so lowercase a-f is
                                                         ; permitted as well as uppercase

; The token itself. A separate grammar: it is the output, not part of the input above.
token             = version 24HEXDIG-LC
version           = DIGIT / %x61-7A                      ; one character
HEXDIG-LC         = DIGIT / %x61-66                      ; 0-9 a-f

No field may contain LF (0x0A). An implementation that finds one MUST NOT derive a token. For conformant input it cannot: a structured-field String is drawn from %x20-7E (RFC 8941 Section 3.3.3), a request target is percent-encoded on the wire, and the remaining fields are decimal digits or base64url. The rule exists so that an implementation which decoded something this section says not to decode fails loudly rather than deriving a forgeable token.

Canonicalization vectors

Each row below is Vector A's signature, created and keyid with only the fetched URL varying, so each is one canonicalization rule stated as a value rather than as a sentence. These are computed by the implementation that serves the validator and asserted in its test suite.

Fetched URLauthoritypathTokenThe rule
https://example.com/pricingexample.com/pricing19be2817c5c441df35d7561bdThe baseline: Vector A.
https://example.com/a/./b/../pricingexample.com/a/pricing1e10a18b4f827fd2542a91b7aDot segments are removed.
https://example.com/%7Esmithexample.com/%7Esmith1f8a61a92b2845edcfc69eeb5Percent-encoded octets are never decoded.
https://example.com/~smithexample.com/~smith1f8bfe1ca5eee14643d9780a9And what arrived undecoded is never encoded.
https://example.com/PRI%2fcingexample.com/PRI%2fcing1b44f222159db92028a7b832bPercent-encoding hex case is preserved, not normalized.
https://example.com/PRI%2Fcingexample.com/PRI%2Fcing140fc2db41a51488eb38bf2e3Which makes this a different token from the row above.
https://example.com/pricing/example.com/pricing/1707fc5160f38c7d1282e0f08A trailing slash is neither added nor removed.
https://exämple.com/pricingxn--exmple-cua.com/pricing1b3fe4af94008f9bf38be42ddAn internationalized host is hashed as its A-label.
https://xn--exmple-cua.com/pricingxn--exmple-cua.com/pricing1b3fe4af94008f9bf38be42ddSo the two spellings of one host agree.
https://EXAMPLE.com:443/pricingexample.com/pricing19be2817c5c441df35d7561bdASCII-lowercased, default port dropped.
https://example.com:8443/pricingexample.com:8443/pricing1c57f0a705fd4fb326f358afcA non-default port stays.
https://example.comexample.com/14ff4d760bd28904123aff813An empty path is /.
http://example.com/pricingexample.com/pricing19be2817c5c441df35d7561bdThe scheme is not an input, so this is the first row's token again.

4.3 Construction

Concatenate the six fields in the order above, separated by a single line feed (0x0A), with no trailing line feed. Encode as UTF-8. Take SHA-256 of that byte string. Take the first 12 bytes of the digest, encode them as lowercase hexadecimal, and prefix the version character 1.

derivation-string = context LF signature LF created LF keyid LF authority LF path
token             = "1" || lowercase-hex( SHA-256( UTF-8( derivation-string ) )[0..11] )

The result is 25 characters: one version character and 24 hexadecimal characters, 96 bits of digest. It is case insensitive, contains only RFC 3986 unreserved characters, never needs percent encoding, and is greppable in a log file.

4.4 Which fetch mints the token

A conversation can produce several signed requests before a link is rendered. Three cases decide which one the token comes from, and getting them wrong makes the join fail on requests that were otherwise perfectly formed.

Redirects. The token is derived from the request whose response supplied the content being cited, which is the final request in a redirect chain. An assistant MUST NOT derive a token from a request that returned a 3xx. Where a redirect crosses authorities, say example.com/x to www.example.com/x, a token derived from the first request names an authority that never served the content and an origin that will never look it up. An assistant MUST NOT emit a token whose authority differs from the authority of the link it is rendering.

Methods. This version is defined for safe methods only. An assistant MUST NOT emit a token derived from a request whose method is other than GET or HEAD. @method is an optional narrowing component in the protocol draft, so a signature covering @authority alone signs one base for every method. Where a signature covers neither @method nor a nonce, a GET and a HEAD of the same resource in the same created second derive the same token and the origin cannot tell them apart. An assistant that needs them distinguished MUST cover @method or include a nonce, which is the same remedy, for the same reason, as Section 4.6.

More than one signature. The protocol draft permits a request to carry more than one Web Bot Auth signature, each under its own label. The assistant MUST derive from the signature it produced. An origin recomputing tokens cannot reliably tell which that was, so it MUST derive one candidate token per web-bot-auth signature on the request and index all of them. An implementation MUST NOT combine signature parameters from one label with signature bytes from another: if the Signature field carries no member under the label being derived from, no token is derivable for that label, and pairing across labels mints a token for a request that never existed.

4.5 What the token deliberately does not encode

The visible token carries no timestamp, no key id, no operator name, and no structure of any kind. Every one of those would invite a reader to parse it and would freeze a format that has to change. The gclid parameter is the cautionary example: two decades of production use with no versioning layer and no defined structure, which is why it cannot safely be changed now. Here the version character is the entire forward-compatibility story. A verifier that does not recognize the version character MUST treat the token as opaque and MUST NOT attempt a match.

4.6 Scope

One token per distinct signed fetch. Not one per person, not one per conversation, not one per click.

Ed25519 is deterministic, so two fetches that produce the same signature base also produce the same signature and therefore the same token. Without a nonce, the base is the same whenever the key, the covered components and the created second are the same, so an assistant fetching one URL twice within a second mints one token for both. An assistant that wants a token per fetch in that case MUST include a nonce, which is the one thing a nonce does for this mechanism that nothing else does. This is the reverse of the usual advice: Section 4.1 explains why the nonce is not a derivation input, and this section explains why a signer may still want one.

When a cached answer is served to many people, they all carry the same token, and the origin sees a count of clicks against one fetch with no way to tell those people apart. That is a real property where it holds, and Section 7.2 sets out where it does not.

5. URL convention

An assistant that mints a token for a link:

An origin that receives more than one wbab MUST treat the token as absent. URLSearchParams.get returns the first value and PHP's $_GET returns the last, so an origin that picks one is picking at random. An origin MUST ASCII-lowercase the received value before comparing it.

If an assistant serves a link from an index rather than a live fetch, there is no bound fetch and therefore no token. It SHOULD still send utm_source, which makes the referral unbound but visible. Section 9.1 describes an extension for this case.

5.1 Error handling

Four cases. In two of them the answer is to recognize the referral and decline the join; in the other two there is no referral to recognize.

A version character this site does not implement. An origin MUST recognize any value that, after the ASCII-lowercasing above, matches version 24HEXDIG-LC as an assistant referral, whatever the version character is, and MUST NOT attempt a match for a version it does not implement. This is the whole of the forward-compatibility mechanism Section 4.5 describes: if a version-1 site treats a version-2 token as malformed, then the first assistant to ship version 2 disappears from every version-1 site's reporting, which is a worse outcome than the one the version character exists to prevent.

A malformed value. Anything not matching that shape is not a referral token, and an origin MUST NOT count it as a referral or attempt a match against it. It is more likely a copied URL fragment or a scanner than a new version.

More than one wbab. Treated as absent, per Section 5.

A well-formed token that matches no logged fetch. This is the ordinary case, not an error. The fetch may have been answered from cache, the headers may never have been logged, the fetch may be older than the join window, or the token may be fabricated. None of those is distinguishable from the others at the origin, and Section 8 says what an unmatched token is worth.

6. Site behavior

This section is guidance. A site implements nothing in this document and receives the parameter anyway, so a site that ignores all of it still sees wbab in its logs and can count assistant referrals it could not count before. The one keyword below binds only a site that chooses to build the join, which is the case Section 3 admits. Where a site's choice has consequences for someone other than the site, the consequence is stated in Section 7 or Section 8 rather than as a keyword here. What an origin does with a token it cannot use is in Section 5.1.

Recognize. The presence of a well-formed wbab value is sufficient to classify the visit as an assistant referral, whatever the Referer header says or fails to say.

Absence means nothing. A visit without a token is not evidence that it did not come from an assistant. Section 5 says an index-served citation carries no token at all, and the measurement in The study found that 26% of tagged clicks had no fetch of that site in the preceding seven days. A site that reads a missing wbab as "not an assistant" will under-count by more than it previously over-counted. The protocol draft has the sentence to borrow, in its Section 6.11: absence of a signal is not evidence about the party that did not send it.

Join, if you want the detail. To learn which operator, which fetch and when, keep a log of verified Web Bot Auth requests with the six derivation fields, recompute the token for each, and look the click's token up. Derive at ingest and index on the token, rather than scanning the window and deriving per row at read time: at a site taking 10 million signed fetches a day, seven days is 7×10⁷ rows and the read-time reading is a full scan on every click. Retain, per token, the created, keyid, signature-agent, authority and path. The signature is the only field the origin does not already have in some form, and it is 64 bytes, exactly 86 characters as unpadded base64url. That is the honest storage cost of the join.

At that index size the truncation holds comfortably. The birthday bound on 96 bits of digest at 7×10⁷ entries is roughly n²/2⁹⁷, about 3×10⁻¹⁴, which is why 12 bytes is enough and 16 would be waste.

An origin implementing the join SHOULD bound it to 7 days from created. That is among the shortest windows in comparable systems (Meta and TikTok default click-through windows, and Safari's cap on script-set cookies; Google's _gl linker parameter is shorter still, at two minutes) and long enough to cover a conversation someone returns to. A longer window makes the index larger without making the token more meaningful.

The join is an opt-in logging change, not a matter of keeping a log. No standard access-log format captures arbitrary request headers. nginx combined, Apache combined and Cloudflare's default Logpush field set carry Referer and User-Agent and nothing else, so Signature and Signature-Input need a custom log format, a header allowlist, or a worker at the edge. Two further cases leave an origin with no row to join against, and neither is the site's fault: a signed fetch of a cacheable page may be answered at the CDN edge and never reach the origin at all, and the protocol draft's Section 6.6 says only that a proxy SHOULD NOT strip these headers.

Do not vary the response. The token is an observation, not an input, and a site that varies its response on it builds a match oracle: anyone can then test a guessed token against the site and learn whether it corresponds to a real fetch. Section 8 sets out why that matters.

Keep it out of the cache key, without keeping it out of your logs. Most CDNs include the full query string in the cache key by default, which means an unfiltered token fragments the cache one entry per fetch. The order of operations decides whether the remedy also destroys the data, so prefer the remedy that does not depend on it.

Stripping the parameter is the fragile choice. On Cloudflare a URL Rewrite Rule using remove_query_args() runs in the http_request_transform phase, and Cloudflare's published phase list puts that phase ahead of both Request Header Transform Rules (http_request_late_transform) and Cache Rules (http_request_cache_settings). Everything downstream of the rewrite sees a URL with no token on it. A deployment that strips the parameter therefore has to establish that its own reader runs before that phase, and "strip it at the edge, read it somewhere else at the edge" is not a safe assumption to make without checking.

Excluding it from the cache key instead leaves the parameter on the request all the way to the origin, so nothing has to run in a particular order. Cache Rules offer two forms. "Ignore query string" is available below Enterprise and is safe for a site whose pages do not vary by query string, though it is blunt: it drops every parameter from the key, not just this one. Per-parameter cache key exclusion is exactly the right instrument and is an Enterprise feature.

Keep it out of the canonical URL. Serve a self-referential <link rel="canonical"> without the token. Never write the token into internal links, sitemaps, og:url, or anything else that gets crawled or shared. Google's own guidance warns against referral parameters and session identifiers in indexed URLs, and the recurring failure mode with gclid is people copying a tagged URL into a link on another site.

Take it out of the address bar. After reading the parameter, call history.replaceState to remove it, before any analytics tag reads the URL. This keeps the token out of copied and pasted links, and it is why the parameter must never be needed on a second page view.

Do not put it in a cookie or join it to a user. See Section 7.

7. Privacy considerations

7.1 What the token reveals

To the origin: that this click came from this fetch. The origin already saw the fetch, and already saw the click. The token joins two events it was party to, and adds no third fact.

To the assistant: nothing new. The assistant learns nothing by minting the token, because it never sees whether the token comes back. There is no callback, no pixel, and no reporting endpoint in this design. An assistant that wants referral numbers has to ask the origin for them, exactly as it does now.

To anyone else: nothing about the person. Like any query parameter, the token is visible to third-party scripts and subresources on the landing page until the site removes it, which is why Section 6 says to take it out of the address bar early.

7.2 Cross-site linkability

None from the token values. The token is bound to one @authority, and a fetch of a different site produces a different signature, a different authority and a different token. Two origins comparing their tokens cannot tell that two of them came from the same conversation, the same person, or the same session, because the token derives from nothing that describes the person. There is no user identifier in the derivation because there is no user identifier available to it.

Two residual channels are worth stating plainly, because the argument above is weaker than it first looks.

The first is metadata neither origin learns from the token itself. Each origin recovers its own created and keyid from its own fetch log, so two colluding origins, or one party that fronts both, can observe that they each saw a bound click from the same key within the same second. The fetch-to-fetch half of that correlation exists in their raw request logs without this document. The click-to-fetch half is what this document adds, so on this point the mechanism does make an existing correlation easier to complete, and a CDN or bot management vendor fronting both origins is the party best placed to do it.

The second is that the anonymity set is sometimes one. The per-fetch property protects a person only when a fetch serves more than one of them. When an assistant fetches a page live while composing an answer for a single person, the set is a single person, and for that fetch the token is a per-person identifier for as long as the join window lasts. The measurements in the study at https://verifyagents.org/study suggest this is not a corner case: 15% of tagged assistant clicks followed a fetch of the same site within two minutes, which is the signature of a live per-query fetch. Implementations that care about this should prefer serving cached answers, and origins should hold to the join window in Section 6 rather than retaining tokens indefinitely. It is the weakest part of this proposal and the place where review is most wanted.

7.3 Why a derived token does not leak the signature

The derivation is a one-way function truncated to 96 bits. It reveals no signature bytes, no key material, and nothing about the private key. An origin that receives a token learns only whether it matches a fetch it already logged.

7.4 The stripping criteria that browsers apply

Browsers that remove tracking parameters from navigation apply a consistent test. Brave strips parameters "known to be specific to: a user, an email address, or an individual click," while explicitly not interfering with campaign-level tracking. Mozilla targets identifiers used "for the purpose of building a user profile." Apple targets "unique identifiers embedded into URLs."

A referral token as specified here is per fetch, not per user and not per click. Many people can share one token, though Section 7.2 sets out the case where only one person does. It cannot be joined to an account, a session, a conversation, or an advertising profile, and it is usable only by the site the fetch was made against. Those properties are what keep it outside the criteria above, and they are properties of this mechanism rather than observations about it. An implementation of this document:

An implementation that violates any of these has built a click identifier, and should expect to be treated as one.

7.5 Relationship to the working group's privacy text

The Web Bot Auth protocol draft carries three passages that a fetch-to-click token has to answer directly.

The working group charter puts "Authenticating the end user of a participating client or agent" and "Tracking or assigning reputation to particular bots" out of scope. Referral Binding does neither. It is defined only for the identifying path, signatures tagged web-bot-auth, and has nothing to say about the anonymous credential work.

8. Security considerations

A token proves nothing on its own. Anyone can put any 25 characters in a URL. A token is meaningful only when it matches a fetch the origin itself logged and verified. An unmatched token is not evidence of anything, and a site MUST NOT grant access, pricing, or content on the strength of one.

Forgery. To forge a matching token an attacker needs the signature bytes of a real bound fetch of the origin, plus its created, keyid, authority and path. Those are visible to the assistant, to the origin, and to any party that terminates TLS or reads the origin's request logs on its behalf, such as a CDN or a bot management vendor. They are not visible to a passive network observer, because the fetch travels over TLS. An attacker who does have them has observed a fetch that already happened, and can at most claim that a click came from a fetch that genuinely occurred. There is nothing to steal by doing so.

Guessing. 96 bits. An attacker with no knowledge of any fetch cannot produce a matching token.

Enumeration. An origin MUST NOT expose a lookup endpoint that answers whether an arbitrary token matches a fetch. The join belongs in the origin's own analytics, not in a public API. The validator in Section 14 is not such an endpoint: it derives a token from signature headers the caller supplies, and it holds no fetch log to look anything up in.

Replay. A token can be replayed as many times as someone likes, in as many URLs as they like. This is not an attack on the origin's accounting so much as noise in it, and it is bounded the same way every click parameter is: by the join window, and by the fact that a replayed token points at a fetch that really happened. Origins that care SHOULD count distinct tokens rather than total tagged clicks.

Not a secret, not a credential. The token is neither. It is safe in logs, in referrer headers and in analytics exports, and it confers nothing on the bearer.

Why the token is derived rather than asserted. An assistant could simply state its identity in a parameter, and that would be far easier to adopt. The study's fetch tests are the argument against it. In four observed requests from one assistant, the user agent, the operating system, the network and the Referer were all asserted, and every asserted value we could check was either absent or false. A self-declared referral marker would be one more field of the same kind: correct when the operator is honest and its client is working, worthless otherwise, and indistinguishable from either case at the origin. Deriving the token from the signature means a claim about a fetch can be checked against that fetch, which is the only property in this design that survives an uncooperative caller.

Heuristics are not a substitute. The Referer and sec-fetch-site contradiction the study records is a real detector and it costs nothing to run, but it names no operator, it produces no join, and it disappears the moment a client emits consistent headers. Origins SHOULD treat such checks as noise filters, never as attribution.

An origin can mint its own tokens, so this is not settlement data. Every derivation input is known to the origin from its own fetch log, which is what lets it recompute the token in the first place. The same property lets it compute tokens for fetches that no human ever clicked and report them as referrals. Nothing in this document detects that, and Section 7.1 deliberately gives the assistant no callback that would. So a token is evidence to the origin about its own traffic, and it is not evidence to a counterparty. Any arrangement that pays on referral counts needs a mechanism this document does not provide: a token only the assistant can produce and anyone can verify, which means a signature over the click rather than a hash of the fetch, at a cost of roughly 86 characters instead of 25. The version character in Section 4.5 is the hook for specifying that, and Section 13 asks whether it should be.

Expiry. The expires parameter of the original signature is a bound on the fetch, not on the token. An origin verifies the signature at fetch time. By the time the click arrives the signature may be long expired, and that is expected. The join window in Section 6 is what bounds the token.

9. Extensions

9.1 Citations served from an index

Assistants cite pages they did not fetch during the conversation. In the measurements behind Section 1, 26% of tagged assistant clicks had no fetch of the site in the previous seven days.

There is no bound fetch to derive from in that case, and this version defines no token for it. The natural extension is to derive the token from the signed crawl request that populated the index, which would let an origin join a click to a crawl that may be weeks old. That trades a much longer join window, and a much larger index on the origin's side, for coverage of the common case. It is left to a future version rather than guessed at here.

9.2 Pre-committing the token at fetch time

An origin that wants to know the token before the click can ask for it. The @query-param component lets an assistant include a parameter in the signed request, and the protocol draft's Accept-Signature mechanism lets an origin state which components it wants covered. An assistant could then fetch with the token already in the URL. This is a heavier arrangement than the recomputation described in Section 6, it requires the origin to opt in, and it is documented here only to note that it is possible.

Two things make it awkward beyond its weight. Cloudflare's verifier does not support the component, spells it @query-params in its documentation, and recommends signing @query instead, so anything built on this path should be tested against real verifiers first. And RFC 9421 Section 2.2.8 does not take the parameter's value from the URL as written: it parses the query per the HTML form-encoding rules into name and value pairs and then re-encodes the value, so what gets signed is a normalized form of the value rather than the bytes on the wire. For a token drawn only from lowercase alphanumerics that normalization is the identity, so it costs nothing today; it means the component's behavior for this parameter is incidental rather than guaranteed, and a longer or differently shaped future token would not get that for free. The same section RECOMMENDS signing the whole query string with @query where an application commonly covers several parameters.

9.3 A signed intent header

The protocol draft observes in Section 5.2.4 that "an agent might include an HTTP header expressing its intent and sign it." A Referral-Binding request header stating that the assistant intends to emit a token for this fetch would let an origin log selectively rather than logging every verified request. If defined, it MUST be an RFC 9651 structured field and MUST be covered by the signature. This document does not define one. The query parameter convention needs no protocol change at all, and that is its main advantage.

10. Relationship to existing work

ExistingWhat it doesWhat it does not do
Cloudflare Attribution Business Insights (2026-07-01)Reports a crawl-to-referral ratio per operator, where a referral is a visit "tracked through UTM parameters"Any per-click join, and nothing at all for operators that do not tag their links. Referral Binding is the assistant-side input this product is missing.
Cloudflare's Radar crawl-to-refer analysis (blog, 2025-07-01)Ratios based on the Referer hostNative app traffic, by Cloudflare's own admission
Forwarded: for="openai";use="reference" (experimental)An operator-level statement of intent on the inbound fetch, where use=reference means "index, excerpt, and link back"Any per-request identifier, and anything on the click. Referral Binding is how an origin would measure whether the linking back actually happened.
utm_source=chatgpt.comThe only vendor-documented AI referral parameter, broader in practice than its documentation claims, and unaffected by the native apps: we logged tagged arrivals carrying no Referer at allAny utm_medium or utm_campaign value, any identification of which fetch produced the link, and any way for the origin to check the claim against its own records. Section 2.1 sets out what this document adds over it, and what it does not
oppref (ChatGPT Ads)An opaque per-click token minted server sideOrganic coverage, and any way for a site to verify it
adview_query_id (Google AI Mode ads)A per-query token on some ad clicksDocumentation, and any confirmed presence on organic links
RSL 1.0Requires "visible credit and a functional link" under its attribution payment typeAny way to measure whether the link was honored
IAB Tech Lab CoMP 1.0License tokens for content useAttribution. The specification does not mention it.

As of 2026-09-07 we found nothing in the working group's drafts, its 333 archived list messages, its two GitHub repositories or its four sets of minutes that addresses referrals, click attribution, or carrying a fetch identifier across to human traffic. The protocol draft does use the word "attribution" throughout, in a different sense: the lookup that attributes a request to a Signature-Agent URL. The nearest neighbor to this document is an unanswered 2025 request in the key directory repository for a pseudonymous session identifier for abuse reporting, which the same token would serve.

11. Test vectors

Both vectors are generated by the reference implementation that serves the validator at https://verifyagents.org/proposals/referral-binding, and are asserted against it in its test suite: the suite parses the headers, the derivation string, both tokens, the stated key id, the published JWK and the rendered link out of this section and compares each against what the code produces, so this section cannot drift away from the implementation without turning the build red.

The key is a throwaway generated for this document and signs nothing real. Both vectors are fixed in time, so their signatures are expired, which is the normal state of a fetch signature by the time a click arrives. Expiry has no part in the derivation. The protocol draft recommends an expiry of no more than 24 hours, so no fixed vector can stay verifiable; a verifier that checks expires against the wall clock will decline to re-confirm these two, and the tokens still derive.

{
  "public": { "kty": "OKP", "crv": "Ed25519", "x": "QJ1Xn3k6gL7dEzCeM95oAweTRmja1VzNHeZfRnQCnAA" },
  "private": { "kty": "OKP", "crv": "Ed25519", "x": "QJ1Xn3k6gL7dEzCeM95oAweTRmja1VzNHeZfRnQCnAA", "d": "-skkN7o8nrSL57bU7u6BWJQGUgeTAZh3rcv2N0O5aXI" }
}

Its RFC 7638 thumbprint, and the keyid in both vectors, is S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM.

11.1 Vector A: no nonce

Request: GET https://example.com/pricing, with Signature-Agent: "https://assistant.example".

Signature-Agent: sig1="https://assistant.example"
Signature-Input: sig1=("@authority" "@path" "signature-agent";key="sig1");created=1788264000;keyid="S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM";alg="ed25519";expires=1788264300;tag="web-bot-auth"
Signature: sig1=:gJZNIMngohl6J4SCJJPCCjWW6B+HRc2QtZ4u5mD4IAWhis1t9jM+puyFUQPlJBYoz560fqGalKZGmPOJ2Xj1Bw==:

Derivation string, with \n shown as a line break:

web-bot-auth-referral-binding/1
gJZNIMngohl6J4SCJJPCCjWW6B-HRc2QtZ4u5mD4IAWhis1t9jM-puyFUQPlJBYoz560fqGalKZGmPOJ2Xj1Bw
1788264000
S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM
example.com
/pricing

Token: 19be2817c5c441df35d7561bd

Rendered link:

https://example.com/pricing?wbab=19be2817c5c441df35d7561bd&utm_source=assistant.example&utm_medium=ai-assistant

Note the signature field in the derivation string: the Signature header carries standard base64 with padding, and the derivation uses unpadded base64url of the same bytes. This is the one place an implementation is most likely to go wrong.

11.2 Vector B: the same request with a nonce

Signature-Agent: sig1="https://assistant.example"
Signature-Input: sig1=("@authority" "@path" "signature-agent";key="sig1");created=1788264000;keyid="S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM";alg="ed25519";expires=1788264300;nonce="AAECAwQFBgcICQoLDA0ODxAREhMUFRYXGBkaGxwdHh8gISIjJCUmJygpKissLS4vMDEyMzQ1Njc4OTo7PD0+Pw==";tag="web-bot-auth"
Signature: sig1=:nCj+XQ3P+4E+gEpjgRXYuBA07LE2rE18PjusmQRLUnChPT8IkDPywnqLEEzX9eJHp3C3RBZHV2FobdlFjYgDCw==:

Token: 1b0a762ac565e10d47c5fa449

The token differs from Vector A even though the derivation never names the nonce, because the nonce is part of the @signature-params line that the signature covers. This is the demonstration behind Section 4.1.

12. Adoption notes

What each operator would have to do, given what it publishes today. Directory status checked 2026-09-08 and tracked at https://verifyagents.org/observatory.

OpenAI. Furthest along on the click side, and the operator with the clearest missing piece on the fetch side. It publishes a live key directory at chatgpt.com with one Ed25519 key and purpose: "ai", and it appends utm_source=chatgpt.com broadly: in a logged-in session on 2026-09-08, both search citations and inline brand links carried the parameter, with rel="noopener" rather than rel="noreferrer". It also runs at least one negotiated per-partner scheme, tagging Yelp links with utm_source=openai&utm_medium=feed_v2&adjust_creative=openai, which shows the rendering path already accommodates a counterparty's parameters.

The fetch side is the gap. Asked to fetch a URL that echoes the request headers back, ChatGPT returned a request carrying User-Agent: ...ChatGPT-User/1.0... and no Signature, Signature-Input or Signature-Agent header at all. The real-time fetcher behind a citation is unsigned; only agent mode signs. Adoption is therefore two steps, not one: sign ChatGPT-User requests, which five verifiers already validate and which needs no new key, and then derive and append the token on the rendering path that already appends parameters. Section 9.1 covers the remaining case where a citation comes from the index with no live fetch behind it.

Google. Signs a subset of requests at agent.bot.goog, with a self-signed directory, and already stamps a per-query token (adview_query_id) on some AI Mode ad clicks. The fetch side is compatible today. The click side is the harder ask, since AI Mode traffic is deliberately not separated from Web search. The narrow version of the pitch is to standardize the token Google already mints, and to apply it in the Gemini apps, whose iOS surface is reported to identify itself in the user agent.

Anthropic. Publishes IP ranges for its fetchers but no key directory, so it would need to stand one up first. It has taken no public position on link tagging. A request for UTM tagging was filed in a Claude Code repository in April 2026 and closed automatically by a bot as off topic for that repository, with no human reply, so nothing has actually been declined.

Perplexity. The furthest from adopting this document as written, and the operator with the most to gain from solving the problem. Comet Plus pays publishers on a mix of "human visits, search citations, and agent actions" with an undisclosed attribution method, so a shared, computable definition of a referral is worth more to Perplexity than to anyone else. Three obstacles sit in front of it, and two of them are structural rather than a matter of will.

First, it publishes IP ranges but no key directory, so it cannot derive a token from a signature it does not produce. Second, its citation surface does not fetch: asked three times to retrieve a URL, against three different targets including one on a zone we control and had built for the purpose, its search mode reported a failed or completed fetch and sent no request in any of the three, which we confirmed against our own edge logs. The third of those is the run recorded in the study. A per-fetch token is undefined for an assistant that answers without fetching, so Section 9.1, the index-derived extension this version defers, is not an optional extra for Perplexity; it is the only form of this proposal that could apply to its citations. Third, its agent surface fetches without identifying itself at all (see the study).

The honest pitch is therefore not "adopt this parameter." It is that Perplexity currently cannot prove a referral to the publishers it pays, that no third party can verify one on its behalf, and that both facts follow from choices it could reverse: publish a directory, sign the agent fetch, and work with us on the index-derived variant.

Microsoft, xAI, Meta, DeepSeek, Mistral. No key directory as of 2026-09-08. Web Bot Auth adoption is the prerequisite, and Referral Binding is a reason to do it.

Verifiers. As of 2026-09-07, Cloudflare, Vercel, Akamai, AWS WAF on CloudFront and DataDome all verify Web Bot Auth signatures, and none of them exposes anything that joins a verified fetch to a later click. Nothing in this document requires a verifier to change: the derivation runs on inputs an origin can log for itself, and a site behind any of these verifiers can implement the join alone, subject to the logging and caching caveats in Section 6. A verifier that exposed the six derivation fields, or the token itself, to the origin it fronts would save every site behind it that work, and would be the single cheapest thing any party could do for this proposal.

13. What we are asking

These are real questions, not rhetorical ones. We have spent more time on the measurements than on the mechanism, and we are a good deal more confident in the first than the second.

The three that matter most, before any of the detail below:

  1. Has this been done? We searched the working group's drafts, its archived list messages, its two repositories and its minutes, and found nothing addressing referrals or carrying a fetch identifier across to human traffic. The nearest neighbor was an unanswered request for a pseudonymous session identifier for abuse reporting. We would rather be told we missed something than duplicate it.
  2. Is the problem worth solving at this layer? Section 2.1 concedes that a UTM convention gets most of the way for a fraction of the effort. If the right answer is "campaign for link tagging and stop there," that is a useful answer.
  3. Does the origin-can-mint limit in Section 8 sink it? Every derivation input is known to the origin, so the token is evidence to a site about its own traffic and not to a counterparty. The fix is a signature over the click instead of a hash of the fetch, at roughly 86 characters instead of 25, and it puts a signing key on the rendering path. We do not know whether that trade is worth making, and it is the single most consequential thing we cannot decide alone.

The narrower ones:

  1. Should the join window be 7 days or 30? Seven matches the shortest comparable windows and keeps the origin's index small. Thirty would match ad platform convention and cover conversations people return to weeks later.
  2. Is a Referral-Binding request header worth defining, so origins can log selectively instead of logging every verified request? Section 9.3 argues not yet.
  3. Should Section 9.1 be folded into this version rather than deferred? Index-served citations are the larger share of the problem, and a specification that covers only live fetches solves the smaller half.
  4. Should this version define the signed, origin-unforgeable token described in Section 8 rather than defer it? A hash is short and needs no verification step; a signature is the only thing that makes the token evidence to anyone but the origin. The answer probably decides whether this is an analytics convention or an accounting one.
  5. Is wbab the right parameter name, or should it be longer? It is unused in the wild, on no published strip list, and names the specification family, but it is opaque to a person reading a URL and, as Section 15 sets out, there is no registry for query parameter names and therefore nothing that prevents a collision. A longer name buys collision resistance at the cost of URL length, and four characters against a 25-character value is not much of a saving.

14. Try it

The validator at https://verifyagents.org/proposals/referral-binding takes a landing URL, the fetch's Signature-Input and Signature headers, and the URL that was fetched. It derives the token, reports whether it matches, and runs the full RFC 9421 signature debug on the fetch. Both vectors above load into it as examples.

The API is one POST:

POST https://verifyagents.org/api/referral-binding/verify
Content-Type: application/json

{
  "click": "https://example.com/pricing?wbab=19be2817c5c441df35d7561bd&utm_source=assistant.example&utm_medium=ai-assistant",
  "signatureInput": "sig1=(\"@authority\" \"@path\" \"signature-agent\";key=\"sig1\");created=1788264000;keyid=\"S08iZl-f8-8esBZmQrrzjj8eBHX48R_IehKH0xqmGvM\";alg=\"ed25519\";expires=1788264300;tag=\"web-bot-auth\"",
  "signature": "sig1=:gJZNIMngohl6J4SCJJPCCjWW6B+HRc2QtZ4u5mD4IAWhis1t9jM+puyFUQPlJBYoz560fqGalKZGmPOJ2Xj1Bw==:",
  "targetUrl": "https://example.com/pricing",
  "signatureAgent": "sig1=\"https://assistant.example\""
}

15. IANA considerations

A registry for the version character. Section 4.5 makes one character carry the whole forward-compatibility story, and a character with no registry behind it is not a mechanism. This document requests a registry, "Web Bot Auth Referral Binding Versions", with:

The parameter name has no registry to go in, and that is a real weakness. There is no IANA registry for URI query parameter names. wbab is claimed by convention and nothing prevents a collision; the mitigations available are that the name is currently unused in the wild, appears on no published parameter-stripping list, and names the specification family. Section 13 asks whether a longer and more collision-resistant name would be a better trade against the cost of a longer URL. A reviewer should read the absence of a registry as an argument for the longer name, not as an argument that the question does not matter.

No other actions. This document defines no new HTTP header field, no media type, no URI scheme and no signature parameter, and it registers nothing in the RFC 9421 or Web Bot Auth registries.

16. Implementation status

This section records the state of implementation at the time of writing, in the form RFC 7942 asks for, and is to be removed before any publication as an RFC. Information current as of 2026-09-08. RFC 7942 places this section before Security Considerations; it sits here because this document is read as a web page first, and it should move before Section 8 on conversion to an Internet-Draft.

One implementation exists, by the editor. It is the code that serves this document. It implements draft-00, this document, and no other version.

17. References

Normative. RFC 9421 (HTTP Message Signatures); BCP 14 (RFC 2119, RFC 8174); RFC 3986 (URI Generic Syntax); RFC 4648 (Base64 encodings); RFC 5234 (ABNF) and RFC 7405 (case-sensitive ABNF strings); RFC 5890 (Internationalized Domain Names); RFC 7638 (JWK Thumbprint); RFC 8941 (Structured Field Values); RFC 9110 (HTTP Semantics); RFC 9651 (Structured Field Values, which obsoletes RFC 8941 and which Section 9.3 refers to); draft-ietf-webbotauth-httpsig-protocol (Web Bot Auth).

This document cites RFC 8941 for structured fields, following RFC 9421, and RFC 9651 only where Section 9.3 describes a future header field. The two agree on everything this document relies on.

Informative. RFC 7942 (Improving Awareness of Running Code); RFC 8032 (EdDSA); Cloudflare's bot verification and Attribution Business Insights documentation; Google Analytics 4 channel-grouping documentation; Brave's and Mozilla's query-parameter stripping policies; Apple's Web AdAttributionKit documentation; RSL 1.0; IAB Tech Lab CoMP 1.0. Measurements cited in Section 1 are the editor's own and are described in Section 18.

18. Acknowledgments and provenance

The measurements summarized in Section 1, and set out in full at https://verifyagents.org/study, come from a fleet of 174 Cloudflare zones instrumented at the edge through 2026, reported at https://completeseo.com. The Web Bot Auth facts are drawn from RFC 9421, the working group's protocol draft, Cloudflare's bot verification documentation, and live checks of operator key directories run on 2026-09-08 and repeated daily at https://verifyagents.org/observatory. Claims in this document that rest on third party reporting rather than a primary source, in particular the details of oppref and adview_query_id, are identified as such in the research notes published alongside it.

This is a proposal from an independent party. It has not been submitted to the IETF, and no operator named in it has endorsed it.

Validator

Derive the token from a signed fetch and check it against a landing URL. The two test vectors above load as examples.

Load