DocsGetting started / Agent traffic

Agent traffic

People now tell an AI assistant "book me a plumber for Thursday", and it goes to a website and tries. This is the reference for the instrument that counts those visits.

The tracker exists to answer one question nobody has an honest number for: how much AI-assistant traffic does an ordinary service business actually get, and how often does the assistant try something and fail?

It observes and classifies. It never executes an action, never writes to the host site, and has no code path to either. There is no Action Graph in it, no endpoint for an agent to call, and no way for it to touch a form.

Four verdicts

What a visit can be classified as
VerdictMeaning
humanA person, on the site, behaving like a person. This is the default.
agentSoftware acting on the site, declared by its user-agent, presenting a signed identity, or strongly fingerprinted.
suspected-agentReal but ambiguous evidence. Deliberately not rounded up to agent.
bot-otherA bot that is not an agent: search crawlers, SEO tools, monitors, and AI training crawlers.

The classifier is a pure, synchronous, dependency-free function that runs identically in the browser, in the collector and in the tests. The verdict that matters for the measurement is agent, an assistant acting for a person. It is line 2 of the report, never line 1.

Five families of evidence

What fires, and what it produces
FamilyWhat firesVerdict
1. Known agent user-agentsChatGPT-User, Claude-User, Perplexity-User, Operator, NotebookLM, an assistant acting for a user.agent, confidence 0.90
1b. AI crawlersGPTBot, ClaudeBot, PerplexityBot, CCBot, Bytespider, Applebot-Extended, Google-Extended, Amazonbot, meta-externalagent.bot-other, platform recorded
2. Web Bot AuthThe Signature, Signature-Input and Signature-Agent headers, a signed agent identity.agent, confidence 0.95–0.97
3. AI referrerschatgpt.com, perplexity.ai, claude.ai, gemini.google.com, copilot.microsoft.com.Stays human, recorded as its own signal
4. Headless and browser-agent fingerprintsA headless user-agent, navigator.webdriver, automation globals, a permissions-surface mismatch, an empty language list, no plugins on desktop Chrome, a mobile user-agent at a desktop viewport, no pointer, scroll or key events, and programmatic or inhumanly fast form fills.agent at two or more signals, otherwise suspected-agent
5. Failed form interactionsWhich form, which field, and how it died: validation-rejected, captcha-blocked, no-submit, abandoned.Contributes nothing to the verdict, captured as evidence

How the signals add up, and why no single one is enough

Signals the classifier reads, strongest first
SignalStrength
Headless user-agent (server-side, trustworthy)Strong
navigator.webdriverStrong
Automation globals presentStrong
Inhuman fill speedMedium
No pointer, scroll or key eventsMedium
Fields set without keystrokesMedium
Zero languages declaredMedium
Mobile user-agent at a desktop viewportMedium
No plugins on desktop ChromeWeak
Permissions-surface anomalyWeak
Strengths add up. It takes at least two independent signals to count a visit as an agent, so no single self-reported signal can do it alone: an automation global on its own lands on suspected-agent, and that is intended.

Why it takes two

1.00agentThe line, and at least two independent signals. The heaviest single weight is 0.70, so one signal never reaches it.0.45suspected-agentReal evidence, carried as ambiguity. Below this line the verdict stays human and the signal is still written down.

  1. One weak signal0.20No plugins on desktop Chrome0.20humanShort of the first line. The default holds and the signal is recorded, changing nothing.
  2. The heaviest signal, alone0.70Headless user-agent0.70suspected-agentRead from the request itself, and still short of the agent line by construction. Counted as suspected, never as agent.
  3. The same signal, and one more0.70Headless user-agent0.35No pointer, scroll or key events1.05agentTwo independent signals, over the line: an agent. With confidence, not certainty, nothing reaches 1.0.
  • A signal in the request itself, the user-agent the server saw. Trusted.
  • A signal the browser reported about itself. Self-reported, untrusted, and never enough alone.
  • The agent line, 1.00: the bar the classifier sets. No single weight reaches it.
Three visits, drawn, and none of them real: worked examples of the rule, not rows from any store. The weights and both lines are illustrative values for the worked examples; the sums beside the drawing are done from them, not written down. A block is exactly as tall as its weight.

The default classification is human. Real evidence is needed to say agent, and ambiguity is carried rather than rounded up.

Server-side evidence, the user-agent and the request headers, is trusted. Client signals are self-reported, so on their own they can never add more than 0.75 confidence. Confidence never reaches 1.0: nothing is verified cryptographically, so the ceiling is 0.97.

The result is that the tracker under-counts agent traffic rather than over-counting it. If a report says agents are visiting, they are. If it says they are not, read that as a floor.

Three ways this could quietly lie, and what stops each

  1. Counting AI crawlers as agents. Training and indexing crawlers visit nearly everything; if they counted, every site would show agent traffic on day one and the measurement would return a guaranteed false positive. They are classified bot-other with the platform recorded, and the report shows them on their own line. That split was put to a decision and confirmed: the headline number counts only assistants acting for a user.
  2. Counting AI referrals as agents. An AI referrer means a person followed an AI answer here. Commercially interesting, and not an agent acting on the site. A referrer never moves the verdict off human.
  3. Counting failed forms as agent evidence. People fail forms constantly. Using failure as evidence of automation would manufacture exactly the false positives that would corrupt the measurement. A failure only becomes a failed agent action once the verdict independently says agent or suspected-agent.

The tracker detects the Signature, Signature-Input and Signature-Agent headers but does not verify the signature: no key fetch, no directory lookup, no cryptographic check.

So the claim it makes is "this request presented a signed agent identity", not "this is a verified agent". The report says "presented (unverified)" and keeps confidence below 1.0 for that reason.

What it collects

  • The site’s own domain name, and a reduced page path. Only the first segment of the URL is kept, and only when it does not look like an identifier. The query string is never seen, so nothing after a ? is ever collected.
  • Which site sent the visitor, as a domain only.
  • The browser’s user-agent string, as every web server already receives on every request.
  • Signals that distinguish a person from software: whether the browser reports itself as automated, how many languages and plugins it declares, the window size, counts of mouse, scroll and keyboard events, and whether the browser’s own notification-permission settings are self-consistent. None of these identify a person; they describe the software.
  • If a form is used: the names and types of the fields, how long each took to fill, how many keystrokes were involved, and whether the form succeeded, was rejected by validation, hit a CAPTCHA, or was abandoned. Where a field was rejected, one word from a fixed list of nine, required, email, number, format, length, range, select, url, other, never the wording of the message.

What it never collects

  • No cookies. Nothing is stored on the visitor’s device, no cookie, no localStorage, no sessionStorage, nothing.
  • No identifier. No visitor ID, no fingerprint hash, no cross-site tracking. Nothing is created or stored whose job is to name a visit or tie it to another one.
  • No IP addresses. Not logged, not stored.
  • No form contents, ever. The names typed in, the phone numbers, the email addresses, the message text, none of it is read, stored or transmitted. The only question the tracker ever asks about a text field is whether it is empty.
  • No password, hidden or file fields at all. These are not observed even as shapes. A keystroke count on a password field approximates the length of somebody’s password, so the field is excluded at the single gate every field observation passes through, and the collector drops any such record a forged beacon tries to supply.
  • No tick-box answers. For checkboxes and radio buttons the ticked state is the answer, whether somebody consented to marketing, whether they have private health cover, so the tracker does not read it. It records that the box was interacted with and when, never which way it was set. The collector enforces this a second time.
  • No error-message wording. A website can put anything into its own validation messages, including a customer’s name, so the message never leaves the page.

Privacy limits

Three limits apply to the data the snippet collects. If you forward the privacy summary to a client, forward these with it.

A one-word top-level path, such as /jane-smith, is stored as written, because it looks the same as /hot-water-repair. A site with person-named top-level pages should hold off installing the tracker until a per-site exclusion list exists.

The page-wide keystroke count includes keys typed into fields the tracker does not observe. On a page whose only input is a password box, that count approximates the password’s length. It stays because dropping it would make a person who only typed a password look like a browser with no human input, which is a false agent verdict.

Rows in a small single-site data file are weakly linkable. Each row holds the full user-agent, the window size, the plugin and language counts and a millisecond timestamp. The tracker creates no identifier and does no linking, but someone with direct access to the raw rows could group repeat visits from the same browser. The accurate claim is that no identifier is created or stored, not that linking is impossible.

What the path reduction does and does not hide

Only the first URL segment survives, and only when it is not identifier-shaped: no run of four or more digits, no @, no more than 40 characters. So /patients/jane-smith-0412345678/appointment is recorded as /patients/*, and /0412345678 as /*.

The first segment can still be sensitive as a category, /patients, /invoices, /accounts are kept deliberately, because knowing that agents fail somewhere under /booking is the entire point of the measurement. That costs detail by design: we can see that agents visited and failed under /booking, not which booking page. The form’s own name is recorded separately, and that is the identifier that actually matters when replaying a failed action.

The reduction runs in the browser, so the full path normally never leaves the visitor’s device, and it runs again in the collector, because a beacon is a string a stranger can POST.

Do Not Track and Global Privacy Control are honoured

If the browser sets navigator.globalPrivacyControl or reports doNotTrack === "1", the snippet does nothing at all: no listeners, no beacon, no request. Only doNotTrack === "1" is matched, not the legacy spellings, and not window.doNotTrack.

Two consequences, both recorded rather than hidden. It undercounts human page views, because agents do not set these signals, so the agent count is unaffected and the agent share is biased slightly upward, the only bias in this package that points that way, which is why it is written down twice. And an automated browser can suppress its own beacon by asserting the signal, a real evasion route accepted in exchange for honouring it, which pushes toward under-counting agent traffic.

Retention: 60 days, enforced by the store

Events older than 60 days are deleted, on by default, with no configuration required. That used to be a promise a person kept, the file was append-only with no expiry, and it was the only claim in the client-forwardable summary that the code did not enforce, which matters because an agency repeats that claim to its client on our behalf.

Nothing survives expiry: the row is deleted outright, with no aggregate, no counter and no rollup. That is what makes "then deleted" literally true rather than approximately true. It costs something, and the cost is stated: a report cannot show a trend older than the retention window.

Deletion is irreversible and this data cannot be re-collected, so the prune is conservative. A row whose timestamp will not parse is kept rather than guessed at; an unreadable or missing file is left alone; survivors are written to a temp file and renamed over the original, so an interrupted prune cannot leave a half-written store; and a retention window that is not a whole number of days of at least one is refused at construction, so a typo or a bad environment variable cannot empty the store. Turning retention off has to be typed out explicitly, "off" should never be somewhere you arrive by leaving a flag blank.

The collector treats every beacon as hostile

A beacon body is a string a stranger POSTed. Anyone can send anything, from a fuzzer to a competitor trying to make the numbers say whatever they want. So the collector assumes exactly that:

  • Nothing from the body is trusted. Only known keys are copied; every string is length-clamped and stripped of control characters; every number must be finite and is clamped; every enumerated value must be a member of its list or it is coerced to a safe default.
  • The classification is computed server-side from the real request headers. A verdict, user-agent, platform or signal list present in the body is ignored outright, a client cannot vote on its own classification.
  • Site attribution prefers the Origin or Referer header over the body’s self-reported site, so a beacon cannot credit its traffic to somebody else’s domain.
  • Prototype-pollution keys are never read, so a __proto__ payload has nothing to reach.
  • The response is always 204, whatever the collector decides. It is seen by a stranger’s browser on somebody else’s site, so it says nothing and never fails visibly.
The clamps, as committed
LimitValue
Body size16 KB
Forms per beacon5
Fields per form40
Automation globals recorded8
Site string120 chars
Path string200 chars
Field name64 chars
Field type24 chars
User-agent256 chars
Any duration86,400,000 ms
The duration cap is a day. Anything larger is nonsense or an attempt to skew an average.

The beacon, on the wire

{ v:1, site, path, ref, cs:{ wd webdriver · ag automationGlobals · pc pluginCount ·
  lc languageCount · vw/vh viewport · tp maxTouchPoints · pe pointerEvents ·
  se scrollEvents · ke keyEvents · ms timeOnPage · pa permissionsAnomaly · fm forms } }
form:  { n name · fc fieldCount · fl fields · tt totalFillMs · oc outcome · ff failedField · ft failureCode }
field: { n name · t type · ms fillMs · k keystrokes · fi filled }
Verbatim from the tracker’s own documentation. Short keys because bytes are the constraint; the collector expands them into the descriptive stored event.

The one table that goes stale

The list of known agent user-agents is the part of this package that ages. New assistant products ship constantly, and adding one is a single-line edit with nothing else in the file changing:

Adding an agent

{ pattern: /NewAgent-User\//i, platform: "Vendor", product: "NewAgent-User", uaClass: "assistant" },
Verbatim from the tracker’s own documentation, "Monthly maintenance, the UA table".

uaClass is the load-bearing field: assistant means acting for a user and counts as agent traffic; ai-crawler means training or indexing and does not. The table records where each entry came from.

The list was last reviewed on 25 August 2026 and the next review is due 25 September 2026. It was compiled by a model whose training data ends in May 2026, so anything that shipped after that is missing until a human re-checks the vendor pages named in the file. A stale table undercounts, which biases toward a false negative, the same direction as every other error in this package.

What this instrument has measured so far

1

site is running the tracker: frontlatch.com itself, since 14 September 2026. The collector is live and answering, and the 60-day measurement clock started on that install.

The tracker is free. You get the answer for your site, and Frontlatch gets the answer across all the sites.

Running it

Inside the repository

npm install
npm test                                              # 108 assertions incl. the byte budget
npm run snippet -- --endpoint https://<host>/b --out dist/t.js   # --endpoint is REQUIRED
npm run snippet -- --endpoint https://<host>/b --readable        # readable build, to show an agency
npm run collector -- --port 8787 --store data/events.ndjson         # 60-day retention, on by default
npm run collector -- --port 8787 --store data/events.ndjson --retention-days off   # opt out, explicitly
npm run report -- --store data/events.ndjson --weeks 4 --out weekly-report
npm run report -- --store data/events.ndjson --json   # same summary as JSON
Verbatim from the tracker’s own documentation, "Run it". The assertion count printed there is 108; the suite currently reports 118 checks with the snippet at 5,689 of its 5,800-byte budget. That line is the one that has not been refreshed, not the number.

The snippet has no default collector endpoint. Building it without --endpoint fails with an error instead of producing something that looks ready to ship, and localhost is not a fallback either.

A default would point at a host someone else could register, and the snippet would then send a business’s visitor telemetry to a stranger. The scanner follows the same rule for its own contact address.

See what an AI agent can do on your site.