Methodology
What Plumb measures, samples, infers, and refuses to guess.
You cannot see inside an AI model. You can see what it does to your site. Every number in Plumb says which of these it is, so an editor, a lawyer or a CFO knows how far to trust it.
01Who it is for
Publishers where one article is worth real money
Plumb closes a loop: measure what AI takes, find the queries worth writing, draft the brief, measure again after it ships. That loop only pays for itself when a single well-placed article earns enough to matter.
General news and programmatic publishers get the measurement side, which is useful on its own for bot policy and licensing conversations, but the commissioning loop will not change their economics. Plumb is open about that rather than pitching to everyone.
02First-partyMEASURED
What is measured
Full coverage, from systems you own. Nothing here depends on asking a third-party model anything.
- Every agent request, at your edge
A Cloudflare Worker on your content routes reports each bot request to Plumb: the user agent, the path, the status, the time. The client IP is hashed at the edge and the raw address is never stored. Raw events are kept for 14 days; daily rollups are kept for good. Not on Cloudflare? Server logs can be uploaded instead.
- Which agents are who they say they are
A User-Agent string is a claim. Plumb checks each request against the IP ranges its vendor publishes (twelve feeds today, refreshed daily) and against RFC 9421 signatures where a vendor signs. A request from inside the ranges is verified. A known agent's name arriving from outside them is spoofed. A vendor that publishes nothing, which today includes Anthropic, leaves its agents unverifiable, and Plumb says so rather than guessing.
- What your robots.txt said, and who listened
Plumb fetches your live robots.txt daily and sets it against the edge log. For each agent it records the rule that applies, the requests to paths that rule disallows, and a verdict: honored, ignored, allowed, or no rule. Agents named in the file that never showed up are listed too.
- Search Console
Impressions, clicks and rank per query, with a flag for queries where an AI Overview rendered. This is how Plumb finds queries where impressions held while clicks fell, without scraping a results page.
- Readers sent back
Two sources. GA4 sessions whose referrer is an AI assistant (chatgpt.com, claude.ai, perplexity.ai and the rest of a maintained list), and referrals the Worker sees directly at the edge. Plumb reports crawls per referral for each platform from the same log, so the two sides of the ratio come from the same place.
03ProbesSAMPLED
What is sampled
Plumb asks the assistants questions on a schedule and records what comes back. It is the only way to see citations from the outside, and it is noisy.
- The weekly citation panel
500 questions across the topics that matter to affiliate and trade publishers, put to ChatGPT, Claude, Gemini and Perplexity every Monday. Plumb records which domains each answer cites and in what order. This is where share of voice and emerging demand come from.
- Why it is noisy
The same question, on the same model, at temperature zero, in the same week, returns different citations hour to hour. Moves under 15 points week over week are inside the noise. Plumb shows trend lines, not single readings, and never uses the panel for an absolute claim about your visibility.
04ModelsINFERRED
What is inferred
Estimates, scores and text that a model produced from the measured and sampled data. Useful for ordering decisions. Never a fact on its own.
- Priority and recoverable value
The priority score on a bleeding query weighs impressions, the size of the click drop and how close you sit to the top. The dollar figure applies your own conversion economics to that. Both are orderings for a planning meeting, not facts, and carry the inferred label wherever they appear.
- What Plumb's models write
Page recaps, the briefing paragraph on the overview, commission briefs, proposed edge rules and answers from Ask Plumb are written by a model that can only call Plumb's own tools. Every line names the tool it called, and the numbers keep the label they had in the tool result. A recap can be wrong about emphasis; it cannot invent a number that is not in the data.
- Readiness scores and story diagnostics
The article readiness score weighs structure, freshness, metadata and similar factors against published citation research. The label on an ignored story (schema, timing, paywall, stale, duplicate) comes from a small decision tree. Treat both as hypotheses to check, not conclusions.
- The hidden AI share of direct traffic
A Bayesian estimate of how many direct sessions began in an AI conversation, from the conversion lift of direct over organic. The output is a range. The sweep chart shows how far it moves when the prior moves, because it should.
05Unknowable
What Plumb will not claim to know
- What a model thinks of your content
Retrieval scores, weights, and why an answer cited a competitor instead of you are private to the model's operator. No outside tool can see them. Plumb can correlate crawl concentration with citation outcomes; it cannot explain the choice.
- What real users are asking
The questions typed into ChatGPT or Perplexity are not shared with publishers. Plumb samples a panel of questions it writes. A product that shows you real user queries is either sampling too or making them up.
- Whether a page ended up in training data
A crawl log shows a bot fetched a page. It says nothing about what happened next. Statements about training-set inclusion are speculation, and Plumb does not make them.
06The panelPANEL
How publishers are compared
When several publishers run Plumb, it can show a first-party median instead of citing someone else's research PDF.
- Three publishers minimum
No cross-publisher figure is shown below three participating tenants. Under that, the dashboard shows the published research figure and names its source.
- Quantiles only
The panel returns a median with the 25th and 75th percentiles. No minimum, no maximum, no per-tenant rows. Your own tenant contributes but cannot be picked out.
- Computed live
Aggregates are computed on request with a five-minute cache. There is no archive. Leave the panel and your data leaves it at the next refresh.
07Independence
Why not the marketplace's dashboard
Two kinds of vendor already show publishers numbers about AI traffic. Both have a stake in what the numbers say.
08Not in scope