The Token Hustler Needs Ten Interrogations Before the Measurement Means Anything

​

Unleash this on:

Original insights by Gregory Druck and Ethan Smith

30-second rundown

Key learnings:

  • One answer is not a measurement: AI systems can answer the same question differently because each response is generated through random sampling.
  • Repeat important prompts: Ten fresh runs can provide a useful directional estimate of how often a brand appears, while more runs improve precision.
  • Measure the uncertainty: Compare results weekly or every two weeks, use confidence ranges, and spot-check the logged-in experience before treating a movement as real.

This week: Choose three commercial comparison questions that matter to your business, ask each one ten times in separate new chats with memory disabled, record whether your brand appears and where, then calculate its appearance rate and flag any result based on only a few mentions.

Coming 3 months: Set up an AI workflow to repeat the approved questions weekly or every two weeks, and collect enough responses to reach a chosen confidence threshold.


Gregory Druck and Ethan Smith have wandered into the AI measurement casino and found marketers celebrating one spin of the wheel as if Moses carried it down from the mountain. Their Graphite study delivers the necessary slap: an AI answer is not a fixed search result, so one response cannot tell you what the machine consistently believes about a brand.

The machine is not lying, exactly. It is improvising under mathematical rules, which is how a sober dashboard can become a carnival mirror before lunch.

1. The Answer Changes Because the Dice Are Inside the Engine

A large language model writes an answer one token at a time. A token is a small piece of text, often a word or part of one, chosen from a probability distribution of plausible next pieces. AI randomness is built into generation, so the same prompt can send the model down different roads even when nothing else appears to change.

An early choice compounds through every choice that follows. One run may recommend your brand, the next may omit it, and a third may bury it below a rival like a body behind a desert motel. Treating any one of those answers as the verdict confuses a sample with the underlying pattern.

2. Ten Runs Can Be Useful, but They Are Not Holy Writ

Druck and Smith tested 200 brand-comparison prompts and collected 400 responses for each prompt from the ChatGPT model named GPT-5.2 Chat. They used 200 responses per prompt as a practical estimate of the underlying result, then repeatedly tested smaller samples against it. That gave them a way to measure how far a quick check might wander from a much larger one.

With ten responses, the average error in brand visibility was 5.6 percentage points overall and 9.1 points for brands that appeared at least 10% of the time. Visibility simply means the share of answers that mention a brand. Ten runs give a useful directional estimate, but a brand shown in five of ten answers should not be treated as if 50% were carved into federal law.

The pattern held when the researchers repeated the test with Gemini 3 Flash: the average ten-response error was 5.5 points overall and 8.6 points for brands above that 10% visibility threshold. Across their prompts, 98.6% of ten-response estimates landed within ten percentage points of the larger estimate. At forty responses, 94.9% landed within five points, which is why more sampling buys tighter confidence instead of mystical certainty.

3. Track a Distribution, Not the Machine’s Daily Mood

The clean method is sequential sampling: keep collecting fresh answers until the confidence range around the result is narrow enough for the decision you need to make. A confidence range is the honest band around an estimate, the statistical equivalent of admitting that the desert is moving under your boots.

Run each test in a new chat with memory disabled so previous conversation does not contaminate the next answer. Measure visibility first for brands that appear inconsistently, then pay closer attention to position once a brand is present often enough for its ordering to mean something. Tools are directional, not ground truth, because API results, logged-out answers, and logged-in answers can differ.

Graphite recommends weekly or every-two-week tracking for most teams. Daily checks can work, but only if nobody mistakes ordinary variation for a strategic earthquake. Add a small set of manual logged-in checks to see whether the customer-facing experience resembles the automated sample before the dashboard priests begin sacrificing the content calendar.

The point is not to eliminate randomness. You cannot bully probability into holding still, and anyone selling that miracle is already reaching for your wallet.

Ask the same important question enough times, record the spread, and make decisions only when the signal escapes the noise. Otherwise you are steering by one flickering headlight while the machine rolls another blank metal ball into the dust.

Unleash this on:
Avatar photo
WH.

All the paranoia of a field correspondent. None of the plane tickets.

WH. has spent 14 years inside the SEO machine and started The Vector Gazette, because he got tired of watching entrepreneurs make catastrophic decisions based on advice from people who discovered GEO last Tuesday.