llmaudit.app

Methodology

How LLM Audit works

Trust is the product. This page is the honest version of what the audit does — which models we query, what the verdict means, and what we deliberately do not claim.

At a glance

An LLM Audit measures whether AI models recommend your brand when buyers ask about your category. We run your buyer-intent prompts against OpenAI, Google Gemini, and Claude (and more providers when enabled), repeat each prompt 3 times, and report how consistently your brand appears against your competitors.

What it is
A verdict on AI answers, not a Google rank and not a score
Models queried
OpenAI, Google Gemini, and Claude, plus opt-in providers
Runs per provider
3, to separate signal from noise
What we report
Presence, confidence count, competitor gap

Models we query

Each audit runs your buyer-intent prompts against these providers' APIs:

  • OpenAI
  • DeepSeekoff by default — opt-in
  • Google Gemini
  • Claude
  • Perplexitywhen enabled — live web-grounded search

We are not affiliated with any of these companies. Provider names refer to the underlying model APIs we call, not the consumer chat products.

Estimated vs. observed signals

The verdict starts from an estimated signal: models assess how likely your brand is to surface for a prompt, based on the site evidence and category context we supply. We do not claim to read live answers from the ChatGPT, Gemini, or Perplexity consumer products, which are closed surfaces that change constantly.

Where a provider performs real retrieval-grounded search (Perplexity, when enabled), its citations are treated as an observed signal — the closest thing to what AI search actually returns — and shown separately from the estimate.

Why there is no score

We used to publish a 0-100 visibility score. We removed it in August 2026, and the reason is worth writing down: the models do not share a ruler. Asked about the same brand on the same day, one model answered 75, another 49 and another 38. Averaging three different rulers gives a number that looks precise and moves when nothing about the brand has moved.

It was worse before we noticed that the prompt never stated a scale at all: one provider was answering with a probability between 0 and 1, which rounded to 1 out of 100 while the same response said the brand does get named. Fixing that cut the swing between identical runs by two thirds. It still left a number we would not defend to a buyer, so it is gone.

What replaced it can be counted and re-counted: a verdict in three bands, how many of the 3 runs each model named you, and how many of your buyer's open questions each model answers with you, a competitor, or nobody. If you want a number, take those counts. They are the ones we would stand behind.

Confidence through repeated runs

LLM outputs are non-deterministic, so a single answer is noise. We query each provider 3 times per audit and report how consistently your brand appears (an X/N confidence count). A brand that appears in 3/3 runs is a far stronger signal than one that appears once.

How the public AI Index is built

The AI Index aggregates audits into a per-category ranking. The order comes from the average of a brand's observations, but the underlying number is not published, for the same reason it is not shown in your own report: it is not stable enough to print next to somebody else's name. The ranking is ordinal. A category is published only once it has at least 5 brands and 10 observations. The Sample column is the raw number of observations behind each row, so you can see how much a position rests on.

Only observations backed by all 3 providers enter the index. This is not a detail. Scores rise with the number of providers that answered, so an audit where one provider timed out or hit a rate limit scores systematically lower for reasons that have nothing to do with the brand. Mixing those into a ranking would measure our own collection reliability and present it as brand visibility. Partial audits are still shown in your own report, where the per-provider breakdown makes the gap visible; they are excluded here, where brands are compared against each other.

A brand needs at least 2 complete observations to be ranked. AI answers are not deterministic: three identical runs of the same brand have returned 45, 30 and 29 here. A single run is a coin flip, and until this rule existed a brand could hold a top position on one. Two runs is a low bar and we say so plainly: it does not make an average precise, it only means one run is not enough to rank someone else's company.

The consequence is that the index is smaller than our dataset, and categories and brands drop out of it when their complete measurements fall below the thresholds. We prefer a narrower index we can defend to a broader one we cannot.

LLM Audit does not rank itself in this index. Not for modesty: our sample is not collected the same way. We re-run our own audit every time we test a change, so we accumulate far more observations than a brand we measured two or three times, under conditions we chose. An average over that sample is not comparable to a competitor's, and it would come from the one company that also controls the method and decides when to measure. We would rather rank a market we are not in than defend a number nobody can check.

How we measure visibility

Short, direct answers to the questions people ask about how the audit works. Each is written to stand on its own.

How does LLM Audit measure AI visibility?

We run your buyer-intent prompts against several model APIs, repeat each prompt 3 times, and record which brands each model names. What you get back is how consistently your brand appears versus your competitors, reported as a count of runs rather than as a score.

What counts as a good result?

A strong result is appearing in most or all runs across multiple providers, because consistency is what separates a durable recommendation from a lucky one-off. A brand named in 3/3 runs is far more visible than one that surfaces once. The gap to your competitors matters more than any single measurement.

What is the difference between an estimated and an observed signal?

An estimated signal is a model assessing how likely your brand is to surface, based on the site evidence and category context we supply. An observed signal comes from a provider that performs real web retrieval and returns citations. We show the two separately so the verdict never overstates what we can actually see.

Why query the same prompt multiple times?

LLM outputs are non-deterministic, so any single answer is noise. Repeating each prompt lets us report how often your brand appears rather than whether it appeared once. Consistency across runs is the signal that matters, and it is what a one-shot check cannot give you.

What we don't claim

  • — We don't claim to query models we don't actually call.
  • — We don't claim to read the live consumer ChatGPT/Gemini/Perplexity apps.
  • — We don't present a single run as ground truth — confidence comes from repetition.
Run a free auditSee pricing