Is one run ever enough?
Only as a first look, and only if it is labeled as one run. It is enough to read what an assistant says about a brand and to find the questions worth measuring properly. It is not enough to state a share, a rank or a change.
Learn
One run of a prompt is a sample, not a measurement: the same question sent to the same model returns a different set of brands often enough that a single answer cannot be read as the model's opinion. How many runs a measurement needs depends on what it is trying to resolve, and the honest answer for most teams is "more than one, and fewer than you fear".
Across 4,500 calls, the same question sent to the same model returned the same set of brands in 15 percent of cases. The first-named brand held in 75 percent of runs. Read together: the leader is fairly stable, the rest of the list is not, and a single run says little about anyone below first place.
With 3 runs, a brand's appearance on a question reads as 0, 1, 2, 3 of 3. That is enough to separate "always named", "never named" and "sometimes named", which is the distinction that decides whether there is a problem. It is not enough to tell 40 percent from 60 percent, to rank brands below the leader, or to treat a change of one appearance between two measurements as movement. Those need many more runs per question, and for most teams that budget is better spent covering more questions.
The snapshot runs 3 times per provider at a low temperature (0.2) with a fixed JSON format. The buyer-question map runs once per question per provider at temperature 0 on every model since 2026-08-23, because it trades depth on one prompt for breadth across the buyer's questions, and every count is shown with the number of runs it came from. Re-runs are self-serve, so a team can repeat a measurement after publishing instead of waiting for a schedule. The methodology page has the models, parameters and limits.
Other people measured the same thing and reached the same place.
Measured
15 percent
of cases, across 4,500 calls, in which the same question to the same model returned the same set of brands; the first-named brand held in 75 percent of runs.
How it was measuredOnly as a first look, and only if it is labeled as one run. It is enough to read what an assistant says about a brand and to find the questions worth measuring properly. It is not enough to state a share, a rank or a change.
Cost, and the trade against breadth. Ten runs of one prompt resolve that prompt well and say nothing about the other eleven questions a buyer asks. The audit puts the budget on the map first and keeps re-runs cheap.
It reduces variation; it does not make the set of brands repeat every time, and the model behind an API can change between calls. Run counts are the honest answer; temperature is a setting.
Other definitions
The free audit reads your site, derives your buyer questions and measures them on three model APIs. No signup, no card.