July 20, 2026
How to measure your AI visibility (before you try to improve it)
You can't fix what you can't see. Here's a repeatable way to measure whether ChatGPT, Claude, Gemini, and Perplexity recommend your product, using prompts you can run today.
By Nahuel Soria
Every team that asks "how do we show up in ChatGPT?" skips the first step: measuring where they stand right now. Without a baseline, you can't tell whether a change helped, hurt, or did nothing.
The problem is that AI answers feel unmeasurable. There's no rank tracker, no impressions chart, no "position 4." But you can measure AI visibility. You just have to do it the way the assistant does it: by asking questions and reading the answers.
Here's a method you can run this afternoon.
Start with the questions a buyer actually asks
Don't test vanity prompts like "is [your brand] good?" An assistant will almost always say something positive when you name a product directly. That tells you nothing.
Test the questions a buyer asks before they know you exist:
- "What's the best tool for [the job you do]?"
- "Alternatives to [your biggest competitor]"
- "[Competitor A] vs [Competitor B] for [use case]"
- "What should I use to [specific outcome] for a small team?"
- "Recommend a [category] tool with [key requirement]"
Write down 10 to 20 of these. This list is your test suite, and it matters more than any single answer. If you only remember one thing from this post: measure the queries where you are not mentioned by name, because those are the ones that decide whether a stranger ever finds you.
Run each prompt across the assistants that matter
Buyers don't all use the same assistant, so a single tool isn't a measurement. Run every prompt across the ones your market actually uses:
- ChatGPT (the default for most)
- Claude
- Gemini (increasingly the answer behind Google)
- Perplexity (search-first, cites sources inline)
Use a fresh chat for each prompt so previous messages don't bias the answer. Assistants also vary run to run, so ask the same prompt two or three times. You're looking for a pattern, not a single verdict.
Score what you see, consistently
For each prompt and assistant, record three things:
- Mentioned? Were you named at all in the answer?
- Position. Were you the first recommendation, one of several, or a footnote?
- Framing. Was the description accurate, or did the assistant get your pricing, features, or category wrong?
That third column is the one teams overlook. Being mentioned with the wrong price or a missing feature can cost you the deal as surely as not being mentioned at all. When an assistant says something false about you, that's a content gap on your own site or in the third-party sources it read.
Turn the results into a simple grid: prompts down the side, assistants across the top, a mark in each cell. Now you have a baseline. A picture of exactly where you're invisible, where you're misrepresented, and where you already win.
Read the sources, not just the answer
Perplexity and, increasingly, the others show the pages they drew from. Read them. Those citations tell you why the answer looks the way it does:
- If competitors' comparison pages show up and yours don't, that's your next page to write.
- If a directory or roundup keeps getting cited and you're not on it, that's your next listing.
- If a Reddit thread is shaping the answer, that's the conversation you're absent from.
The sources are the instruction manual for what to fix. An answer tells you the score; the citations tell you the play.
Turn the baseline into a short list
Once the grid is filled in, the work sorts itself. Rank your gaps by intent:
- High-intent prompts where you're invisible (comparisons, "best tool for X", alternatives). Fix these first, they're closest to a purchase.
- Prompts where you're mentioned but misrepresented. Correct the record on your own pages so assistants have accurate material to pull from.
- Category and informational prompts where you're absent. Broader reach, lower urgency.
Re-measure on a schedule
A baseline is only useful if you check it again. Re-run the same test suite every few weeks, especially after you ship a comparison page, land a directory listing, or earn a review. That's how you learn what actually moved the needle instead of guessing.
This is the whole loop: define the prompts, run them across assistants, score mention and accuracy, read the citations, fix the highest-intent gaps, and re-measure. It's manual but honest, and it's the same thing an AI visibility audit automates so you don't have to run dozens of prompts by hand every month.
You can't improve what you haven't measured. Start with the baseline.