Why you have to do this manually
Search engines report your rankings. Assistants report nothing. There is no console, no impression count, and no notification when a model starts recommending a competitor instead of you.
The only reliable method available is to ask the questions your buyers ask and record what comes back. It is unglamorous, it takes about an hour a month, and it is the single highest-value hour in most content programmes right now because almost nobody does it.
The five query shapes
Category discovery — "best X". Highest volume, most contested, and the one every competitor is already watching.
Segmented — "best X for Y". Lower volume, far higher conversion, and much easier to win because the answer depends on niche discussion rather than brand size.
Competitor alternatives — "alternatives to Z". Buyers with budget who are actively leaving somewhere.
Evaluation — "is Z worth it", "is X legit". These answers come overwhelmingly from community threads, which makes them the clearest signal of whether your off-domain presence is working.
Problem-first — "how do I stop...". The buyer does not know your category exists yet. Consistently the most neglected and the least contested.
How to run the test properly
Use a fresh session with no memory or personalisation for each run, otherwise you are measuring your own history rather than the model's default answer. Run each question two or three times, because outputs vary meaningfully between runs and a single result tells you very little.
Record four columns: whether you were named, which competitors were named, the order they appeared in, and — most importantly — which sources were cited. Run the same set monthly. Weekly produces noise.
Turning the results into work
The source column is where the value is. It converts "we need better AI visibility" into a concrete list, and that list almost always sorts into three piles.
Pages you control that are cited, which you make more quotable. Third-party editorial that is cited and omits you, where getting included is a bounded outreach task. And community threads, which are usually the largest pile for segmented and evaluation questions.
That last pile is the one with no shortcut. You cannot edit those threads, and replying to old ones achieves little. The only lever is being present when new ones appear — which is a monitoring problem rather than a content problem.
FREQUENTLY ASKED QUESTIONS
Questions about measuring AI visibility
Does this tool query ChatGPT for me?
No. It builds the question set; you run it. Querying assistants server-side would require API keys and would return different results than the consumer products your buyers actually use, which would make the measurement misleading.
How often should I run the set?
Monthly. Run-to-run variance in model outputs is large enough that weekly measurement mostly captures noise rather than change.
What counts as a good result?
There is no universal benchmark. Track your mention rate across the set over time, and your share relative to the competitors named on the same questions.
Why include questions about my competitor?
Alternatives and evaluation queries are where buyers with budget actually are, and being absent from the answer to "alternatives to X" is a far more expensive gap than missing a generic category query.
WORK THE LARGEST SOURCE PILE
Community threads feed the answers. Be in them.
MentionSpot monitors Reddit and X for the conversations that AI assistants cite in your category, scored by buying intent.
Get the 7-day passCONTINUE READING