Visora
Sign up
← Back to blog

· Visora

GEOAuditMeasurement

How many prompts do you need to test AI citation coverage properly?

Short answer: aim for roughly 30 to 50 prompts before you believe a coverage number. Below that, the difference between "we are cited on 40% of questions" and "we are cited on 55% of questions" is mostly sampling noise, not a change in your citation position. Above roughly 100, you are usually buying precision you will not act on differently. The number matters less than the shape of the set: prompts should be spread across the jobs your buyers are actually trying to finish, not concentrated around your top product name.

## Why one or two prompts is worse than no measurement

Most stores that "check AI visibility" run a handful of prompts: the brand name, the category term, maybe one product-specific question. That produces a number, and a number feels like evidence. It usually is not.

A citation is per-prompt, not per-store. An assistant answering "best lightweight travel backpack under $100 for a 15-inch laptop" and an assistant answering "travel backpack brands with repair programmes" are running on different retrieval, different page evidence, and often different cited sources from your own catalogue. Measuring one and generalising to the other is the classic mistake.

If you test five prompts and get two citations, you do not know whether your coverage is 40% — you know that two specific questions went your way. The right reading is: coverage is somewhere in a wide band, and the band is too wide to prioritise fixes from.

## What a defensible prompt set looks like

Treat the prompt set as a survey instrument, not a checklist. Four things decide whether it holds up.

1. Spread across intent types. Buying-intent prompts ("best X for Y"), comparison prompts ("X vs Z for Y"), constraint and filter prompts ("X under $N with next-day delivery"), and factual prompts about a specific product all pull on different page evidence. A set made only of "best X" prompts will over-report citation for stores with strong category pages and miss everything else. 2. Spread across your actual catalogue. If 70% of your SKUs sit in accessories and you test only your hero product, you are measuring your hero product. 3. Include prompts you expect to lose. Control prompts are what make a coverage number meaningful. A set of 30 questions you chose because you already rank well on them is not a measurement; it is a victory lap. Aim for a set where you would honestly guess the win rate is between 30% and 45%. 4. Write them the way buyers write them, not the way merchants write them. Buyers write constraints, mixed languages, and half-formed needs. Merchants write keywords. Those are not the same distribution, and prompt-shaped queries are exactly where assistants now do most shopping work.

A practical starting shape: 8-12 buying-intent prompts, 6-10 comparison prompts, 6-10 constraint-and-filter prompts, 5-8 product-fact prompts, and 5 control prompts you expect to lose. That lands you near 30-50 without trying to.

## Sizing the set to your catalogue, not your ambition

A common failure is testing a 40-prompt set against a 900-SKU catalogue and concluding coverage is fine. Coverage measured at the prompt level and coverage measured at the catalogue level are different quantities.

If your prompt set is representative, prompt-level and catalogue-level coverage should roughly track each other. When they diverge sharply — high prompt coverage, low catalogue coverage — the usual cause is that your prompts cluster on a small, well-optimised subset of products. The fix is not more prompts. It is a prompt set that samples products the way a buyer would: weighted toward what people actually ask about, not toward what you have already fixed.

This is the same discipline as the coverage-not-pages framing: you are trying to find the invisible shelf, and a narrow prompt set is precisely the instrument that hides it. [Run a scan on /audit](/audit) to get the page-level and catalogue-level picture, then use prompts to test whether the fixes changed what assistants actually say.

## How to avoid fooling yourself

  • Do not compare a 10-prompt run to a 40-prompt run. Different instrument, different number. Coverage only compares against a fixed set.
  • Record the exact prompt text. Paraphrases are different prompts. If your set drifts week to week, so does your trend line.
  • Test across at least two assistants. Citation behaviour differs enough that one assistant is a sample, not a trend.
  • Re-run the same set before changing it. Improve the set when you have a reason, not because the last number was unflattering.
  • Treat a single-run change under about 10 points as unconfirmed until a second run reproduces it. Prompt sets of this size are not precise to the point.

## Where Visora fits

Visora is built around exactly this problem: instead of testing a handful of prompts and guessing, it scans your catalogue for the structural reasons an assistant cannot cite a page — missing or inconsistent structured data, facts that only exist in images, spec information that is present but not extractable — and reports them at catalogue scale. You then use a fixed prompt set to check whether the assistant's answers moved. Diagnostics and prompts answer different questions, and you want both. The [FAQ](/faq) covers what the scan does and does not measure.

## FAQ

Isn't 30-50 prompts a lot of manual work?

It is the first time. Build the set once, keep it in a spreadsheet, and re-run it. The set is an asset; reusing it is what makes week-over-week numbers comparable.

Can I just use a keyword tool to generate prompts?

Keyword tools give you search-engine demand. Prompt-shaped buying questions are related but not identical, and they skew longer, more constraint-heavy, and more conversational. Use keywords as a starting point, then rewrite them as questions a person would actually say out loud.

Does high prompt coverage mean I will get AI traffic?

No. Coverage tells you whether you are eligible to be cited. Whether an assistant actually picks you when you are eligible is a separate question, influenced by competition and by how decisively your page states facts. Coverage is the part you can verify and fix.

How often should I re-run the set?

Monthly is enough for a trend. Re-run sooner only when you have shipped a specific fix and want to know whether it landed.

Put this into practice

Audit your PDP or category page with Visora, then fix schema and FAQ gaps that block AI citations.

Run a free GEO audit →

https://geovisora.com/en/blog/how-many-prompts-to-test-ai-citation-coverage-2026