First-party measurement · September 12, 2026
How Often Two AI Answer Samples Disagreed
A reproducible analysis of 15 matched prompt-engine pairs showing answer, mention, competitor, and citation variation between two live samples.
Observed result
All 15 matched question-engine pairs produced different answer text. The mention verdict changed in 3 pairs, and the cited-domain set changed in 6 pairs.
Pair-level disagreement
| Compared field | Pairs differing | Pairs compared |
|---|---|---|
| Answer text | 15 | 15 |
| VisiScan mention label | 3 | 15 |
| Extracted competitor set | 2 | 15 |
| Cited-domain set | 6 | 15 |
Method
Each pair holds the question and engine constant and compares sample 0 with sample 1. Exact answer strings, boolean mention labels, normalized competitor sets and normalized cited-domain sets are compared independently.
The scan used five questions and three available engines on one date. Its OTHER industry classification produced a weak prompt panel, so these rates describe this run only.