Originally published on angeo.dev. Full tables, p-values and the sealed plan are there.
Most claims about AI visibility are untestable by design: publish the signals, wait, attribute anything good that happens to the signals.
I wanted a version I could not fudge, so I wrote the analysis plan first, hashed it, and sent the hash to the other party before I had any data.
Do businesses AI assistants name repeatedly differ, on observable technical signals, from businesses the same assistants name once?
Every business in the corpus was named at least once, so this says nothing about how to enter an answer. It compares repeat against one-off mentions inside a named-business corpus.
| Signal | Check | |---|---| | Crawler access | Does robots.txt block any of 8 AI crawlers | | Content map | Does the site serve /llms.txt | | Structured data | Does a product page emit JSON-LD Product | | Buyability | Does that node carry offers.availability |
The answers came from a partner (connexion.me), who ran 44 product-level home-decor buying questions across ChatGPT, Gemini and Perplexity, twice, in two arms — 264 answers per arm.
Blinding was deliberate. I did not write the questions and did not see their store list until my plan was sealed; they never saw my frame, my scan results or my thresholds.
Cases: 3+ mentions across both runs and present in both. Controls: exactly one mention across both runs. Head excluded first — anything in 53+ of 264 answers (Amazon, Etsy, Wayfair, Target, Home Depot).
Sealed 10 August, SHA-256 9b4ccf12629e…: Under 15% of named businesses would be Magento No signal would separate the groups by more than 15 points Refutation condition: any signal differing by 20+ points with the named group higher
Two-sided Fisher exact. Percentages are of stores where the signal was observable.
| Signal | Cases | Controls | Diff | p | |---|---|---|---|---| | Blocks an AI crawler | 1/55 (2%) | 7/181 (4%) | −2.0 | .685 | | Serves llms.txt | 10/55 (18%) | 61/181 (34%) | −15.5 | .030 | | JSON-LD Product | 3/10 (30%) | 15/46 (33%) | −2.6 | 1.000 | | offers.availability | 3/10 (30%) | 13/46 (28%) | +1.7 | 1.000 |
The refutation condition was not met — in this arm, in the context arm, or pooled. No signal separates the repeatedly named group from the once-named group in the direction the field assumes.
llms.txt is more common among businesses named once than among those named repeatedly. Same direction in all five cuts, diffs of −15.3 to −17.4, nominal p between .012 and .030. Four signals across five cuts, no multiplicity correction — report them, do not treat them as confirmatory.
Before reading that as "llms.txt hurts you", I fetched the flagged files and read the first line of each:
113 of 159 llms.txt files — 71% — matched the same generated heading, differing only in the brand name. Not 113 independent decisions to publish. One generator.
Narrowed to the analysis universe: 146 of 455 serve the file (32%); 104 of those 146 are the template. Strip it and adoption is 9%, against 11% on my separately measured frame of 762 Magento stores.
I was about to publish 32% as adoption. Most of the apparent difference between the two populations was boilerplate.
