the ai citation gap: we scanned 51 businesses, most are invisible to ai.
everyone is writing think pieces about ai search. we wanted numbers instead, so on july 20 we ran a field scan of 55 real business websites and measured whether an ai system could actually read, extract, and cite them.
the ai citation gap is the distance between having a website and having a website ai systems can cite. in our july 2026 scan of 51 businesses (55 attempted, 51 fetched), only 5 of 51 (10%) addressed ai crawlers in robots.txt at all, and only 6 of 51 (12%) carried the faq or article schema that answer engines extract quotes from.
the median readiness score was 70 of 100, the mean 67. that sounds passable until you look at what's missing: the two checks that most directly decide whether chatgpt or perplexity can quote you are exactly the two that almost nobody passes. here is the full data.
what we scanned, and how
we scanned 55 public business websites across five industries: creator-economy platforms, ai-visibility and seo tools, micro-saas, course and community platforms, and content marketing agencies. 51 fetched successfully; 4 failed to respond to a standard fetch and are excluded from every stat below. for each site we ran the same check families our free ai visibility scan runs:
- ai crawler access: does robots.txt name GPTBot, ClaudeBot, PerplexityBot, or Google-Extended, and does it allow or block them?
- llms.txt: does the site publish a machine-readable context file?
- structured data: is there any JSON-LD, an Organization block, and, critically, FAQ or Article schema?
- extractability basics: a clean title under 70 characters, a meta description, exactly one H1.
each site got a weighted score out of 100. single scan, homepage only, one point in time (july 20, 2026). this measures citation *readiness*, the structural half of ai visibility, not domain authority or third-party presence. limits are limits, and we'd rather state them than inflate the claim.
the headline numbers
| check | pass rate | what it means |
|---|---|---|
| names any ai crawler in robots.txt | 5 of 51 (10%) | 90% have made no decision about ai at all |
| faq or article schema on homepage | 6 of 51 (12%) | almost nobody gives ai an answer it can extract |
| blocks an ai crawler | 1 of 51 (2%) | blocking is not the problem. indifference is |
| has llms.txt | 22 of 51 (43%) | the trendy file is spreading faster than the fundamentals |
| any JSON-LD structured data | 30 of 51 (59%) | 4 in 10 sites are unreadable as data |
| exactly one H1 | 34 of 51 (67%) | a third fail the most basic extractability check |
| scored under 50 of 100 | 12 of 51 (24%) | a quarter of businesses are structurally invisible |
sit with the first two rows. the fear narrative around ai search is "should we block the bots?" but in the field, one site in fifty-one blocks anything. the real story is that nine in ten businesses have never addressed ai crawlers at all, and nearly nine in ten give ai systems no extractable answer to quote. the machines are ready to cite. the websites are not ready to be cited.
llms.txt is winning the wrong race
43% of scanned sites ship an llms.txt file, while 12% ship faq or article schema. read that again: the optional context file with no confirmed ranking weight is being adopted almost four times as often as the schema that answer engines demonstrably extract from. that is what following hype instead of mechanics looks like: teams are adding the fashionable file and skipping the boring markup that actually moves citations.
scores by industry
| industry (n) | median score | llms.txt | faq/article schema |
|---|---|---|---|
| creator platforms (11) | 85 | 8 of 11 | 3 of 11 |
| marketing agencies (10) | 75 | 2 of 10 | 2 of 10 |
| ai + seo tools (10) | 70 | 3 of 10 | 0 of 10 |
| course platforms (9) | 60 | 5 of 9 | 0 of 9 |
| micro-saas (11) | 55 | 4 of 11 | 1 of 11 |
two findings worth naming. first, the top of the table: creator-economy platforms like collabstr, joinbrands, and grin scored 95, and beehiiv matched them, the strongest cluster in the study. the platforms that live on discovery invested in being discoverable. second, the industry selling ai visibility placed third. not one of the ten ai and seo tools we scanned carries faq or article schema on its homepage, and the single lowest fetched score in the entire study, 15 of 100, belongs to a vendor whose product is literally ai-visibility monitoring. the cobbler's children have no shoes, and in this market the cobbler is charging a subscription for them.
what this means if you run a business
the competitive read is simple: ai engines answer buyers' questions from the sites they can read and extract, and 88% of the market is not giving them anything to extract. that is not a threat, it is white space. the structural fixes are boring and fast: allow the ai crawlers explicitly, put a direct answer in your first paragraph, mark up your faqs, keep one h1. the full checklist is public. most of your competitors will not do it, which is exactly why it works.
the takeaway
the ai citation gap in one sentence: only 10% of businesses have addressed ai crawlers and only 12% give ai an extractable answer, so the ones that do both are competing for citations against almost nobody. the gap will close eventually. the businesses that close it first inherit the answers.
find out which side of the gap you're on
the same checks this study ran, pointed at your site: crawler access, extractable answers, structured data, and the exact gaps to close. free, no email to see your score.
run the free ai visibility scan