field notes · ai visibility · original data

the ai citation gap: we scanned 51 businesses, most are invisible to ai.

elisabeth hitz · july 21, 2026 · 8 min read · original research

everyone is writing think pieces about ai search. we wanted numbers instead, so on july 20 we ran a field scan of 55 real business websites and measured whether an ai system could actually read, extract, and cite them.

the ai citation gap is the distance between having a website and having a website ai systems can cite. in our july 2026 scan of 51 businesses (55 attempted, 51 fetched), only 5 of 51 (10%) addressed ai crawlers in robots.txt at all, and only 6 of 51 (12%) carried the faq or article schema that answer engines extract quotes from.

the median readiness score was 70 of 100, the mean 67. that sounds passable until you look at what's missing: the two checks that most directly decide whether chatgpt or perplexity can quote you are exactly the two that almost nobody passes. here is the full data.

what we scanned, and how

we scanned 55 public business websites across five industries: creator-economy platforms, ai-visibility and seo tools, micro-saas, course and community platforms, and content marketing agencies. 51 fetched successfully; 4 failed to respond to a standard fetch and are excluded from every stat below. for each site we ran the same check families our free ai visibility scan runs:

  • ai crawler access: does robots.txt name GPTBot, ClaudeBot, PerplexityBot, or Google-Extended, and does it allow or block them?
  • llms.txt: does the site publish a machine-readable context file?
  • structured data: is there any JSON-LD, an Organization block, and, critically, FAQ or Article schema?
  • extractability basics: a clean title under 70 characters, a meta description, exactly one H1.

each site got a weighted score out of 100. single scan, homepage only, one point in time (july 20, 2026). this measures citation *readiness*, the structural half of ai visibility, not domain authority or third-party presence. limits are limits, and we'd rather state them than inflate the claim.

the headline numbers

checkpass ratewhat it means
names any ai crawler in robots.txt5 of 51 (10%)90% have made no decision about ai at all
faq or article schema on homepage6 of 51 (12%)almost nobody gives ai an answer it can extract
blocks an ai crawler1 of 51 (2%)blocking is not the problem. indifference is
has llms.txt22 of 51 (43%)the trendy file is spreading faster than the fundamentals
any JSON-LD structured data30 of 51 (59%)4 in 10 sites are unreadable as data
exactly one H134 of 51 (67%)a third fail the most basic extractability check
scored under 50 of 10012 of 51 (24%)a quarter of businesses are structurally invisible

sit with the first two rows. the fear narrative around ai search is "should we block the bots?" but in the field, one site in fifty-one blocks anything. the real story is that nine in ten businesses have never addressed ai crawlers at all, and nearly nine in ten give ai systems no extractable answer to quote. the machines are ready to cite. the websites are not ready to be cited.

llms.txt is winning the wrong race

43% of scanned sites ship an llms.txt file, while 12% ship faq or article schema. read that again: the optional context file with no confirmed ranking weight is being adopted almost four times as often as the schema that answer engines demonstrably extract from. that is what following hype instead of mechanics looks like: teams are adding the fashionable file and skipping the boring markup that actually moves citations.

scores by industry

industry (n)median scorellms.txtfaq/article schema
creator platforms (11)858 of 113 of 11
marketing agencies (10)752 of 102 of 10
ai + seo tools (10)703 of 100 of 10
course platforms (9)605 of 90 of 9
micro-saas (11)554 of 111 of 11

two findings worth naming. first, the top of the table: creator-economy platforms like collabstr, joinbrands, and grin scored 95, and beehiiv matched them, the strongest cluster in the study. the platforms that live on discovery invested in being discoverable. second, the industry selling ai visibility placed third. not one of the ten ai and seo tools we scanned carries faq or article schema on its homepage, and the single lowest fetched score in the entire study, 15 of 100, belongs to a vendor whose product is literally ai-visibility monitoring. the cobbler's children have no shoes, and in this market the cobbler is charging a subscription for them.

what this means if you run a business

the competitive read is simple: ai engines answer buyers' questions from the sites they can read and extract, and 88% of the market is not giving them anything to extract. that is not a threat, it is white space. the structural fixes are boring and fast: allow the ai crawlers explicitly, put a direct answer in your first paragraph, mark up your faqs, keep one h1. the full checklist is public. most of your competitors will not do it, which is exactly why it works.

the takeaway

the ai citation gap in one sentence: only 10% of businesses have addressed ai crawlers and only 12% give ai an extractable answer, so the ones that do both are competing for citations against almost nobody. the gap will close eventually. the businesses that close it first inherit the answers.

find out which side of the gap you're on

the same checks this study ran, pointed at your site: crawler access, extractable answers, structured data, and the exact gaps to close. free, no email to see your score.

run the free ai visibility scan

original research by the closer method. method: automated field scan, july 20, 2026, of 55 public business homepages across five industries (51 fetched, 4 excluded as fetch failures); checks mirror the closer method ai visibility scan's check families (ai-crawler handling in robots.txt, llms.txt, JSON-LD, faq/article schema, title, meta description, h1), scored 0-100 weighted. raw per-site data retained on file. named scores are for fetched, publicly accessible homepages on the scan date; sites change, so treat every number as a snapshot. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork.