a funded b2b saas company told AI not to read their website.
a teardown, with the scan attached. verified against the live site before publishing.
a funded b2b saas company blocked claude-web and ccbot in robots.txt while running a content marketing programme. the effect: AI answer engines cannot read their site to cite them. the check takes thirty seconds: open yourdomain.com/robots.txt and look for user-agent lines naming AI crawlers. it is the most common and most expensive silent failure in b2b ai visibility, and it is free to fix.
the finding
i ran a mechanical ai visibility scan across a set of funded b2b saas sites in july 2026. one of them, a well-funded revenue technology company, was blocking claude-web and ccbot in robots.txt. same site: zero relevant structured data, two h1 tags on a single page.
they publish content constantly. they have a content team. and they had instructed a significant share of the ai crawler population not to look. i am not naming them to dunk on them. i am naming the pattern, because it is everywhere and nobody is checking.
why this is worse than it sounds
when a buyer asks an engine for a recommendation, the engine returns three to five names with reasons attached. not ten links, no page two. to be one of them you need to be retrievable, describable, and corroborated. blocking a crawler kills the first, and the other two never get a chance.
the compounding part: engine answers draw on both live lookups and the training corpus. a block does not just cost you today's answer. it costs you presence in the material that shapes tomorrow's answers too.
the thirty second check
open yourdomain.com/robots.txt and look for User-agent: lines naming any of: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-Web, anthropic-ai, PerplexityBot, CCBot, Google-Extended. then look at what follows each. Disallow: / under any of them means that engine has been told to stay out.
a note on why this happens: most of these blocks were not decisions. they were a default in a security template, a wordpress plugin, a cloudflare bot-fight setting, or a 2023 policy about training data nobody revisited when ai search became a buying channel. the person who set it is often no longer at the company. that is why nobody catches it. there was never a meeting.
the other two failures on that same site
zero relevant structured data. schema is how you hand a machine your facts without asking it to infer them. no schema means every fact about the company has to be guessed from prose.
two h1 tags on one page. a small thing that signals the larger thing: nobody has looked at this page as a document rather than a design.
none of these are hard. all three together are the difference between being retrievable and being invisible, and the fix is measured in hours.
check yourself before you enjoy this too much
i ran the same scan on my own site. nine numeric claims, three external sources: i failed my own citable-authority check at the time. a normal miss rather than an embarrassment, and i am publishing it because a teardown from someone who has not run the check on themselves is not a teardown, it is marketing. the scan is the same for everyone. that is the point of having one.
what to do this week
1. read your robots.txt. thirty seconds. if an ai crawler is disallowed, decide deliberately whether you meant it.
2. check whether your site renders without javascript. content loaded entirely through javascript often means agents see fragments. server-side rendering or static html fixes it.
3. count your h1s. one per page.
4. add organization schema. name, url, description, sameAs. an afternoon.
5. then measure. the four above make you readable. they do not tell you whether any engine actually names you when a buyer asks. that needs a run, not a checklist. in a july 2026 censuswide survey, 84% of business leaders were not systematically tracking ai search performance and 53% had no ai visibility strategy at all, so measuring puts you ahead of most of your category, including the funded ones.
the fixed-scope version
the audit runs all of this against your real competitors.
mechanical checks plus buyer-intent query runs across five engines against your three named competitors, delivered as a dated report with the open method attached so your team can re-run it.
see the audit →more field notes on x · instagram · tiktok
written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork.
scan run july 2026 using the open method at CVI v1.0. survey figures: democracy pr / censuswide, "the power of i", 1 july 2026. ai crawler directives change; re-run before relying on any finding here. related: why chatgpt doesn't recommend your business