Original research · 2026-09-20

We measured 100 SaaS sites for AI crawler access.

Almost nobody blocks AI crawlers on purpose: only 2 of the 90 sites that answered disallow a retrieval agent. The problems are quieter. 76.7% name no AI crawler at all, 34.4% publish no Organization entity, and 8.9% serve under 200 words before JavaScript runs. Ten refused this crawler outright.

10 of 100

refused an ordinary crawler outright

Ten sites answered 403, 429 or 400 to a plain, identified, rate-limited request. Bot protection does not distinguish between a research crawler and an answer engine as reliably as its buyers assume, and the same defences that stop scrapers can stop citations.

76.7%

name no AI crawler in robots.txt at all

69 of 90 reachable sites leave every AI agent to fall through to a wildcard rule. That usually works, and it means nobody made a decision. Several crawlers read only their own group when one exists, so a file with both a named group and a wildcard can behave differently from how it reads.

57.8%

publish an llms.txt

52 of 90. Higher than expected, and the clearest sign in this dataset that AEO has moved from an idea to a default for well-resourced marketing teams. No engine has committed to reading the file, which makes the adoption rate a statement about where these teams think search is going.

34.4%

have no Organization entity on the homepage

31 of 90 publish no Organization, LocalBusiness or Corporation markup anywhere on their front page. A model resolving a brand name has nothing authoritative to resolve it against, which is how a confident answer ends up describing a different company.

8.9%

serve under 200 words before JavaScript runs

8 of 90, and 4 serve under fifty. To a crawler that executes no client-side code, and none of the major retrieval crawlers do, those homepages are close to blank. The median across the sample is 910 words.

2.2%

deliberately block a retrieval crawler

Only two of the reachable sites disallow one of the six agents from the root. Outright blocking is rare and mostly confined to companies whose product is the content itself. The invisibility problem in this sample is not refusal, it is everything above.

The sample, stated before the conclusions

A convenience sample of 100 well-known SaaS, developer-tool and marketing-tool companies. Not a random sample and not a ranking. These are companies whose marketing sites are public, well-resourced and widely imitated, which makes them the interesting case: if sites like these get AI crawler access wrong, the long tail is worse rather than better.

Of 100 domains, 90 answered and 10 did not. Every percentage on this page divides by 90, and that is worth repeating because a rate over an unstated denominator is the single most common way a study of this kind misleads.

Method

  • Two requests per domain, both public: robots.txt and the homepage. Nothing else was crawled, nothing was stored about anyone, and there was a delay between sites.
  • Crawler access is read from robots.txt for the six retrieval agents that decide whether a brand can appear in an AI answer today: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot and Google-Extended. Training-only crawlers are excluded, because refusing those is a defensible business decision while refusing these makes you invisible.
  • Everything else is read from the HTML the server returned, before any JavaScript ran, because that is what a retrieval crawler receives.
  • An llms.txt counts only when the response is 200, served as text/plain, and begins with a Markdown heading. A first pass without those conditions reported 70%; the real figure is 57.8%.

The fourth point is the one worth dwelling on. A first pass of this study reported that 70% of the sample published an llms.txt, which would have been a striking headline and a wrong one: plenty of sites answer 200 with an HTML shell at any path, and one answered 403 to a browser and 200 to this client within the same minute. Requiring a text/plain response and a Markdown heading moved the real figure to 57.8%. A false positive is more expensive than a miss, and in published research it is fatal.

Who actually blocks a retrieval crawler

3 sites in the sample disallow at least one of the six agents from the root. The pattern is consistent: companies whose product is the content or the canvas itself.

SiteAgents disallowed from /
figma.comGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended
canva.comGPTBot, ClaudeBot
loom.comGPTBot

The homepages a crawler sees as almost blank

None of the major AI retrieval crawlers executes JavaScript. These are the word counts in the HTML the server actually returned, which is the whole of what those crawlers receive. The median across the sample is 910 words.

SiteWords servedJSON-LD blocks
jfrog.com00
render.com133
snowflake.com280
recurly.com434
notion.so1300
vercel.com1403
trello.com1700
1password.com1990

A low number here is not proof that a company is invisible in AI answers. It means the homepage carries almost nothing an answer engine can quote, and that whatever authority the brand has is being earned somewhere other than its own front page. The measurement behind the JavaScript claim is a separate page, with the sources.

The sites that refused

10 of 100 answered 403, 429 or 400 to a plain, identified, rate-limited request at a delay of well under one request per second. They are listed because excluding them silently would change every percentage on this page without saying so.

SiteStatus
canva.com403
iterable.com403
make.com403
github.com400
docker.com403
hashicorp.com429
fivetran.comno response
contentful.com429
convertkit.com403
thinkific.com403

Every measurement, per site

All 100 domains, in the order they were read. Anyone can reproduce a row with two requests.

SiteWordsJSON-LDOrg entityllms.txtNames agentsBlocks
stripe.com16363YesYes0none
notion.so1300NoNo0none
figma.com2421YesNo66
linear.app15610NoYes0none
vercel.com1403YesYes0none
netlify.com12686YesYes6none
supabase.com11664YesYes4none
railway.app9874YesYes0none
render.com133YesNo6none
fly.io7737YesYes0none
planetscale.com18201YesYes0none
cloudflare.com10023YesYes4none
digitalocean.com8120NoNo0none
heroku.com11396YesNo0none
twilio.com11183YesNo0none
sendgrid.com2780NoNo0none
mailchimp.com12971NoYes0none
hubspot.com8525YesYes0none
salesforce.com8921YesNo0none
zendesk.com18806YesYes0none
intercom.com11209YesYes0none
drift.com4921NoNo0none
segment.com2970NoNo0none
amplitude.com5961NoYes6none
mixpanel.com12030NoNo0none
heap.io4540NoNo0none
hotjar.com6600NoNo0none
optimizely.com5413YesYes0none
semrush.com8911YesYes0none
ahrefs.com15421YesNo0none
moz.com15045YesNo1none
similarweb.com11713YesYes0none
sproutsocial.com6542YesYes0none
buffer.com7031YesYes0none
hootsuite.com13640NoYes6none
later.com8821YesNo0none
canva.com403NoNo32
webflow.com23145YesYes0none
squarespace.com23995YesYes3none
wix.com23775YesYes0none
shopify.com5691YesYes0none
bigcommerce.com8543YesYes0none
woocommerce.com5727YesNo0none
klaviyo.com10776YesYes5none
attentive.com7595YesNo0none
braze.com11781NoYes0none
iterable.com403NoNo0none
customer.io5961YesYes0none
airtable.com8190NoNo0none
asana.com10364YesYes0none
monday.com21965YesYes5none
clickup.com18351YesYes0none
basecamp.com8003NoNo0none
trello.com1700NoNo0none
miro.com16890NoNo0none
loom.com3451YesYes11
calendly.com13323YesYes0none
zapier.com21656YesYes6none
make.com403NoYes0none
retool.com4531YesYes5none
posthog.com11192YesYes0none
sentry.io24782YesYes0none
datadoghq.com2415YesYes0none
newrelic.com13273YesYes0none
pagerduty.com14154YesNo0none
gitlab.com6115YesNo0none
github.com400NoNo0none
bitbucket.org14301NoNo0none
circleci.com19971YesYes1none
jfrog.com00NoNo0none
docker.com403NoNo0none
hashicorp.com429NoNo0none
mongodb.com9102YesYes0none
elastic.co4821NoNo0none
snowflake.com280NoNo0none
databricks.com9450NoYes0none
dbt.com6880NoYes0none
fivetran.comn/aNoYes0none
confluent.io15824YesNo3none
auth0.com18496YesYes0none
okta.com7891YesYes0none
1password.com1990NoNo0none
lastpass.com19821NoNo6none
cloudinary.com12166YesYes0none
contentful.com429NoNo0none
sanity.io12740NoNo0none
storyblok.com8254YesYes0none
prismic.io16187YesYes0none
strapi.io6065YesNo0none
ghost.org6281NoNo0none
substack.com2490NoNo0none
beehiiv.com9056YesYes1none
convertkit.com403NoNo0none
teachable.com14984YesNo5none
thinkific.com403NoYes0none
gumroad.com6270NoYes0none
lemonsqueezy.com14390NoNo0none
paddle.com3293YesNo0none
chargebee.com179812NoYes6none
recurly.com434YesYes6none

Questions

How was this measured?

Two public requests per domain: robots.txt and the homepage. Crawler access is read from robots.txt for the six retrieval agents that decide whether a brand appears in an AI answer. Everything else is read from the HTML the server returned before any JavaScript ran, because that is what a retrieval crawler receives.

Is this a random sample?

No, and it is not a ranking either. It is a convenience sample of well-known SaaS, developer-tool and marketing-tool companies. The argument for it is that these are well-resourced sites whose patterns get copied, so problems here are a floor rather than a ceiling.

Why does the percentage not divide by 100?

Ten of the hundred sampled sites refused the crawler outright with a 403, 429 or 400. Every rate on this page divides by the 90 that answered, and that number is stated next to each one.

Can I reproduce this?

Yes, and that is the point. The script is in the repository, the sample list is a plain text file next to it, and every measurement is two public requests anyone can make from a browser or a terminal.

Why can you afford to publish this when nobody else does?

Because the audit uses no language model and no data vendor, so a run costs bandwidth and nothing else. Every competitor measuring the same thing pays a model call or a data credit per site. That is an architecture difference before it is a marketing one.

Run the same measurements on your own site

The crawler access check and the extractability check are free, need no account, and use the same engine that produced this dataset.