We measured 100 SaaS sites for AI crawler access.
Almost nobody blocks AI crawlers on purpose: only 2 of the 90 sites that answered disallow a retrieval agent. The problems are quieter. 76.7% name no AI crawler at all, 34.4% publish no Organization entity, and 8.9% serve under 200 words before JavaScript runs. Ten refused this crawler outright.
refused an ordinary crawler outright
Ten sites answered 403, 429 or 400 to a plain, identified, rate-limited request. Bot protection does not distinguish between a research crawler and an answer engine as reliably as its buyers assume, and the same defences that stop scrapers can stop citations.
name no AI crawler in robots.txt at all
69 of 90 reachable sites leave every AI agent to fall through to a wildcard rule. That usually works, and it means nobody made a decision. Several crawlers read only their own group when one exists, so a file with both a named group and a wildcard can behave differently from how it reads.
publish an llms.txt
52 of 90. Higher than expected, and the clearest sign in this dataset that AEO has moved from an idea to a default for well-resourced marketing teams. No engine has committed to reading the file, which makes the adoption rate a statement about where these teams think search is going.
have no Organization entity on the homepage
31 of 90 publish no Organization, LocalBusiness or Corporation markup anywhere on their front page. A model resolving a brand name has nothing authoritative to resolve it against, which is how a confident answer ends up describing a different company.
serve under 200 words before JavaScript runs
8 of 90, and 4 serve under fifty. To a crawler that executes no client-side code, and none of the major retrieval crawlers do, those homepages are close to blank. The median across the sample is 910 words.
deliberately block a retrieval crawler
Only two of the reachable sites disallow one of the six agents from the root. Outright blocking is rare and mostly confined to companies whose product is the content itself. The invisibility problem in this sample is not refusal, it is everything above.
The sample, stated before the conclusions
A convenience sample of 100 well-known SaaS, developer-tool and marketing-tool companies. Not a random sample and not a ranking. These are companies whose marketing sites are public, well-resourced and widely imitated, which makes them the interesting case: if sites like these get AI crawler access wrong, the long tail is worse rather than better.
Of 100 domains, 90 answered and 10 did not. Every percentage on this page divides by 90, and that is worth repeating because a rate over an unstated denominator is the single most common way a study of this kind misleads.
Method
- Two requests per domain, both public: robots.txt and the homepage. Nothing else was crawled, nothing was stored about anyone, and there was a delay between sites.
- Crawler access is read from robots.txt for the six retrieval agents that decide whether a brand can appear in an AI answer today: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot and Google-Extended. Training-only crawlers are excluded, because refusing those is a defensible business decision while refusing these makes you invisible.
- Everything else is read from the HTML the server returned, before any JavaScript ran, because that is what a retrieval crawler receives.
- An llms.txt counts only when the response is 200, served as text/plain, and begins with a Markdown heading. A first pass without those conditions reported 70%; the real figure is 57.8%.
The fourth point is the one worth dwelling on. A first pass of this study reported that 70% of the sample published an llms.txt, which would have been a striking headline and a wrong one: plenty of sites answer 200 with an HTML shell at any path, and one answered 403 to a browser and 200 to this client within the same minute. Requiring a text/plain response and a Markdown heading moved the real figure to 57.8%. A false positive is more expensive than a miss, and in published research it is fatal.
Who actually blocks a retrieval crawler
3 sites in the sample disallow at least one of the six agents from the root. The pattern is consistent: companies whose product is the content or the canvas itself.
| Site | Agents disallowed from / |
|---|---|
figma.com | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended |
canva.com | GPTBot, ClaudeBot |
loom.com | GPTBot |
The homepages a crawler sees as almost blank
None of the major AI retrieval crawlers executes JavaScript. These are the word counts in the HTML the server actually returned, which is the whole of what those crawlers receive. The median across the sample is 910 words.
| Site | Words served | JSON-LD blocks |
|---|---|---|
jfrog.com | 0 | 0 |
render.com | 13 | 3 |
snowflake.com | 28 | 0 |
recurly.com | 43 | 4 |
notion.so | 130 | 0 |
vercel.com | 140 | 3 |
trello.com | 170 | 0 |
1password.com | 199 | 0 |
A low number here is not proof that a company is invisible in AI answers. It means the homepage carries almost nothing an answer engine can quote, and that whatever authority the brand has is being earned somewhere other than its own front page. The measurement behind the JavaScript claim is a separate page, with the sources.
The sites that refused
10 of 100 answered 403, 429 or 400 to a plain, identified, rate-limited request at a delay of well under one request per second. They are listed because excluding them silently would change every percentage on this page without saying so.
| Site | Status |
|---|---|
canva.com | 403 |
iterable.com | 403 |
make.com | 403 |
github.com | 400 |
docker.com | 403 |
hashicorp.com | 429 |
fivetran.com | no response |
contentful.com | 429 |
convertkit.com | 403 |
thinkific.com | 403 |
Every measurement, per site
All 100 domains, in the order they were read. Anyone can reproduce a row with two requests.
| Site | Words | JSON-LD | Org entity | llms.txt | Names agents | Blocks |
|---|---|---|---|---|---|---|
stripe.com | 1636 | 3 | Yes | Yes | 0 | none |
notion.so | 130 | 0 | No | No | 0 | none |
figma.com | 242 | 1 | Yes | No | 6 | 6 |
linear.app | 1561 | 0 | No | Yes | 0 | none |
vercel.com | 140 | 3 | Yes | Yes | 0 | none |
netlify.com | 1268 | 6 | Yes | Yes | 6 | none |
supabase.com | 1166 | 4 | Yes | Yes | 4 | none |
railway.app | 987 | 4 | Yes | Yes | 0 | none |
render.com | 13 | 3 | Yes | No | 6 | none |
fly.io | 773 | 7 | Yes | Yes | 0 | none |
planetscale.com | 1820 | 1 | Yes | Yes | 0 | none |
cloudflare.com | 1002 | 3 | Yes | Yes | 4 | none |
digitalocean.com | 812 | 0 | No | No | 0 | none |
heroku.com | 1139 | 6 | Yes | No | 0 | none |
twilio.com | 1118 | 3 | Yes | No | 0 | none |
sendgrid.com | 278 | 0 | No | No | 0 | none |
mailchimp.com | 1297 | 1 | No | Yes | 0 | none |
hubspot.com | 852 | 5 | Yes | Yes | 0 | none |
salesforce.com | 892 | 1 | Yes | No | 0 | none |
zendesk.com | 1880 | 6 | Yes | Yes | 0 | none |
intercom.com | 1120 | 9 | Yes | Yes | 0 | none |
drift.com | 492 | 1 | No | No | 0 | none |
segment.com | 297 | 0 | No | No | 0 | none |
amplitude.com | 596 | 1 | No | Yes | 6 | none |
mixpanel.com | 1203 | 0 | No | No | 0 | none |
heap.io | 454 | 0 | No | No | 0 | none |
hotjar.com | 660 | 0 | No | No | 0 | none |
optimizely.com | 541 | 3 | Yes | Yes | 0 | none |
semrush.com | 891 | 1 | Yes | Yes | 0 | none |
ahrefs.com | 1542 | 1 | Yes | No | 0 | none |
moz.com | 1504 | 5 | Yes | No | 1 | none |
similarweb.com | 1171 | 3 | Yes | Yes | 0 | none |
sproutsocial.com | 654 | 2 | Yes | Yes | 0 | none |
buffer.com | 703 | 1 | Yes | Yes | 0 | none |
hootsuite.com | 1364 | 0 | No | Yes | 6 | none |
later.com | 882 | 1 | Yes | No | 0 | none |
canva.com | 403 | No | No | 3 | 2 | |
webflow.com | 2314 | 5 | Yes | Yes | 0 | none |
squarespace.com | 2399 | 5 | Yes | Yes | 3 | none |
wix.com | 2377 | 5 | Yes | Yes | 0 | none |
shopify.com | 569 | 1 | Yes | Yes | 0 | none |
bigcommerce.com | 854 | 3 | Yes | Yes | 0 | none |
woocommerce.com | 572 | 7 | Yes | No | 0 | none |
klaviyo.com | 1077 | 6 | Yes | Yes | 5 | none |
attentive.com | 759 | 5 | Yes | No | 0 | none |
braze.com | 1178 | 1 | No | Yes | 0 | none |
iterable.com | 403 | No | No | 0 | none | |
customer.io | 596 | 1 | Yes | Yes | 0 | none |
airtable.com | 819 | 0 | No | No | 0 | none |
asana.com | 1036 | 4 | Yes | Yes | 0 | none |
monday.com | 2196 | 5 | Yes | Yes | 5 | none |
clickup.com | 1835 | 1 | Yes | Yes | 0 | none |
basecamp.com | 800 | 3 | No | No | 0 | none |
trello.com | 170 | 0 | No | No | 0 | none |
miro.com | 1689 | 0 | No | No | 0 | none |
loom.com | 345 | 1 | Yes | Yes | 1 | 1 |
calendly.com | 1332 | 3 | Yes | Yes | 0 | none |
zapier.com | 2165 | 6 | Yes | Yes | 6 | none |
make.com | 403 | No | Yes | 0 | none | |
retool.com | 453 | 1 | Yes | Yes | 5 | none |
posthog.com | 1119 | 2 | Yes | Yes | 0 | none |
sentry.io | 2478 | 2 | Yes | Yes | 0 | none |
datadoghq.com | 241 | 5 | Yes | Yes | 0 | none |
newrelic.com | 1327 | 3 | Yes | Yes | 0 | none |
pagerduty.com | 1415 | 4 | Yes | No | 0 | none |
gitlab.com | 611 | 5 | Yes | No | 0 | none |
github.com | 400 | No | No | 0 | none | |
bitbucket.org | 1430 | 1 | No | No | 0 | none |
circleci.com | 1997 | 1 | Yes | Yes | 1 | none |
jfrog.com | 0 | 0 | No | No | 0 | none |
docker.com | 403 | No | No | 0 | none | |
hashicorp.com | 429 | No | No | 0 | none | |
mongodb.com | 910 | 2 | Yes | Yes | 0 | none |
elastic.co | 482 | 1 | No | No | 0 | none |
snowflake.com | 28 | 0 | No | No | 0 | none |
databricks.com | 945 | 0 | No | Yes | 0 | none |
dbt.com | 688 | 0 | No | Yes | 0 | none |
fivetran.com | n/a | No | Yes | 0 | none | |
confluent.io | 1582 | 4 | Yes | No | 3 | none |
auth0.com | 1849 | 6 | Yes | Yes | 0 | none |
okta.com | 789 | 1 | Yes | Yes | 0 | none |
1password.com | 199 | 0 | No | No | 0 | none |
lastpass.com | 1982 | 1 | No | No | 6 | none |
cloudinary.com | 1216 | 6 | Yes | Yes | 0 | none |
contentful.com | 429 | No | No | 0 | none | |
sanity.io | 1274 | 0 | No | No | 0 | none |
storyblok.com | 825 | 4 | Yes | Yes | 0 | none |
prismic.io | 1618 | 7 | Yes | Yes | 0 | none |
strapi.io | 606 | 5 | Yes | No | 0 | none |
ghost.org | 628 | 1 | No | No | 0 | none |
substack.com | 249 | 0 | No | No | 0 | none |
beehiiv.com | 905 | 6 | Yes | Yes | 1 | none |
convertkit.com | 403 | No | No | 0 | none | |
teachable.com | 1498 | 4 | Yes | No | 5 | none |
thinkific.com | 403 | No | Yes | 0 | none | |
gumroad.com | 627 | 0 | No | Yes | 0 | none |
lemonsqueezy.com | 1439 | 0 | No | No | 0 | none |
paddle.com | 329 | 3 | Yes | No | 0 | none |
chargebee.com | 1798 | 12 | No | Yes | 6 | none |
recurly.com | 43 | 4 | Yes | Yes | 6 | none |
Questions
How was this measured?
Two public requests per domain: robots.txt and the homepage. Crawler access is read from robots.txt for the six retrieval agents that decide whether a brand appears in an AI answer. Everything else is read from the HTML the server returned before any JavaScript ran, because that is what a retrieval crawler receives.
Is this a random sample?
No, and it is not a ranking either. It is a convenience sample of well-known SaaS, developer-tool and marketing-tool companies. The argument for it is that these are well-resourced sites whose patterns get copied, so problems here are a floor rather than a ceiling.
Why does the percentage not divide by 100?
Ten of the hundred sampled sites refused the crawler outright with a 403, 429 or 400. Every rate on this page divides by the 90 that answered, and that number is stated next to each one.
Can I reproduce this?
Yes, and that is the point. The script is in the repository, the sample list is a plain text file next to it, and every measurement is two public requests anyone can make from a browser or a terminal.
Why can you afford to publish this when nobody else does?
Because the audit uses no language model and no data vendor, so a run costs bandwidth and nothing else. Every competitor measuring the same thing pays a model call or a data credit per site. That is an architecture difference before it is a marketing one.
Run the same measurements on your own site
The crawler access check and the extractability check are free, need no account, and use the same engine that produced this dataset.