What robots directives major websites actually use

Two censuses run on this site, both measured line by line rather than surveyed: 83 robots.txt files fetched and parsed with an RFC 9309 matcher (2026-08-24), and the same 83 homepages fetched for tag coverage (2026-09-06). Every figure below is a count over the measured hosts, stated as a fraction.

robots.txt, 83 of 93 hosts responding

Sitemap line present
62 of 83 (75%)
Wildcard * rules
4,218 rules across the set; 89 of the 83 hosts define a * group
Dollar-anchored rules
9,017 rules end in $, exact-match anchors
Crawl-delay set
15 of 83 (18%). Google ignores it; Bing honors it
Host over Google's 500 KiB limit
0. Largest file: figma.com at 372 KiB
Rules named after AI crawlers
39 of 83 hosts (47%) name at least one AI crawler

AI crawler rules, named hosts out of 83

CrawlerHostsShare
perplexity3643%
ccbot3036%
claudebot2935%
gptbot2935%
google-extended2530%
omgili2530%
meta-externalagent2429%
bytespider2328%
anthropic2227%
amazonbot2125%
applebot-extended2125%

What this means for you: if roughly half of major sites now write explicit AI-crawler rules, a site with no position at all is the outlier, not the default. The robots txt checker parses your file with the same matcher used here.

Garbage lines Google's parser still has to eat

Unknown directives are counted the way Google's own parser treats them: kept, attached to the current group, and ignored at match time. The census found 20 hosts serving crawl-delay inside googlebot groups where Google ignores it, 2 hosts writing a noindex line inside robots.txt (a directive that does not exist there), and 1 host serving an entire HTML document with inline JavaScript as its robots.txt. Two hosts wrote license metadata lines. A robots.txt that returns 200 with HTML in it is treated by crawlers as an empty file of unknown rules; it does not block anything.

Homepage tag coverage, 58 of 83 hosts measured

Any Open Graph tag
51 of 58 (88%)
All four core og tags (title, type, image, url)
40 of 58 (69%)
Any twitter card tag
43 of 58 (74%)
twitter:card value summary_large_image
25 of 43 measured card hosts
Any hreflang
23 of 58 (40%)
hreflang with x-default
14 of 23 (61%)
hreflang with self-reference
21 of 23 (91%)
Invalid or duplicated hreflang codes
2 hosts

What this means for you: Open Graph is near-universal at the top of the market but a third of measured hosts miss at least one core tag. Twitter card tags exist on 74% and no host with a twitter card lacks Open Graph entirely, which matches the practice of deriving Twitter tags from og: tags. The Open Graph checker and Twitter card validator test your pages against the same field list.

Homepage text, the 48 hosts with at least 200 words

Median word count
949 words
Median top-term density
3.86%
90th percentile density
7.82%
Maximum density
12.49% (cloudflare.com)
Hosts under 200 words served
10 of 58, homepages that render by script

Method and limits

robots.txt census: 93 well-known hosts fetched over HTTPS on 2026-08-24 with a 15 second timeout; 83 returned a body and every figure uses those 83. Parsing implements RFC 9309 and the behaviors in Google's robotstxt robots.cc, including sitemap lines not terminating a group. Homepage census: same host list fetched 2026-09-06, one GET each, served HTML only, nothing rendered by script; 58 returned 200 and only they are counted. Text figures cover the 48 hosts whose HTML held at least 200 words.

Limits: a fixed list of large hosts, not a random sample of the web. Homepages only, which are the least typical page of any site. hreflang checked by syntax, not reciprocity. Density figures say nothing about ranking.

Related pages