What robots directives major websites actually use
Two censuses run on this site, both measured line by line rather than surveyed: 83 robots.txt files fetched and parsed with an RFC 9309 matcher (2026-08-24), and the same 83 homepages fetched for tag coverage (2026-09-06). Every figure below is a count over the measured hosts, stated as a fraction.
robots.txt, 83 of 93 hosts responding
- Sitemap line present
- 62 of 83 (75%)
- Wildcard
*rules - 4,218 rules across the set; 89 of the 83 hosts define a
*group - Dollar-anchored rules
- 9,017 rules end in
$, exact-match anchors - Crawl-delay set
- 15 of 83 (18%). Google ignores it; Bing honors it
- Host over Google's 500 KiB limit
- 0. Largest file: figma.com at 372 KiB
- Rules named after AI crawlers
- 39 of 83 hosts (47%) name at least one AI crawler
AI crawler rules, named hosts out of 83
| Crawler | Hosts | Share |
|---|---|---|
| perplexity | 36 | 43% |
| ccbot | 30 | 36% |
| claudebot | 29 | 35% |
| gptbot | 29 | 35% |
| google-extended | 25 | 30% |
| omgili | 25 | 30% |
| meta-externalagent | 24 | 29% |
| bytespider | 23 | 28% |
| anthropic | 22 | 27% |
| amazonbot | 21 | 25% |
| applebot-extended | 21 | 25% |
What this means for you: if roughly half of major sites now write explicit AI-crawler rules, a site with no position at all is the outlier, not the default. The robots txt checker parses your file with the same matcher used here.
Garbage lines Google's parser still has to eat
Unknown directives are counted the way Google's own parser treats them: kept, attached to the current group, and ignored at match time. The census found 20 hosts serving crawl-delay inside googlebot groups where Google ignores it, 2 hosts writing a noindex line inside robots.txt (a directive that does not exist there), and 1 host serving an entire HTML document with inline JavaScript as its robots.txt. Two hosts wrote license metadata lines. A robots.txt that returns 200 with HTML in it is treated by crawlers as an empty file of unknown rules; it does not block anything.
Homepage tag coverage, 58 of 83 hosts measured
- Any Open Graph tag
- 51 of 58 (88%)
- All four core og tags (title, type, image, url)
- 40 of 58 (69%)
- Any twitter card tag
- 43 of 58 (74%)
- twitter:card value summary_large_image
- 25 of 43 measured card hosts
- Any hreflang
- 23 of 58 (40%)
- hreflang with x-default
- 14 of 23 (61%)
- hreflang with self-reference
- 21 of 23 (91%)
- Invalid or duplicated hreflang codes
- 2 hosts
What this means for you: Open Graph is near-universal at the top of the market but a third of measured hosts miss at least one core tag. Twitter card tags exist on 74% and no host with a twitter card lacks Open Graph entirely, which matches the practice of deriving Twitter tags from og: tags. The Open Graph checker and Twitter card validator test your pages against the same field list.
Homepage text, the 48 hosts with at least 200 words
- Median word count
- 949 words
- Median top-term density
- 3.86%
- 90th percentile density
- 7.82%
- Maximum density
- 12.49% (cloudflare.com)
- Hosts under 200 words served
- 10 of 58, homepages that render by script
Method and limits
robots.txt census: 93 well-known hosts fetched over HTTPS on 2026-08-24 with a 15 second timeout; 83 returned a body and every figure uses those 83. Parsing implements RFC 9309 and the behaviors in Google's robotstxt robots.cc, including sitemap lines not terminating a group. Homepage census: same host list fetched 2026-09-06, one GET each, served HTML only, nothing rendered by script; 58 returned 200 and only they are counted. Text figures cover the 48 hosts whose HTML held at least 200 words.
Limits: a fixed list of large hosts, not a random sample of the web. Homepages only, which are the least typical page of any site. hreflang checked by syntax, not reciprocity. Density figures say nothing about ranking.