Free browser tool / Googlebot text search
Noindex vs robots.txt: which one blocks what
These two files solve different problems, and mixing them up is the most common indexing bug we see in audits. Robots.txt controls whether a crawler may fetch a URL. Noindex controls whether Google may keep that URL in its index. A page blocked by robots.txt can still appear in Google results, because Google lists URLs it knows about even when it cannot read them.
The one-line answer
- Keep a page out of Google's index
- Use
noindex(meta or X-Robots-Tag). The page must be crawlable. - Stop crawlers from wasting your bandwidth
- Use robots.txt. The page may still be indexed by URL.
- Remove an already-indexed page
- noindex, crawlable, and wait for recrawl. Then optionally robots.txt.
Side by side
| Property | robots.txt | meta robots | X-Robots-Tag |
|---|---|---|---|
| Controls | Crawling (fetching) | Indexing (after fetch) | Indexing (after fetch) |
| Where it lives | /robots.txt at site root | head of the HTML | HTTP response header |
| Works on | Any URL type | HTML pages only | Any file: PDFs, images, HTML |
| Can deindex | No | Yes, once crawled | Yes, once crawled |
| Read at | Crawler, before fetch | Parser, after fetch | Parser, after fetch |
The classic failure: robots.txt blocking a noindex page
You put Disallow: /staging/ in robots.txt and a noindex meta on the staging pages, and figure you are covered twice. You are covered zero times. Google never fetches the page, so it never reads the noindex. If links point at those URLs, Google can show them as URL-only results with no snippet, and they can sit there for months. The order of operations is always fetch first, read directives second. Blocking the fetch removes the directive's chance to run.
The reverse order works: let Google crawl the noindex page until it drops out of the index, then add the robots.txt block if you also want to save bandwidth. Google's own documentation warns about this exact combination.
Real syntax, both mechanisms
robots.txt, blocking one path for all crawlers:
User-agent: * Disallow: /private/ Sitemap: https://example.com/sitemap.xml
The noindex equivalent, in the page head:
<meta name="robots" content="noindex">
Or as a response header, which also works on PDFs and images:
X-Robots-Tag: noindex
And the Google-specific crawl block, in robots.txt:
User-agent: googlebot Disallow: /private/
Common mistakes, in the order audits find them
- noindex in the body
- The tag only counts inside the head. Google reads a second head block anywhere in the document, but a tag dropped into body markup is ignored by validators and some parsers.
- noindex plus robots.txt block
- Described above. The block wins, the noindex never loads.
- noindex plus canonical to another URL
- Mixed signals. Canonical says index the other one, noindex says index nothing. Google usually picks one and ignores the other. Pick one intent yourself.
- Blocking right after noindexing
- Safe only after the URL is fully out of the index. Verify with site: or URL Inspection before adding the block.
- Missing X-Robots-Tag on PDFs
- Meta tags do nothing on non-HTML files. Only the header reaches them.
Decision flow
- Step 1
- Is the URL already in Google's index? Check first. If yes and you want it gone, you need a crawlable noindex.
- Step 2
- Gone forever or just unlisted for Google? Gone: noindex, keep crawlable, wait. Unlisted but crawlable for other engines: robots.txt rules for specific user-agents only.
- Step 3
- Is the resource an HTML page? If not, the X-Robots-Tag header is your only indexing control.
- Step 4
- Want to save crawl budget too? After deindexing completes, add the robots.txt rule.
FAQ
- Does robots.txt prevent indexing?
- No. It prevents crawling. Google can still index a blocked URL from external links, showing the URL without a snippet.
- Can I use both noindex and robots.txt at once?
- Only after the URL has left the index. Before that, the block prevents the noindex from ever being read.
- How long does noindex take?
- Until the next recrawl of that page, days to weeks. Requesting a recrawl in Search Console speeds it up.
- Which is better for a whole directory?
- For crawl control, one robots.txt line. For index control on non-HTML files, the X-Robots-Tag header on every response.
Check what your pages actually serve: the meta robots checker reads the head directives and the robots txt checker parses your robots.txt with RFC 9309 matching.