Autonomous Agentic AI Pipeline

Free browser tool / Googlebot text search

Noindex vs robots.txt: which one blocks what

These two files solve different problems, and mixing them up is the most common indexing bug we see in audits. Robots.txt controls whether a crawler may fetch a URL. Noindex controls whether Google may keep that URL in its index. A page blocked by robots.txt can still appear in Google results, because Google lists URLs it knows about even when it cannot read them.

The one-line answer

Keep a page out of Google's index
Use noindex (meta or X-Robots-Tag). The page must be crawlable.
Stop crawlers from wasting your bandwidth
Use robots.txt. The page may still be indexed by URL.
Remove an already-indexed page
noindex, crawlable, and wait for recrawl. Then optionally robots.txt.

Side by side

Propertyrobots.txtmeta robotsX-Robots-Tag
ControlsCrawling (fetching)Indexing (after fetch)Indexing (after fetch)
Where it lives/robots.txt at site roothead of the HTMLHTTP response header
Works onAny URL typeHTML pages onlyAny file: PDFs, images, HTML
Can deindexNoYes, once crawledYes, once crawled
Read atCrawler, before fetchParser, after fetchParser, after fetch

The classic failure: robots.txt blocking a noindex page

You put Disallow: /staging/ in robots.txt and a noindex meta on the staging pages, and figure you are covered twice. You are covered zero times. Google never fetches the page, so it never reads the noindex. If links point at those URLs, Google can show them as URL-only results with no snippet, and they can sit there for months. The order of operations is always fetch first, read directives second. Blocking the fetch removes the directive's chance to run.

The reverse order works: let Google crawl the noindex page until it drops out of the index, then add the robots.txt block if you also want to save bandwidth. Google's own documentation warns about this exact combination.

Real syntax, both mechanisms

robots.txt, blocking one path for all crawlers:

User-agent: *
Disallow: /private/
Sitemap: https://example.com/sitemap.xml

The noindex equivalent, in the page head:

<meta name="robots" content="noindex">

Or as a response header, which also works on PDFs and images:

X-Robots-Tag: noindex

And the Google-specific crawl block, in robots.txt:

User-agent: googlebot
Disallow: /private/

Common mistakes, in the order audits find them

noindex in the body
The tag only counts inside the head. Google reads a second head block anywhere in the document, but a tag dropped into body markup is ignored by validators and some parsers.
noindex plus robots.txt block
Described above. The block wins, the noindex never loads.
noindex plus canonical to another URL
Mixed signals. Canonical says index the other one, noindex says index nothing. Google usually picks one and ignores the other. Pick one intent yourself.
Blocking right after noindexing
Safe only after the URL is fully out of the index. Verify with site: or URL Inspection before adding the block.
Missing X-Robots-Tag on PDFs
Meta tags do nothing on non-HTML files. Only the header reaches them.

Decision flow

Step 1
Is the URL already in Google's index? Check first. If yes and you want it gone, you need a crawlable noindex.
Step 2
Gone forever or just unlisted for Google? Gone: noindex, keep crawlable, wait. Unlisted but crawlable for other engines: robots.txt rules for specific user-agents only.
Step 3
Is the resource an HTML page? If not, the X-Robots-Tag header is your only indexing control.
Step 4
Want to save crawl budget too? After deindexing completes, add the robots.txt rule.

FAQ

Does robots.txt prevent indexing?
No. It prevents crawling. Google can still index a blocked URL from external links, showing the URL without a snippet.
Can I use both noindex and robots.txt at once?
Only after the URL has left the index. Before that, the block prevents the noindex from ever being read.
How long does noindex take?
Until the next recrawl of that page, days to weeks. Requesting a recrawl in Search Console speeds it up.
Which is better for a whole directory?
For crawl control, one robots.txt line. For index control on non-HTML files, the X-Robots-Tag header on every response.

Check what your pages actually serve: the meta robots checker reads the head directives and the robots txt checker parses your robots.txt with RFC 9309 matching.

Related pages