Autonomous Agentic AI Pipeline

Free browser tool / Googlebot text search

X-Robots-Tag: the HTTP header that controls indexing

X-Robots-Tag is a response header carrying the same directives as the meta robots tag. Its advantage is reach: it works on every response your server emits, including PDFs, images, and any file type a meta tag can never touch. It is also the only indexing control that applies site-wide from one config line.

Basic syntax

X-Robots-Tag: noindex, nofollow

Directives match the meta list exactly: noindex, nofollow, noarchive, nosnippet, max-snippet, max-image-preview, max-video-preview, notranslate, noimageindex, unavailable_after, and indexifembedded. Multiple headers on one response are allowed; Google applies the most restrictive union, the same rule as multiple meta tags.

Per-user-agent headers

Use the Googlebot variant of the header name to scope rules:

X-Robots-Tag: googlebot: noindex
X-Robots-Tag: otherbot: nofollow

Like meta tags, each bot line is independent. Unnamed directives apply to every crawler.

Server configurations that work

nginx

location ~* \.(pdf|jpg|png)$ {
    add_header X-Robots-Tag "noindex, nofollow";
}

Site-wide noindex for a staging host:

add_header X-Robots-Tag "noindex, nofollow" always;

Apache

<FilesMatch "\.(pdf|docx)$">
    Header set X-Robots-Tag "noindex, nofollow"
</FilesMatch>

Cloudflare Workers

export default {
  async fetch(request) {
    const res = await fetch(request);
    const headers = new Headers(res.headers);
    headers.set("X-Robots-Tag", "noindex");
    return new Response(res.body, { status: res.status, headers });
  }
}

Express

app.use((req, res, next) => {
  res.set("X-Robots-Tag", "noindex, nofollow");
  next();
});

When to prefer the header over the meta tag

Non-HTML resources
PDFs, images, spreadsheets. No meta tag can reach them.
Site-wide or directory-wide rules
One config line instead of editing every file. Staging environments especially.
Pages you cannot edit
Legacy systems, third-party responses, proxied content.
Response-level logic
Conditional headers by status code, query parameter, or auth state.

When to prefer meta: single-page exceptions where the template cannot inject headers, and host environments like static file hosting where you control HTML but not headers.

Verify what your server actually sends

curl -sI https://example.com/report.pdf | grep -i x-robots

Empty output means no header, which means default indexing behavior. Watch for duplicate headers from layered configs: nginx add_header inside a location block replaces inherited headers unless both lines are present, and a CDN can strip or duplicate headers upstream. Confirm the final response, not the origin one. The meta robots checker reads the response header and the head tags together and shows the combined directive set Google will apply.

Common mistakes

Setting the header on 404 responses
Harmless but noisy; 404s are already excluded from the index.
Conflict between header and meta
Google unions them with the most restrictive result. A meta index plus header noindex means noindex. Diagnose double sources before "fixing" the wrong layer.
Caching the header in
A noindex header cached by a CDN keeps deindexing pages after you remove the config. Purge on change.
Wrong case or typo
The header is case-insensitive by HTTP but directive spelling is not. "no-index" is an unknown value and is dropped.

Related pages