Free browser tool / Googlebot text search
X-Robots-Tag: the HTTP header that controls indexing
X-Robots-Tag is a response header carrying the same directives as the meta robots tag. Its advantage is reach: it works on every response your server emits, including PDFs, images, and any file type a meta tag can never touch. It is also the only indexing control that applies site-wide from one config line.
Basic syntax
X-Robots-Tag: noindex, nofollow
Directives match the meta list exactly: noindex, nofollow, noarchive, nosnippet, max-snippet, max-image-preview, max-video-preview, notranslate, noimageindex, unavailable_after, and indexifembedded. Multiple headers on one response are allowed; Google applies the most restrictive union, the same rule as multiple meta tags.
Per-user-agent headers
Use the Googlebot variant of the header name to scope rules:
X-Robots-Tag: googlebot: noindex X-Robots-Tag: otherbot: nofollow
Like meta tags, each bot line is independent. Unnamed directives apply to every crawler.
Server configurations that work
nginx
location ~* \.(pdf|jpg|png)$ {
add_header X-Robots-Tag "noindex, nofollow";
}
Site-wide noindex for a staging host:
add_header X-Robots-Tag "noindex, nofollow" always;
Apache
<FilesMatch "\.(pdf|docx)$">
Header set X-Robots-Tag "noindex, nofollow"
</FilesMatch>
Cloudflare Workers
export default {
async fetch(request) {
const res = await fetch(request);
const headers = new Headers(res.headers);
headers.set("X-Robots-Tag", "noindex");
return new Response(res.body, { status: res.status, headers });
}
}
Express
app.use((req, res, next) => {
res.set("X-Robots-Tag", "noindex, nofollow");
next();
});
When to prefer the header over the meta tag
- Non-HTML resources
- PDFs, images, spreadsheets. No meta tag can reach them.
- Site-wide or directory-wide rules
- One config line instead of editing every file. Staging environments especially.
- Pages you cannot edit
- Legacy systems, third-party responses, proxied content.
- Response-level logic
- Conditional headers by status code, query parameter, or auth state.
When to prefer meta: single-page exceptions where the template cannot inject headers, and host environments like static file hosting where you control HTML but not headers.
Verify what your server actually sends
curl -sI https://example.com/report.pdf | grep -i x-robots
Empty output means no header, which means default indexing behavior. Watch for duplicate headers from layered configs: nginx add_header inside a location block replaces inherited headers unless both lines are present, and a CDN can strip or duplicate headers upstream. Confirm the final response, not the origin one. The meta robots checker reads the response header and the head tags together and shows the combined directive set Google will apply.
Common mistakes
- Setting the header on 404 responses
- Harmless but noisy; 404s are already excluded from the index.
- Conflict between header and meta
- Google unions them with the most restrictive result. A meta index plus header noindex means noindex. Diagnose double sources before "fixing" the wrong layer.
- Caching the header in
- A noindex header cached by a CDN keeps deindexing pages after you remove the config. Purge on change.
- Wrong case or typo
- The header is case-insensitive by HTTP but directive spelling is not. "no-index" is an unknown value and is dropped.