Robots vs indexability controls: Use `robots.txt` to block Googlebot from crawling URLs; Use `noindex` to keep accessible pages out of Google Search; A canonical signal favours one URL among duplicates
Image: Search Marketing Desk

Indexation

Part of Website crawling and indexability

Robots directives versus indexability controls

Choose between robots.txt, noindex, canonical signals and access controls by the outcome needed for a URL.

Choose a control by the outcome needed for the URL. Use robots.txt to control Googlebot requests to matching paths. Use noindex when an accessible page should stay out of Google's index.

Use a canonical signal when duplicate or very similar URLs should have one representative. These controls solve different problems; a crawl block can hide the indexing instruction Google needs to read.

Decide what should happen to the URL

Intended outcomeRelevant controlImportant limit
Prevent Googlebot crawling of matching URLsrobots.txt disallow ruleA discovered blocked URL can still appear in Google Search without its page content.
Keep an accessible page out of Google Searchnoindex in an HTML robots meta tag or X-Robots-Tag headerGoogle must be allowed to crawl the URL to read it.
Prefer one URL among duplicates or near duplicatesRedirect if the alternate should send visitors elsewhere; canonical annotation if it should remain accessibleGoogle ultimately chooses the canonical; an annotation is a signal.
Restrict access to private informationAuthentication or another access controlSearch directives are not privacy controls.
Show that content is gone with no suitable replacementA genuine missing-page responseA working-page response should not conceal a missing page.

A sitemap supports discovery and provides a weak canonical signal. It is neither an index command nor a guarantee that a listed URL will be crawled or indexed.

Avoid hiding noindex behind a crawl block

Suppose a public confirmation page should stay available to people who have its address but should not appear in Google Search. A crawlable noindex may suit that aim. If the same URL is disallowed in robots.txt, Google cannot fetch its meta tag or response header to read the instruction.

The blocked URL may still be known through links. For a public page intended to use noindex, allow the crawl and verify the actual directive on the response.

For HTML, inspect the robots meta tag and the HTTP response for an X-Robots-Tag. A server or CDN may apply a header across many files, regardless of a CMS setting. A PDF cannot carry an HTML meta tag, so its indexing instruction must use a response header. Check the final response rather than relying only on a settings screen.

Treat canonicalisation separately

A canonical annotation asks Google to favour one URL in a similar-content group. It does not forbid crawling the alternate. A redirect moves requests to another URL and may be appropriate when the old page should no longer remain accessible. Neither control is a substitute for restricting private information.

Before changing a setting, record the intended visitor and search outcome. Then check the applicable crawl rule, final response, robots meta tag, response headers and canonical signals. Use Search Console’s URL Inspection tool to test a specific page.

More from Indexation