
Indexation
Website crawling and indexability
Understand discovery, crawling and indexing, then check the access rules and page signals that matter for important website pages.
For an important URL to appear as a representative page in Google Search, Google needs to find and access its content. The page must also be eligible for indexing. These are separate conditions.
A page that loads for you may be blocked for Googlebot. A page Googlebot fetches may still be excluded or treated as a duplicate. Passing the technical checks does not guarantee indexing or visibility.
Follow the page from discovery to indexing
Google can discover a URL through links or a sitemap, then decide whether to crawl it. It fetches and renders pages during crawling. During indexing, it assesses the content and may choose a representative URL from a group of similar pages. What appears in results also depends on the search.
Start with pages that matter to the business, such as current service pages, categories, products and useful guides. Record the URL intended to represent each one. An intentionally retired page or duplicate filter URL needs a different decision from a missing main service page.
| Question | What to inspect | What the answer tells you |
|---|---|---|
| Can Google discover it? | Crawlable site links and, where used, the sitemap | Whether the URL has a discovery path |
| Can Googlebot request it? | Applicable robots.txt rules, public access and the response | Whether crawling is allowed and the page can be fetched |
| Can Google use the content? | Rendered main content and indexing directives | Whether the fetched page contains the intended answer and permits indexing |
| Is this the selected URL? | Redirects, canonical signals and Search Console’s indexed view | Whether another URL represents the content in Google Search |
Sitemap inclusion can help discovery but does not guarantee crawling or indexing.
Make important pages discoverable
Google can usually discover important pages when internal links from the home page make them reachable, such as through navigation. A sitemap is particularly useful for a large or complex site, a new site with few external links, or a site with substantial video, image or news content.
A sitemap is not essential for every site. Google describes a small site as about 500 pages or fewer, counting only pages intended to appear in search results. A sitemap may not be needed if those pages are comprehensively linked internally and there is little relevant media or news content.
Googlebot determines algorithmically which sites to crawl, how often to visit and how many pages to fetch. It tries to avoid overloading a site and may slow down in response to server errors such as HTTP 500 responses.
Sitemap inclusion: benefits and limitations
- ProsHelps Google discover pages on large, complex, or new sites; useful for video, image, and news content.
- ConsDoes not guarantee crawling or indexing; not essential for small sites with strong internal linking.
Check the minimum conditions for indexing
Google’s technical requirements for eligibility include an accessible page that is not blocked to Googlebot and returns HTTP 200. A page behind a login is not publicly accessible to Googlebot, while client and server error pages are not indexed.
The page also needs indexable content: its text must use a file type supported by Google Search, and the content must comply with Google’s spam policies. Meeting these minimum technical requirements makes a page eligible for indexing.
Separate access rules from index instructions
A robots.txt rule controls which matching URLs Googlebot may request. It does not reliably remove a known URL from search results. A noindex instruction tells Google not to index a page, but Google must be allowed to crawl the URL to read it. An HTML robots meta tag or an X-Robots-Tag response header can carry the instruction; the header also works for non-HTML files.
A canonical annotation expresses a preference among duplicate or very similar URLs. Google makes the final selection. A redirect sends visitors and crawlers to another URL. A page that is genuinely gone without a suitable replacement should return an appropriate missing-page response.
Access rules vs index instructions: what each controls
- `robots.txt` ruleControls whether Googlebot may request the URL; does not remove a page from search results.
- `noindex` instructionPrevents indexing; requires Googlebot to be able to crawl the page first.
- `rel="canonical"` annotationSignals preference among duplicate or similar URLs; Google considers but does not always follow.
- Redirect (e.g., 301)Sends crawlers and users to another URL; strong signal for the preferred version.
Understand how Google selects a representative URL
For duplicate or very similar pages, Google can consider redirects, canonical annotations and sitemap inclusion as signals. Redirects are a strong signal towards the target URL, and a rel="canonical" annotation is also a strong signal; sitemap inclusion is weaker.
These signals can be combined to increase the chance that the preferred URL appears in results, but none is required and Google makes the final choice. When no canonical preference is specified, Google can select the version it considers objectively best for users.
Treat eligibility as a condition, not a promise
Google Search is automated: crawlers regularly explore the web, and most pages in search results are found and added to the index without manual submission.
Google does not accept payment to crawl a site more frequently or to rank it higher.
Check an important page
Follow a normal site route and record the final address. Google generally extracts links from anchors with an href; a destination available only through a script action may be less reliable for discovery. Check the applicable crawl rule, public access and final response. Then review the rendered main content, noindex instructions and canonical signals against the page’s intended role.
For a property you control, use Search Console’s URL Inspection tool to test a specific page. The Page indexing report can help identify pages that are inaccessible to Google.
If a priority page is blocked, identify the rule or access setting before changing it. If Google selects another URL, compare the pages and their canonical signals. If a page was crawled but not indexed, examine the page and its reported status rather than repeatedly requesting a crawl. Record the intended outcome, observed condition, change and recheck result for each affected URL.
In this guide
- Checking whether important pages can be crawledCheck discovery links, robots rules, the final URL and its response, then compare Search Console’s indexed and live views.
- Robots directives versus indexability controlsChoose between robots.txt, noindex, canonical signals and access controls by the outcome needed for a URL.
- Diagnosing a page excluded from search indexingUse URL Inspection, a page’s indexing reason, its live response and canonical details to decide whether exclusion needs a fix.



