Check crawlability of key pages: Use exact URLs and verify Googlebot can reach them; Check robots.txt rules for the final URL’s host and port; Ensure final page returns HTTP 200 and serves indexable content
Image: Search Marketing Desk

Indexation

Part of Website crawling and indexability

Checking whether important pages can be crawled

Check discovery links, robots rules, the final URL and its response, then compare Search Console’s indexed and live views.

To check a page's crawlability, start with its exact URL and intended final destination. Confirm Google has a discovery route, Googlebot is allowed to request the destination, and the final page returns usable content. Opening it in your browser establishes only what happened under your conditions.

Choose URLs that matter

Make a short list of current pages that should be available in Google Search, such as a main service page, a category and a useful guide. Record each intended final URL, not just its title.

For each page, follow a route from the homepage or a relevant category. Google generally extracts crawlable links from HTML anchors with an href.

A menu item available only through a script action may work for visitors without providing a reliable discovery link. A sitemap offers another discovery route, but listing a URL there does not guarantee that Google will crawl it.

Check the request path

Use this sequence for each URL:

  1. Follow the link.Record redirects and the final destination. Check that the destination answers the original page’s purpose.
  2. Check the applicable rule.Read the robots.txt file for the final URL’s protocol, host and port, and determine whether its rules allow Googlebot to request the page. A rule on another host does not apply automatically.
  3. Check public access and response.Confirm that the final page does not require a login and serves the intended content. Google’s minimum technical requirements for an indexable page include a successful HTTP 200 response and indexable content; the starting URL may instead redirect to that page.
  4. Inspect the fetched content.If the main answer depends on JavaScript or other resources, check the rendered result. A successful HTTP response alone does not show that Google received the intended content.

When several important URLs fail in the same way, compare their template, host and recent changes. For one failing URL, inspect its route and settings.

Google’s technical requirements for indexable pages

HTTP Status Code
200 OK
Indexable Content
Visible text and links in HTML
No Login Required
Publicly accessible
Rendered Content
JavaScript-dependent content must be fully rendered

Use Search Console for follow-up

In a Search Console property you control, use URL Inspection to test the exact URL and check its HTTP status code. Treat this as a follow-up to your direct request-path check.

For redirects, record each URL in the route and check the final destination separately.

Google's technical guidance recommends using both the Page indexing and Crawl Stats reports to find pages that are inaccessible to Google; the reports may contain different information.

Record the URL's intended role, discovery route, applicable rule, final response, observation date and proposed change. Retest after a change. A crawlable result is an access finding, not a promise of indexing.

More from Indexation