Deindexing pages on purpose
Site owners remove pages that shouldn’t appear in search, such as thin pages, duplicates, internal search results, thank-you pages or outdated content:
- Noindex: a
noindexrobots meta tag or X-Robots-Tag header. The page must stay crawlable, so Google can see the instruction. - 404 or 410: delete the page, and it drops out as Google recrawls it.
- 301 redirect: send the URL to a better page, which takes its place in the index.
- Removals tool: in Google Search Console, hides a URL from results for about six months while a permanent fix takes effect.
Blocking a page in robots.txt doesn’t deindex it: Google can still list the URL if other sites link to it.
Unintended deindexing
- A noindex tag left in place after a redesign, or “Discourage search engines” switched on in WordPress
- Robots.txt blocking the whole site, or server errors and timeouts
- Wrong canonical tags that point pages at other URLs
- Google choosing not to index low-quality or duplicate pages
- A manual action for spam, or a hacked site
How to check
The Pages report and the URL Inspection tool in Google Search Console show which pages are indexed and why others aren’t. A site: search gives a rough idea, but it isn’t complete.