“Excluded by ‘noindex’ tag”: how to find where the noindex comes from

“Excluded by ‘noindex’ tag” means Google crawled the page and found an instruction not to index it, so it did not. For admin pages, thank-you pages, internal search results and duplicate filters that is correct. If the list includes pages you want in Google, the noindex is a mistake and only you can remove it.

Where a noindex can come from

  1. A meta tag in the HTML: <meta name="robots" content="noindex"> (or none, or a googlebot meta tag).
  2. An HTTP header: X-Robots-Tag: noindex, set by the server, a CDN or a security plugin. It is invisible in the page source.
  3. A CMS setting: WordPress “Discourage search engines from indexing this site” (Settings → Reading), an SEO plugin's per-type setting (for example tags, archives or a custom post type set to noindex), Shopify or Wix page visibility settings.
  4. A staging leftover: the site was launched with the staging environment's noindex still on. This is the most common cause of a whole site disappearing.
  5. JavaScript: a script that adds a noindex tag after load. Google follows it.

How to find it on a page

  • URL Inspection → Test live URL → View tested page shows the HTML Google received; search it for noindex. The More info tab shows the HTTP headers.
  • In a browser, view the page source and search for robots. For the header, open developer tools → Network → the page request → Response Headers.
  • From a terminal: curl -sI https://example.com/page | grep -i x-robots for the header, and curl -s https://example.com/page | grep -i noindex for the tag.

How to fix it

  1. Decide per page type: should these pages be in Google? Leave the noindex on the ones that should not.
  2. Remove the noindex where it is wrong: in the CMS or SEO plugin setting for that page type, in the template, or in the server or CDN rule that adds the header.
  3. Make sure robots.txt does not block those pages; Google has to crawl a page to see that the noindex is gone.
  4. Link to the pages from the site and list them in the sitemap.
  5. Inspect a fixed URL and Request indexing, then Validate fix on the reason.

Two mistakes to avoid

  • Noindex on pages you link from the main navigation. It sends mixed signals and wastes crawling. Either make such pages indexable or remove them from navigation.
  • Noindex together with a robots.txt block. Google cannot crawl the page, so it never sees the noindex, and the URL may still appear in results without a description.

How to check that it worked

  • URL Inspection shows Indexing allowed? Yes.
  • The URL moves from this reason to Indexed after validation.

What SignalCrawler checks here

The audit reads both the meta robots tag and the X-Robots-Tag header on every crawled page. It reports a homepage that cannot be indexed, sites where most pages are noindex, noindex pages linked from navigation, sitemaps that list noindex pages, and canonical tags that point to noindex pages. On a staging audit, a sitewide noindex is listed as a pre-launch item to remove at launch.