“Blocked by robots.txt” means Google did not crawl the URL because a rule in your robots.txt forbids it. “Indexed, though blocked by robots.txt” means Google could not crawl the page but indexed the URL anyway, usually because other pages link to it, so it shows up in results without a description.
The two have different fixes, and the second one surprises many people: robots.txt controls crawling, not indexing. To keep a page out of Google, it must be crawlable and carry a noindex.
Find the rule that blocks a URL
- Open your robots.txt at
https://yourdomain.com/robots.txt. - Find the group that applies to Googlebot: a
User-agent: Googlebotgroup if there is one, otherwiseUser-agent: *. Googlebot follows only the most specific group that matches it. - In that group, the longest matching
DisalloworAllowpath wins; if they are the same length,Allowwins.Disallow: /blocks everything; an emptyDisallow:blocks nothing. - Search Console's robots.txt report (Settings → robots.txt) shows the version Google fetched and any parsing errors. URL Inspection tells you whether a specific URL is blocked.
“Blocked by robots.txt”: what to do
- The block is intended (admin, cart, internal search, filter parameters): nothing to do.
- The block is a mistake (a whole section, product pages, CSS or JavaScript files): remove or narrow the rule. Common errors are a leftover
Disallow: /from staging, a rule likeDisallow: /pthat also matches/products, and blocking/wp-content/or/assets/, which hides the files Google needs to render pages.
“Indexed, though blocked by robots.txt”: what to do
Decide whether you want the page in Google.
- Yes: remove the robots.txt rule so Google can crawl it and show a proper title and description.
- No: remove the robots.txt rule and add
<meta name="robots" content="noindex">(or anX-Robots-Tag: noindexheader). Once Google crawls the page and sees the noindex, it drops it. Keeping the block prevents Google from ever seeing the noindex.
How to check that it worked
- URL Inspection shows Crawl allowed? Yes.
- The URL leaves the reason after Validate fix; pages you set to noindex move to Excluded by ‘noindex’ tag.
Do not forget the other crawlers
The same file decides what Bingbot and AI crawlers such as GPTBot, OAI-SearchBot or PerplexityBot may fetch. Our free AI crawler checker shows which of them your robots.txt allows.
What SignalCrawler checks here
The audit parses robots.txt the way Google does and reports a file that blocks the whole site, rules that block linked sections, internal links that point to blocked URLs, sitemaps that list blocked URLs, lines search engines cannot parse, a robots.txt that answers with a server error, and a missing reference to the sitemap.