The technical part of an SEO audit: can search engines reach, read and index your pages, and do they get one clear version of each? Work through it once per site and again after every big release. Download it as a spreadsheet with a status column.
Most checks are faster with a crawler than by hand; the right-hand notes say where to look.
Crawling
- robots.txt answers 200 and does not block pages you want indexed — open /robots.txt; check Search Console's robots.txt report
- CSS and JavaScript files are not blocked — Google needs them to render pages
- Important pages are reachable within three or four clicks from the homepage — crawl depth report
- No pages are only in the sitemap without internal links — orphan pages
- Parameter URLs (sort, filter, session) do not multiply into a crawl trap — count URL variants per path
- Server responses are fast and free of 5xx errors — Search Console Crawl stats
- Missing pages answer 404, not 200 — request a made-up address
Indexing
- The homepage and key pages are indexable — no noindex in meta robots or X-Robots-Tag
- No noindex pages are linked from the main navigation — mixed signals
- Search Console's Pages report has no unexplained growth in a reason — review each reason
- Internal search result pages are not indexable — unless chosen on purpose
- Thin and near-duplicate pages are merged, improved or excluded — by template
Redirects and status codes
- HTTP and www variants 301 to one HTTPS address in one hop — test all four
- No redirect chains or loops — every redirect is one hop to a 200
- Internal links point to final URLs, not redirects — fix in templates
- No internal links to 4xx pages — broken links report
- One trailing-slash form; the other redirects — /page and /page/ do not both answer 200
Canonicals
- Every indexable page has a self-referencing canonical — absolute URL
- Only one canonical tag per page — theme and plugin may both add one
- Canonicals do not point to redirects, errors or noindex pages — check targets
- Paginated pages do not canonicalize to page 1 — each page points to itself
- Not many pages canonicalize to the homepage — template bug
Sitemaps
- An XML sitemap exists and is listed in robots.txt — and submitted in Search Console
- The sitemap lists only canonical, indexable URLs that answer 200 — no redirects, 404s or noindex
- lastmod dates are real, not all the same — generated from content changes
- No URLs from another host — staging or old domain
On-page basics
- Unique titles that fit in about 600 px — check with the snippet checker
- Unique meta descriptions that fit in about 920 px — and are not too short
- One H1 per page that matches the topic — and agrees with the title
- The page language is declared — html lang
- Open Graph tags for social sharing — og:title, og:description, og:image
International
- hreflang annotations are reciprocal — every language version links back
- hreflang includes the page itself and uses valid codes — en-GB, not en-UK
- hreflang targets answer 200, are indexable and canonical to themselves — no redirects or noindex
- HTML and sitemap hreflang agree — if both are used
Structured data
- JSON-LD parses without errors — Rich Results Test
- Required properties are present for the types used — Product, Article, Organization, BreadcrumbList
- Organization data has name, logo and sameAs profiles — on the homepage
- Marked-up FAQ content is visible on the page — no hidden FAQ markup
Rendering
- The main content is in the HTML source, not only after JavaScript — view source
- JavaScript does not change titles, canonicals or robots tags after load — compare raw and rendered HTML
- Links are real a href links, not click handlers — so crawlers can follow them
Page experience
- Core Web Vitals pass for real users — LCP, INP and CLS in Search Console
- Pages are usable on phones with no horizontal scroll — 320 to 430 px
- HTTPS everywhere with no mixed content — browser console
- No intrusive pop-ups covering the content on phones — on landing from search
AI crawlers
- AI search crawlers are allowed or blocked on purpose — OAI-SearchBot, PerplexityBot, Claude-SearchBot
- Training crawlers are allowed or blocked on purpose — GPTBot, ClaudeBot, CCBot, Google-Extended