SEO and AI-search checklist

The 2026 on-page SEO and AI-search checklist, with every item automated

Most SEO checklists are a list of advice. This one is the list of checks SEOAST actually runs, in the order they matter: first whether the page can be fetched and indexed, then whether its tags and headings describe it, then whether its content is in the HTML and can be lifted into an answer. Work through it by hand, or run the audit and get the same list back with your page’s results filled in.

No signup required to start — 10 audits a day without an account.

The 23 checks behind this

Each one is a check that runs, described by the engine itself. If a check changes, this page changes with it.

  • HTTP status

    http_status

    The status of the final response after redirects. Fails on anything outside the 2xx range, because a URL that does not answer 200 has no document to index; warns on a 2xx that is not 200. Reports not_measurable when the audit was run against supplied HTML rather than a live fetch.

    Fixing it: Gives search engines a document to index at this URL at all.

  • Redirect chain

    redirect_chain

    Every hop followed to reach the page. Fails on a loop, on any hop through plain http, and at 3 or more hops; warns above 1. Each hop costs crawl budget and adds latency for every visitor.

    Fixing it: Cuts latency for every visitor and stops crawl budget being spent on hops.

  • HTTPS

    transport_security

    The scheme of the final URL and whether a Strict-Transport-Security header was returned. Plain http fails; https without HSTS warns, because the first request of a session can still go out in the clear.

    Fixing it: Removes the browser "not secure" warning and satisfies a confirmed ranking signal.

  • X-Robots-Tag header

    x_robots_tag

    Indexing directives delivered in the response header rather than the markup, including any bot-name prefix. A noindex or none fails; nofollow warns. This is checked separately from the robots meta tag because the two disagree often and a header-level noindex is invisible in the HTML.

    Fixing it: Removes a header-level block that keeps the page out of the index entirely.

  • robots.txt

    robots_txt

    The origin /robots.txt, parsed to RFC 9309: consecutive user-agent lines share a group, the most specific group wins, and the longest matching pattern decides with Allow winning ties. Fails when the audited path is disallowed for Googlebot; warns when the file is absent or declares no Sitemap. Distinct from the robots meta tag: a disallow here stops the crawl, not just the indexing.

    Fixing it: Lets crawlers read the page, so every other signal on it can be seen.

  • Robots meta directives

    robots_meta

    Directives across meta robots, googlebot, bingbot and googlebot-news tags. A noindex or none directive fails the check outright because it removes the page from search regardless of every other signal; nofollow warns because it strips the page of its outbound crawl value.

    Fixing it: Removes a directive that keeps the page out of search regardless of its content.

  • XML sitemap

    sitemap

    Discovery via the robots.txt Sitemap directive then the conventional paths, followed by XML validation. Fails when nothing retrieved parses, or when a file breaks the 50,000 URL or 50 MB protocol limits; warns on relative or cross-host <loc> values, on partial parse failures, and when a valid sitemap does not list the audited page.

    Fixing it: Shortens the delay between publishing a page and it being discovered.

  • Canonical URL

    canonical

    Presence of <link rel="canonical">, whether its href is absolute in the markup, whether it resolves to the audited URL itself, and whether more than one canonical is declared. A missing or conflicting canonical lets parameter and trailing-slash variants compete with the page.

    Fixing it: Consolidates duplicate URL variants onto one address, so ranking signals stop being split.

  • Title tag

    title

    Presence, length and structure of <title>. Fails when absent or empty; warns outside 15-60 characters (Google truncates around 60) or when a separator-delimited segment such as the brand name is repeated within the same title.

    Fixing it: Improves the strongest on-page relevance signal and the line people click.

  • Meta description

    meta_description

    Presence, uniqueness and length of <meta name="description">. Fails when absent or empty; warns outside 70-160 characters or when more than one description tag is present. Does not rank, but it is the copy that earns the click.

    Fixing it: Raises click-through on the ranking you already have. Does not affect ranking itself.

  • Heading hierarchy

    heading_hierarchy

    The h1-h6 outline: how many h1 elements exist, whether any heading is empty, and whether the document skips a level (an h2 followed directly by an h4). Zero or multiple h1 elements and skipped levels both make the page outline ambiguous to crawlers and screen readers.

    Fixing it: Gives crawlers and screen readers an unambiguous outline of what the page covers.

  • H1 and title alignment

    h1_title_alignment

    Overlap between the meaningful terms in <title> and in the first <h1>, as a fraction of the smaller term set. At or above 0.5 the two describe the same topic; at or above 0.25 they are loosely related; below that the search result promises one thing and the page delivers another. Reports not_measurable when either element is missing or empty.

    Fixing it: Makes the search result and the page agree about the topic, reducing bounce.

  • Server-rendered content

    render_dependency

    Visible text as a share of the response body, plus inline script share and server-rendered word count. Fails under 60 words alongside a script-heavy document, or below a 2% text ratio; warns below 5%. SEOAST executes no JavaScript, so this measures what a crawler sees before deciding whether to queue a rendered pass.

    Fixing it: Puts the content in the first response, where every crawler reads it, instead of behind a render pass that is neither fast nor guaranteed.

  • Content depth and readability

    content_readability

    Visible copy taken from <main>, <article>, <body> or the document, then measured for word count and Flesch Reading Ease. Under 300 words reads as thin; a reading ease below 30 reads as academic or legal prose to a general audience.

    Fixing it: Gives the page enough substance to be a credible answer to the query.

  • Question-and-answer extractability

    qa_content

    Headings that end in a question mark, and summary/details blocks, in the main content — each paired with the copy beneath it, plus any FAQPage or QAPage JSON-LD. Fails when marked-up questions are absent from the visible copy, or when every question on the page has under 15 words beneath it; warns when questions carry no markup, when some go unanswered, or when an answer runs past 200 words before reaching the point. A page with no questions on it reports not_measurable — this measures how liftable existing answers are and does not require a page to have any.

    Fixing it: Lets an answer engine lift a complete answer off this page instead of a fragment of one.

  • Structured data (JSON-LD)

    structured_data

    Every <script type="application/ld+json"> block: whether it parses as JSON, whether it declares an @type (walking arrays and @graph), and which types are present. A block that fails to parse is worse than no block at all, because it forfeits rich-result eligibility silently.

    Fixing it: Makes the page eligible for rich results. Eligible, not guaranteed — Google decides whether to show them.

  • Open Graph tags

    open_graph

    Presence of og:title, og:description and og:image, plus whether the image URL is absolute. Relative og:image values are not resolved by most social crawlers, so the preview renders blank.

    Fixing it: Makes shared links render with a title, description and image instead of a bare URL.

  • X (Twitter) card

    twitter_card

    Presence and validity of twitter:card (summary, summary_large_image, app or player) and the accompanying title, description and image. Complete Open Graph tags are treated as an acceptable fallback, because X reads them when the twitter:* equivalents are absent.

    Fixing it: Makes links shared on X render as a card rather than plain text.

  • Image alt text

    image_alt

    Share of <img> elements with no alt attribute at all. An explicit alt="" is counted as a deliberate decorative marker, not a defect. More than 25% of images missing alt fails; a smaller share warns. Reports not_measurable when the page has no images.

    Fixing it: Makes images usable by screen readers and eligible for image search.

  • Image dimensions (layout shift)

    image_dimensions

    Share of <img> elements missing both width and height attributes. Without intrinsic dimensions the browser cannot reserve space before the image loads, which is a direct input to Cumulative Layout Shift. Reports not_measurable when the page has no images.

    Fixing it: Reserves space before images load, improving Cumulative Layout Shift.

  • Internal linking

    internal_linking

    Anchors classified into same-site, external, same-page fragment, non-navigation (mailto:, tel:, javascript:) and href-less. Fewer than 3 internal links makes the page a crawl dead end. Reports not_measurable when the final URL could not be parsed, since there is no origin to judge "same site" against.

    Fixing it: Gives crawlers a path onward from this page and spreads authority through the site.

  • HTML lang attribute

    html_lang

    Presence and shape of the lang attribute on <html>, checked loosely against BCP 47 (a two- or three-letter primary subtag plus optional script, region and variant subtags). Drives language targeting, hreflang consistency and screen-reader pronunciation.

    Fixing it: Tells search engines and screen readers which language the page is written in.

  • Viewport meta tag

    viewport

    Presence of <meta name="viewport">, whether it sets width=device-width, and whether it blocks pinch zoom via user-scalable=no or a maximum-scale below 2. A missing viewport fails mobile-first indexing; a zoom lock is an accessibility defect.

    Fixing it: Makes the page usable on mobile, which is the index Google ranks from.

How this works in practice

Start with access, because nothing else counts without it

Status code, redirects, HTTPS, robots.txt, the robots meta tag, the X-Robots-Tag header and the canonical decide whether a page is eligible at all. A noindex left in a response header after a migration cancels every other item on this list, which is why these come first.

Then the parts Google now calls main content

In October 2026 Google updated its helpful-content documentation to list page titles and headings as part of main content. Check that the title says what the page delivers, that there is one clear H1 in agreement with it, and that the headings summarize the page on their own.

Then what a machine can lift from the page

Copy that only appears after JavaScript runs is invisible to fetchers that do not render. Confirm the main text is in the server HTML, that questions are answered directly under their headings, and that your JSON-LD parses. SEOAST does not query ChatGPT, Perplexity, Gemini or Google AI Overviews, and it does not track citations or produce an AI visibility score.

What this will not do

On-page checks for one URL at a time. It does not measure Core Web Vitals, backlinks or rankings.

Questions

How many checks are on the list?

This page lists the 23 on-page checks shown above. Local businesses can add the local audit, which runs its own set of name, address, phone and location checks.

Is there a downloadable version?

Not yet. The list on this page is the complete checklist, and the audit returns it with your results.

Does the checklist cover Core Web Vitals?

No. SEOAST reads server HTML and does not measure field performance. It does check image dimensions, which is one cause of layout shift.

  • Free SEO audit and website SEO checker

    Free SEO audit for any URL, no signup. SEOAST checks indexability, on-page tags, headings, structured data, server-rendered content and AI-crawler access, shows the evidence for each result and ranks the fixes.

  • Technical SEO audit

    Run a technical SEO audit on any URL. SEOAST checks HTTP status, redirects, HTTPS, robots.txt, the robots meta tag, X-Robots-Tag, canonicals and your XML sitemap, then ranks what to fix first.

  • Answer engine optimization (AEO) audit

    AEO audit for any URL. SEOAST checks whether your answers can be extracted: question headings, direct answers beneath them, FAQPage and QAPage schema, heading structure, readability and server-rendered copy.

Audit your page

No signup required to start.