Skip to content

Technical SEO Checklist 2026: Built for AI Crawlers

The technical SEO checklist for 2026, rebuilt for AI crawlers, crawlability, Core Web Vitals, structured data, llms.txt, and bot-specific robots.txt rules.

30 Jul 20267 min read
  • AEO
Code on a developer's screen, illustrating Technical SEO Checklist 2026: Built for AI Crawlers

A complete technical SEO checklist for 2026 covers classic fundamentals, crawlability, sitemaps, Core Web Vitals, mobile-friendliness, canonical tags, plus a new layer specific to AI crawlers: structured data built for LLM consumption, clean semantic HTML, server-rendered content, explicit robots.txt directives for bots like GPTBot and PerplexityBot, and an optional llms.txt file. Below is the full checklist, grouped so you can work through it section by section.

Part 1: Crawlability Fundamentals (Still Non-Negotiable)

  • [ ] robots.txt is present, valid, and not accidentally blocking key content. Check it with Google Search Console's robots.txt tester or a manual fetch, a stray Disallow: / left over from staging is more common than you'd think.
  • [ ] XML sitemap exists, is submitted to Search Console, and is actually current. Stale sitemaps listing removed pages (or missing new ones) waste crawl budget and delay indexing.
  • [ ] No orphan pages. Every important page should be reachable via internal links, not just the sitemap, crawlers (AI and traditional) weight internally-linked pages more heavily.
  • [ ] HTTP status codes are clean. Audit for unexpected 404s, redirect chains (A→B→C instead of A→C), and soft 404s (pages that return 200 but show "not found" content).
  • [ ] Crawl budget isn't wasted on low-value URLs. Parameter-heavy URLs, faceted navigation, and duplicate filtered views should be canonicalized or excluded via robots.txt where appropriate.

Part 2: Indexing and Canonicalization

  • [ ] Canonical tags are set correctly on every page, including self-referencing canonicals on pages that don't have duplicates.
  • [ ] No conflicting signals, e.g., a page canonicalized to URL A but also included in the sitemap as URL B, or a noindex tag on a page you actually want indexed.
  • [ ] Pagination and duplicate content are handled deliberately (canonical to a "view all" version where feasible, or clear canonical chains for paginated series).
  • [ ] HTTPS is enforced site-wide with proper redirects from HTTP, and no mixed-content warnings.
  • [ ] International/multi-language sites use hreflang correctly, if applicable.

Part 3: Core Web Vitals and Page Experience

  • [ ] Largest Contentful Paint (LCP) is under ~2.5 seconds on real-world mobile conditions, not just lab tests.
  • [ ] Interaction to Next Paint (INP) is under ~200ms, this replaced First Input Delay as the responsiveness metric and is worth checking specifically if you haven't audited since the metric changed.
  • [ ] Cumulative Layout Shift (CLS) is under 0.1, check for late-loading ads, images without dimensions, and web fonts causing layout jumps.
  • [ ] Images are compressed and served in modern formats (WebP/AVIF) with proper lazy-loading for below-the-fold content.
  • [ ] Third-party scripts (chat widgets, analytics, ad tags) are audited for their performance cost, these are a common, under-checked source of Core Web Vitals failures.

Part 4: Mobile-Friendliness

  • [ ] Site passes Google's mobile-friendly checks, responsive layout, readable font sizes without zooming, tap targets sized appropriately.
  • [ ] No mobile-specific blocked resources (CSS/JS files accidentally disallowed for mobile crawling).
  • [ ] Mobile page speed is tested separately from desktop, mobile network conditions and device processing power are still the bottleneck for most users.

Part 5: Structured Data, Classic and LLM-Oriented

  • [ ] Schema.org markup is implemented for core page types: Article/BlogPosting for content, Organization and Person for author/entity data, Product for product pages, FAQPage for genuine FAQ sections, and BreadcrumbList for navigation context.
  • [ ] Structured data validates cleanly in Google's Rich Results Test or Schema.org's validator, broken JSON-LD is worse than none, since it can create confusing signals.
  • [ ] Author and organization entities are clearly and consistently marked up across pages, this reinforces the E-E-A-T signals that both traditional ranking and AI-citation selection appear to weight.
  • [ ] FAQ and Q&A structured data is used honestly, matching visible on-page Q&A content, not stuffed with unrelated queries purely to game rich results.
  • [ ] Consider llms.txt as a low-effort addition, a markdown file at /llms.txt summarizing your site for LLM consumption. Evidence that major AI crawlers consistently use it is limited and unconfirmed as of 2026, so treat it as a cheap supplement, not a core strategy.

Part 6: Clean, Semantic HTML

  • [ ] Use real heading hierarchy (single H1, logically nested H2s/H3s) rather than styled <div>s pretending to be headings, this is what both search crawlers and AI systems use to understand your content's structure.
  • [ ] Core content is in actual HTML text, not images or embedded documents, wherever possible, text baked into an image is invisible to every crawler, AI or otherwise.
  • [ ] Semantic tags are used appropriately (<article>, <nav>, <main>, <figure>) rather than a soup of generic <div>s, this helps automated systems parse what's primary content versus navigation/boilerplate.
  • [ ] Alt text is descriptive and accurate on meaningful images, not just present for compliance's sake.
  • [ ] Tables are used for actual tabular data with proper <th> headers, since AI systems parsing for facts/comparisons rely on this structure.

Part 7: AI-Crawler-Specific Accessibility

  • [ ] Critical content is server-rendered or statically generated, not client-side-only. Many AI crawlers (GPTBot, ClaudeBot, PerplexityBot, CCBot) are generally reported to have limited or no JavaScript rendering capability, unlike Googlebot. If your key pages depend on client-side rendering to display content, test this directly (see below) rather than assuming it works.
  • [ ] Test what each crawler actually sees. Fetch key pages with curl -A "GPTBot" [url] (and equivalents for other bots) and compare the raw HTML response to what a browser renders. If your core content is missing from the raw response, that crawler likely can't read it.
  • [ ] robots.txt includes explicit, deliberate directives for AI bots. Common bots to consider explicitly allowing or disallowing: GPTBot (OpenAI), ChatGPT-User (OpenAI, used for live browsing), ClaudeBot/anthropic-ai (Anthropic), PerplexityBot (Perplexity), CCBot (Common Crawl, which feeds many AI training sets), and Google-Extended (controls Google's AI training use separately from standard Googlebot indexing). Decide deliberately per bot, blocking all of them protects content from AI training/citation entirely, which may or may not match your visibility goals.
  • [ ] Check server/CDN logs for AI bot activity to confirm crawlers are actually reaching your site, not just theoretically allowed to.
  • [ ] Consider a pre-rendering service (e.g., Prerender.io or equivalent) as a stopgap if a full SSR/SSG migration isn't immediately feasible for a JS-heavy site.

Part 8: Content Extractability (The AEO Layer)

  • [ ] Every page opens with a direct, self-contained answer in the first 2-3 sentences after the H1, this is what AI systems most reliably lift into summaries.
  • [ ] Headers are phrased as real questions where relevant ("How does X work?" rather than a vague label), matching how people actually query AI systems.
  • [ ] Claims are specific and sourced, dates, numbers, and named examples make content harder to generate away without attribution and more likely to be selected as a citation.
  • [ ] Duplicate or near-duplicate content across your own site is consolidated, competing pages for the same query confuse both traditional ranking and AI source-selection.

Part 9: Ongoing Monitoring

  • [ ] Google Search Console is set up and checked regularly for coverage errors, manual actions, and Core Web Vitals reports.
  • [ ] A rank-tracking or SEO platform with AI Overview/citation tracking is in place if budget allows, since Google Search Console doesn't natively report AI Overview citation performance.
  • [ ] Server logs are periodically reviewed for both traditional and AI crawler activity patterns.
  • [ ] A recurring (quarterly, at minimum) technical audit re-checks this whole list, technical SEO decays quietly as sites get new features, migrations, and third-party scripts added over time.

FAQ

What's genuinely new in this checklist compared to a pre-2024 technical SEO checklist? Parts 7 and 8 are the newest additions, explicit AI-bot handling in robots.txt, JS-rendering verification for non-Google crawlers, and content structured for direct extraction rather than just ranking. Everything else (crawlability, Core Web Vitals, structured data, semantic HTML) was already best practice, just now doing double duty for AI systems too.

Do I need to block AI crawlers like GPTBot if I don't want my content used for AI training? That's a legitimate choice, but it's a trade-off: blocking GPTBot or Google-Extended can reduce your chances of being cited in tools built on those systems. There's no universally correct answer, decide based on whether you value AI visibility or content protection more for your specific business.

How often should I run a full technical SEO audit in 2026? Quarterly is a reasonable baseline for most sites, with lighter spot-checks after any major site migration, framework upgrade, or CMS change, since rendering and crawlability issues often get introduced silently during development work.

Is structured data still worth the implementation effort if AI crawlers might not use llms.txt? Yes, structured data (schema.org/JSON-LD) has far more established support across both traditional search and AI systems than llms.txt does, so it remains one of the higher-confidence investments on this checklist.

What's the single highest-priority item for a JS-heavy startup site? Verifying that critical content is actually present in the raw server HTML response, not just rendered client-side after JavaScript runs. It's the most common blind spot on modern React/Next.js marketing sites and the one most likely to silently make you invisible to non-Google AI crawlers.


I run this exact checklist as part of technical SEO/AEO audits for edtech and startup clients, it's usually where the highest-leverage, lowest-effort fixes turn up, especially the JS-rendering and structured-data gaps that are easy to miss on fast-moving product sites.