Introduction: The Foundation Hasn’t Moved. The Consequences of Ignoring It Have.
Here’s the uncomfortable truth most technical SEO content as we head towards 2027 tries to avoid: none of the fundamentals have actually changed. Fast, crawlable, well-structured websites were the foundation of search visibility a decade ago, and they still are today. What’s changed is who’s reading them, how unforgiving the audience has become, and what it costs you when you get it wrong.
Google’s own guidance is unambiguous on this point: technical SEO is the discipline of helping search engines find, crawl, understand, and index your pages. Ahrefs’ own technical SEO primer is clear that this now extends beyond traditional search: AI search still depends on crawlable, well-structured, trustworthy web pages. A JavaScript-rendered product catalogue that Googlebot has spent a decade learning to parse is often a blank page to the crawlers powering ChatGPT, Claude and Perplexity. A slow, technically messy site that used to just underperform in rankings can now simply not exist in an AI-generated answer at all.
This guide extends our BrightonSEO 2026 recap post on technical SEO in the age of AI, and pulls in current guidance from across the industry – Google’s own Search Central documentation, Ahrefs, Moz, Semrush, Screaming Frog, Search Engine Land, Neil Patel and Bruce Clay – to give you a single, prioritised reference for what actually matters in 2027, and what’s noise.
By the end of this guide, you’ll understand:
- Why technical SEO has become more important, not less, as AI search has grown
- Crawlability and indexability fundamentals, including how AI bots like GPTBot actually behave
- Where Core Web Vitals stand in 2027 and which thresholds still matter
- How structured data earns both rich results and AI citations
- The mobile-first indexing and JavaScript rendering mistakes still costing sites visibility
- Platform-specific checklists for Shopify, WordPress and custom builds
- A realistic 90-day technical SEO action plan, with links to every supporting guide in this series
Book a free Technical SEO Health Check with Digital Hothouse to see exactly where your site stands against this checklist – details at the end of this guide.
Why Technical SEO Matters More, Not Less, in an AI-Search World
There’s a lazy assumption doing the rounds in marketing circles: that as AI search grows, the old technical rules matter less, because AI is “smarter” than a traditional crawler and can just figure things out. BrightonSEO 2026 comprehensively debunked this, and the industry data since backs it up.
The most important finding from the conference’s technical sessions, echoed across multiple speakers, is one that reframes the entire brief: Google has spent close to a decade teaching Googlebot to render JavaScript, interpret dynamic content and understand complex page architectures. ChatGPT, Claude, Gemini, Perplexity and the crawlers behind them have not had that decade. If your website relies on JavaScript to render the words, headings, products or pricing that matter, and that content only appears after a script executes in a browser, many AI crawlers simply cannot read it. Your content, from their perspective, doesn’t exist.

This isn’t a fringe concern. Independent research reviewed by Search Engine Land confirms the same pattern our own team saw at BrightonSEO: AI systems now operate through families of distinct crawlers with different purposes and different technical capabilities, and many sites are inadvertently blocking the very bots through which they now want to gain visibility, simply because their robots.txt hasn’t been reviewed in years.
Moz’s own framing of how search engines work is a useful mental model to carry through this entire guide: everything runs through a crawl, index, and rank sequence, and if you fail at the first stage, nothing downstream matters. That was true when it only applied to Google. It’s now true for every AI platform mediating your customers’ first impression of your brand as well.
The reassuring part, and it’s worth holding onto this while working through the rest of this guide: a website that is technically clean, loads quickly, uses structured data properly and has consistent brand signals across the web will perform well for Google. It will also perform well for AI crawlers. It will also be more likely to be cited by AI platforms in their answers. The audience has expanded. The fundamentals remain. What’s changed is the cost of ignoring them.
Crawlability & Indexability Fundamentals (Including AI Bots Like GPTBot)
Crawlability and indexability are two different things, and the distinction matters more than ever now that you’re managing access for multiple, very differently-behaved bot families.
Crawlability is whether a bot can reach a page at all. Indexability is whether, once reached, that page qualifies to be stored and potentially shown. A page can be crawlable but not indexable (a noindex tag, a stray canonical, thin content that trips a quality filter). A page can also fail at the earlier stage entirely: no internal links pointing to it, a blocked resource, or a robots.txt rule that quietly excludes it.
The AI crawler landscape has fragmented, and treating it as one bot is now actively costly
The single biggest mistake we see heading into 2027 is businesses treating “AI crawlers” as one undifferentiated group and either blocking all of them or allowing all of them without understanding the distinction. That’s no longer a defensible default.
The major AI vendors have split their crawlers by function. OpenAI runs GPTBot for model training and a separate bot, OAI-SearchBot, for live ChatGPT search results. Anthropic runs multiple distinct bots for training versus retrieval. Google, Amazon and Apple all follow the same pattern. As Search Engine Land’s crawler guide explains, OpenAI specifically maintains OAI-SearchBot for building a search and citation index, separately from GPTBot, which exists purely to gather training data.
The practical implication: blocking GPTBot in your robots.txt does not remove you from ChatGPT’s live search results. But if you accidentally block OAI-SearchBot along with it, using a blanket “block all OpenAI bots” rule, you do. The old advice to just block or allow “AI bots” as a single category is now more likely to hurt your visibility than protect your content.
Our own BrightonSEO research reinforces this from the diagnostic side. Different AI crawlers behave very differently once they do reach your site: GPTBot, Anthropic’s crawlers, Google’s AI Overview crawler and Perplexity’s crawler each prioritise different pages, explore to different depths, and follow distinct frequency patterns. Understanding which crawlers are actually visiting which pages, something only visible in your server log files, tells you a great deal about where your content is and isn’t being surfaced for AI retrieval.
A pragmatic 2027 robots.txt starting point
- Allow standard search bots (Googlebot, Bingbot) without exception, unless you have a specific reason to exclude a section of the site.
- Decide deliberately on training bots (GPTBot, Google-Extended, ClaudeBot, CCBot) based on your own position on AI training data, not by default.
- Allow retrieval and search bots (OAI-SearchBot, PerplexityBot, and equivalents) if you want to appear in live AI answers, which for most commercial businesses is the visibility that actually drives leads.
- Verify your CDN or firewall agrees with your robots.txt. A block enforced at the WAF or CDN level overrides a permissive robots.txt rule, and the two layers disagreeing is a common, invisible cause of AI invisibility.
- Review, don’t set and forget. This landscape has moved fast over the past two years and will keep moving. A robots.txt reviewed in 2024 is very likely wrong today.
On llms.txt: a clear recommendation
You’ll have seen llms.txt proposed as a quick AI-optimisation win: a markdown file, similar in spirit to robots.txt, intended to guide AI agents to your best content. It’s a reasonable-sounding idea, and it generated real excitement across the industry, including from us. The data doesn’t support prioritising it.

Thomas Peham’s controlled BrightonSEO testing found that only 0.1% of AI bot traffic accessed llms.txt files, with no positive correlation to improved AI visibility. That finding holds up against independent research: Search Engine Land’s own review of the standard notes that it’s still not clear how AI crawlers use the file or what impact it has on training and indexing, and cites a separate practitioner’s tracked test across ten sites that found llms.txt made no measurable difference. It isn’t actively harmful, and it’s cheap to implement on most modern CMSs, so there’s no reason to avoid adding one. But if it’s sitting anywhere near the top of your technical SEO priority list, that’s the wrong call. Structured data, crawlability, entity consistency and content quality are where the evidence says your time belongs.
Core Web Vitals and Page Experience in 2027
There was a period a couple of years ago when Core Web Vitals started to feel like solved, old news. It isn’t. It’s table stakes that both traditional and AI crawlers increasingly filter on before anything else about your site gets evaluated.
Google’s own current guidance, via Search Central, sets out the three metrics and their thresholds clearly, and these figures haven’t moved:
- Largest Contentful Paint (LCP) – measures loading performance. Google’s benchmark: occur within the first 2.5 seconds of the page starting to load.
- Interaction to Next Paint (INP) – measures responsiveness. Google’s benchmark: under 200 milliseconds. INP replaced First Input Delay as the responsiveness metric in March 2024, and it measures the worst interaction latency across the full session, not just the first click.
- Cumulative Layout Shift (CLS) – measures visual stability. Google’s benchmark: a score under 0.1.
Google evaluates all three at the 75th percentile of real visitor data (not lab data, not a single test run), which means consistency across your actual traffic matters more than a single good score in PageSpeed Insights.

Why this now sits underneath your AI visibility too
BrightonSEO’s technical sessions drew a direct line from Core Web Vitals to AI crawl behaviour: AI crawlers deprioritise slow, technically inaccessible pages in much the same way Google’s own crawl-budget prioritisation works. A page that loads slowly, shifts around as it loads, or is sluggish to respond to interaction is a page AI crawlers spend less time on, and potentially don’t index fully. That creates a clear hierarchy for any brand pursuing both traditional rankings and AI citations: fix crawlability and JavaScript rendering first, then make sure the pages you most want cited are technically healthy by Core Web Vitals standards. Everything else – structured data, entity architecture, content optimisation – sits on top of that foundation.
The commercial argument for prioritising this isn’t abstract, either. Ahrefs’ own technical SEO advisor Patrick Stox makes the point bluntly: more crawling doesn’t mean you’ll rank better, but not being crawled means you can’t rank. The same logic extends cleanly to AI citation: a technically healthy page isn’t guaranteed to be cited, but a page an AI crawler struggles to load or parse effectively removes itself from contention before content quality is even assessed.
Where to check your numbers
- Google Search Console – the Core Web Vitals report, based on real Chrome user data (CrUX) over a rolling 28-day window.
- PageSpeed Insights – combines field data with lab data from Lighthouse, and gives specific, actionable fixes per URL.
- Third-party audit tools (Semrush, Ahrefs, Screaming Frog) – useful for site-wide monitoring and flagging regressions before they show up in Search Console.
If you haven’t checked these in the last quarter, you’re flying blind on one of the fundamentals underneath both your traditional rankings and your AI visibility.
Structured Data and Schema Markup: The Signal AI Actually Rewards
This is the technical area where the evidence is most consistently encouraging, because it confirms an investment we’ve recommended to clients for years is paying dividends that extend well beyond traditional rich results.
Google itself is direct about what structured data does and doesn’t do. It is not a direct ranking factor, but it is, in Google’s own words, part of how rich results and Search understand a page’s content and context, and the compounding downstream effects – better entity understanding, richer SERP appearance, and eligibility for AI Overview citation – add up to meaningful organic growth over time.
Thomas Peham’s controlled BrightonSEO testing found structured data produces significant, measurable improvements in Google visibility specifically. Its direct effect on other LLMs is less certain, but the indirect chain – better Google performance leading to more citations, which in turn feeds AI training data – makes it worth implementing regardless of which platform you’re optimising for.

The types worth prioritising
You don’t need to implement all 800-plus Schema.org types. Prioritise:
- Organization schema on your homepage, including accurate sameAs links to every verified profile your brand controls (more on this below).
- Article or BlogPosting schema on content pages, to reinforce authorship and publication data.
- Product, Review and LocalBusiness schema wherever genuinely applicable, since these carry the most AI citation weight for commercial queries.
- FAQPage schema where a page genuinely answers discrete questions, though with a caveat: Google has reduced FAQ rich results in standard search, so don’t add FAQ markup purely to chase a rich result that may no longer appear. It can still support AI-driven answer systems when the content itself is genuinely useful.
Neil Patel’s own SEO framework places structured data firmly inside the technical SEO discipline, alongside site speed, mobile-friendliness and Core Web Vitals, not as a nice-to-have on-page extra, precisely because it’s foundational infrastructure rather than content polish.
Entity architecture: the strategic layer structured data sits inside
Schema markup on your own site is necessary but not sufficient. Sam Davis’s BrightonSEO session on entity architecture argued, backed by data from his work at Yext, that entity-rich structured data across your entire web presence, not just on-site schema, is the single most important technical signal for AI search performance when you look at the full picture.
The concept that crystallised this for our team is the entity confidence score: an informal but genuinely useful way of thinking about how confident an AI model is in its understanding of who your brand is. It’s built from consistency across every platform your brand appears on: the same business name, description, category and contact details on your website, Google Business Profile, LinkedIn, Wikipedia or Wikidata (if you qualify), Trustpilot, and relevant industry directories.

The technical mechanism for connecting these signals explicitly is Schema SameAs mapping, a property that links your website’s structured data directly to your verified profiles across every platform above. It isn’t complex to implement, but the majority of small and mid-market websites still don’t have it in place. It’s now standard on our own implementation checklist for every client.
Paul’s take: “Darko Brzica’s session on links and AI signals showed branded anchor text now consistently outperforming exact-match keyword anchors in link effectiveness. It’s the same underlying shift as entity architecture: the web is moving from recognising keywords to recognising brands. Every technical and link-building decision now needs to reinforce a consistent brand entity, not just target a phrase.”
Validate everything through Google’s Rich Results Test and the Schema Markup Validator after any content update, since errors and warnings creep in quietly, particularly after a CMS or theme change.
Mobile-First Indexing and JavaScript Rendering Pitfalls
Two separate but related technical risks sit in this section, and BrightonSEO’s technical sessions ranked the second one as the single highest-priority issue for any business that cares about AI discoverability specifically.
Mobile-first indexing: still the baseline, still getting missed
Google indexes and ranks based on your site’s mobile version, not desktop. Bruce Clay’s SEO team frames the current state plainly in their 2026 strategic guidance: technical performance and a logical, mobile-first architecture aren’t about pleasing the crawler in isolation; they’re about creating an environment that keeps real visitors engaged, and that same infrastructure is what determines whether AI-driven results can parse your content easily in the first place.
The most common mobile mistakes we still see in our latest client audits:
- Content hidden behind tabs or accordions that isn’t fully accessible to the mobile crawler, or that renders differently than the desktop version.
- Intrusive interstitials that block content on the transition from a mobile search result, which Google has penalised for years and AI crawlers handle no better.
- Structured data present on desktop but missing on mobile, a surprisingly common gap when sites use separate templates.
- Font sizes and tap targets that fail basic mobile usability testing, which increasingly correlates with poor Core Web Vitals scores on real mobile traffic.
JavaScript rendering: the highest-priority AI visibility issue on the table
This is the finding we’d put at the top of any 2027 technical priority list, because multiple independent speakers at BrightonSEO converged on the same point from different angles, and the risk is compounding as more businesses build sites using AI tools that generate JavaScript-heavy front-ends by default.
If your website relies on JavaScript to render its content, meaning text, headings, products or pricing only appear once a script executes in a browser, AI crawlers frequently cannot read it. Google has spent a decade solving this problem for its own crawler and still advises against relying on it. Screaming Frog’s own technical guidance on the issue is direct: Google advises using server-side rendering or pre-rendering rather than a purely client-side approach, because it’s difficult to process JavaScript, and not all search engine crawlers are able to process it successfully or immediately. If that’s still true for Google’s own crawler after a decade of investment, it’s considerably more true for the crawlers behind ChatGPT, Claude and Perplexity.
The diagnostic is simple and free: render your key pages with JavaScript disabled in your browser. If the content disappears, an AI crawler is very likely having the same experience. That’s your starting point for a remediation conversation between your development and SEO teams, whether the fix is server-side rendering, pre-rendering, or a broader architectural change.
Paul’s take: “This came up in multiple sessions and roundtables, not just as a technical problem but as a strategic vulnerability. Businesses that moved to JavaScript-heavy frameworks in the last few years for perfectly good UX reasons now have a hidden technical debt in terms of AI visibility. The audit is the first step, but the remediation conversation needs to happen between development and SEO together, not in isolation.”
Platform-Specific Checklists: Shopify, WordPress, Custom Builds
Most technical SEO principles are platform-agnostic, but each major platform has its own recurring failure patterns worth checking specifically.
Shopify
- URL structure: Shopify’s fixed /products/, /collections/ architecture can create duplicate content across collections. Use canonical tags deliberately rather than relying on defaults.
- App bloat: every installed app adds JavaScript, and Shopify stores are particularly prone to the client-side rendering issues covered in Section 5. Audit installed apps quarterly and remove anything not earning its keep.
- Collection page thinness: collection pages with little unique content beyond a product grid are common casualties of quality filters. Add genuine, useful copy above the fold.
- Page speed: theme and app stacking is the single biggest speed killer on Shopify. Run a Core Web Vitals check after every theme or major app change, not just at launch.
WordPress
- Plugin bloat: the WordPress equivalent of Shopify’s app problem. Every plugin is a potential speed and security liability; audit and consolidate regularly.
- Caching configuration: a properly configured caching plugin is close to mandatory for meeting LCP thresholds on shared or budget hosting.
- SEO plugin conflicts: running more than one SEO plugin (Yoast alongside a theme’s built-in SEO features, for example) frequently produces duplicate or conflicting schema markup. Audit your rendered structured data, not just your plugin settings.
- Image optimisation: WordPress media libraries accumulate unoptimised, oversized images by default. This is consistently one of the fastest wins available on the platform.
Custom builds
- Rendering strategy is the first decision, not an afterthought: decide on server-side rendering, static generation, or a hybrid approach before development starts, with both traditional and AI crawlability in mind, not just user experience.
- Sitemap and robots.txt are developer responsibilities, not marketing afterthoughts: build them into your deployment pipeline so they’re never accidentally left in a staging configuration (a surprisingly common cause of a site vanishing from search entirely after launch).
- Structured data should be templated, not hand-coded per page: build schema generation into your CMS or component library so it stays consistent as the site scales, rather than degrading page by page.
Across all three platforms, the same underlying principle from Sections 2 through 5 applies: the platform doesn’t change what matters; it changes where the risk of getting it wrong hides.
Your 90-Day Technical SEO Action Plan
You don’t need to fix everything simultaneously. Here’s the priority order our own team runs for clients, pulling together the most consistent, high-impact recommendations from BrightonSEO and the wider industry guidance covered in this post.
Days 1–14: Diagnose
- Run the JavaScript rendering test on your key pages (disable JavaScript, see what remains).
- Review your robots.txt against the crawler-by-crawler framework in Section 2, and verify your CDN or firewall settings agree with it.
- Pull your Core Web Vitals report in Google Search Console and flag any pages rated “poor” or “needs improvement.”
Days 15–40: Fix the foundation
- Remediate any JavaScript rendering gaps identified in the diagnosis, prioritising your highest-value commercial pages first.
- Address Core Web Vitals failures, starting with LCP on your top-traffic landing pages.
- Run your structured data through Google’s Rich Results Test and the Schema Markup Validator, and fix any errors or warnings.
Days 41–65: Build the layers on top
- Implement Organization schema with full sameAs mapping across your verified profiles.
- Conduct an entity consistency audit: check your brand name, description and category data across Google Business Profile, LinkedIn, key directories and social profiles, and correct any mismatches.
- Run the relevant platform-specific checklist from Section 6 against your own site.
Days 66–90: Monitor and systemise
- Set up recurring Core Web Vitals and crawl-error monitoring, rather than one-off checks.
- If you have access to server log files, run a basic log file analysis to see which AI crawlers are actually visiting your site, and which pages they prioritise.
- Document the whole process so the next site update, theme change or migration doesn’t quietly undo the work.
Supporting guides in this pillar
This is the first post and the pillar of our Technical SEO Series over the next 12 months. Here is a snapshot of the upcoming posts to keep an eye on – make sure you follow us on LinkedIn so you never miss a post.
- Core Web Vitals in 2027: What’s Changed and What Still Matters
- Crawlability & Indexability: Is Your Site Even Visible to Google and AI Bots?
- llms.txt Explained: Should Your Website Have One?
- Site Speed Audit: A Step-by-Step Guide for Non-Developers
- Mobile-First Indexing: Common Mistakes Costing You Rankings
- Structured Data 101: The Schema Types Every Website Needs
- Technical SEO for Shopify Stores: The Complete Checklist
- Technical SEO for WordPress Sites: The Complete Checklist
- How Broken Internal Links Are Quietly Killing Your SEO
Get a Clear Picture of Where You Actually Stand
Everything in this guide compounds in the same order it’s written: crawlability and indexability first, then Core Web Vitals, then structured data and entity signals, then platform-specific polish. Skipping ahead to schema markup while your JavaScript rendering silently hides your content from AI crawlers is effort spent on the wrong layer.
Book a free Technical SEO Health Check with Digital Hothouse. We’ll run an initial diagnostic crawl of your site and then we can talk though a full diagnostic covered in this guide against your own site, benchmark you against the Core Web Vitals and crawlability standards that matter in 2027, and hand you a prioritised, 90-day action plan, not just a list of errors.

