August 28, 2026

How we continuously monitor SEO + GEO health across 27 checks

How we continuously monitor SEO + GEO health across 27 checks

Our team maintains a 27-point SEO + GEO health checklist designed to catch the technical and structural issues that can affect whether search engines and AI assistants can find, read, understand and cite a site. Every check links to the published documentation that says it matters. As of August 2026, the checklist draws from thirty-four resources published by nineteen organizations.

Sixteen of the twenty-seven checks re-run automatically on the first of every month. The remaining eleven require human judgment, so they are reviewed manually, dated, and signed by the person who made the call. The result is not a one-time SEO audit. It is a continuously updated record of what is healthy, what needs attention, what has not yet been checked, and exactly why.

Strong GEO starts with strong SEO

Seventeen rows are about search engines, five are about AI assistants specifically, and five count for both. GEO — generative engine optimization — is the part aimed at AI assistants: whether one can retrieve, read and cite your pages. SEO is the part aimed at the search results themselves. Most of what serves one serves the other, which is why the two live on one list rather than two.

The five assistant rows exist because those levers are documented only by the assistant vendors, and a Google-sourced checklist misses them entirely. A site can be wide open to Googlebot and invisible to ChatGPT, Claude and Perplexity at the same time, because those crawlers are separate and are controlled separately. When a row concerns assistants, it cites the assistants.

Every row answers the same five questions

Open any row on the checklist and it tells you five things, in the same order, every time:

  • Standard — what "good" means for this row, stated plainly enough that you can disagree with it.
  • Why it matters — what it affects, claiming no more than the source claims.
  • Evidence — what we actually saw: the URL fetched, and what came back.
  • Checked — automatically once a month, or by a named person on a named date.
  • Source — the published resource behind the standard, linked.

The verdict beside each row is one of seven, and the distinctions are the point. Good and Needs attention are the two you would expect. Not yet audited is for a row nobody has checked yet, in its own colour, because a checklist that hides its unchecked rows is a checklist that lies by omission. Not yet connected is for a check that is red only because a setup step is unfinished — a new site is not broken because nobody has connected Search Console. Not enough visitor data yet is for Core Web Vitals on a site Google has not published enough real-visitor numbers about — we will not grade its lab test as though it were the same measurement. Could not check is for a fetch that failed, which our engineering team is prompted to investigate. Not applicable is also an option.

Sixteen checks that re-run themselves, eleven a person signs

Sixteen checks require nothing more than a public web address to execute. No plugin, no tracking script, no access to your CMS or your server. A machine reads your robots.txt, finds your sitemap wherever it actually lives, fetches pages that the sitemap lists, reads the homepage's <head> and response headers, and inspects the certificate. Then it writes down what it saw, with the URL beside it, and does the same thing again next month.

Running the same question against the same site on a schedule is what separates this from a point-in-time resource that turns stale over time. Sites drift. A plugin update adds a directive nobody chose, a redesign drops a canonical tag, a security setting starts blocking crawlers that were welcome last quarter.

The eleven person-audited rows split into two kinds, and it is worth knowing which is which before you budget for them.

Five are set up once and then stay done. Registering in Bing Webmaster Tools, and verifying Instagram, TikTok, X and YouTube as platform properties in Search Console.

Six are re-judged as the site changes, because they are conclusions rather than readings: are there pages nothing links to, is any content "thin" enough to be worth removing, does each practice area have a main page with its supporting articles pointing back at it, has the site been reviewed against WCAG.

The sixteen automatic checks

Nine of these cover search engines, four cover AI assistants, and three count for both.

Check Our standard, and who publishes it
Sitemap sitemap.xml exists, parses, and lists the site's pages — Google Search Central
robots.txt Crawlers are not blocked from the site — Google Search Central
AI assistant crawlers The crawlers Claude, ChatGPT, Perplexity, Apple, DuckDuckGo and Copilot use to retrieve and cite your site are not blocked in robots.txt — OpenAI · Anthropic · Perplexity · Apple · DuckDuckGo · Google · Bing
HTTPS redirect Plain http:// permanently redirects to https:// — Google Search Central
One address The www and bare forms of the domain collapse to one host with a permanent redirect — Google Search Central
Canonical tags The homepage declares a canonical URL on this domain — Google Search Central
Page titles and descriptions Every page listed in the sitemap has a <title> and a <meta name="description">, and neither is left blank — Google Search Central · Google Search Central
Structured data The homepage carries JSON-LD that parses and has its fields filled in — not a blank template — Google Search Central
Organization schema and entity links The homepage identifies the business as an Organization in JSON-LD, and every profile URL that block lists under sameAs still resolves — Google Search Central · schema.org
Link previews The homepage carries og:title and og:image, so a link to it renders as a preview card — The Open Graph protocol
AI citation directives No noarchive, nocache or nosnippet directive is suppressing the site in AI answers — Bing Webmaster Guidelines · Apple
SSL certificate The certificate is valid and not close to expiry — Google Search Console Help
Core Web Vitals Google's measurements from real visitors to the homepage on a phone meet its published thresholds: LCP 2.5 seconds or less, INP 200 milliseconds or less, CLS 0.1 or less, at the 75th percentile. Where Google has not published enough real-visitor data for the site, the row says so rather than grading a lab test as though it were the same thing — Google Search Central · web.dev
llms.txt Reported, never graded: whether a guide for AI assistants exists at /llms.txt — Google Search Central · OpenAI
Agent capabilities Reported, never graded: whether the site publishes a machine-callable capability surface — an API catalogue or an MCP server card — that an assistant can discover — IETF RFC 9727 · Cloudflare
Search Console Google Search Console is connected and answering queries — Google Search Console Help

The five you set up once

Check Our standard, and who publishes it
Bing Webmaster The site is registered and verified in Bing Webmaster Tools — Microsoft Bing Webmaster
Instagram in Search The firm's Instagram account is verified as a platform property in Search Console — Google Search Console Help
TikTok in Search The firm's TikTok account is verified as a platform property in Search Console — Google Search Console Help
X in Search The firm's X account is verified as a platform property in Search Console — Google Search Console Help
YouTube in Search The firm's YouTube channel is verified as a platform property in Search Console — Google Search Console Help

Google's platform properties let you see how content you publish on those accounts performs in Google Search, alongside News and Discover. The setup is verifying the account as its own Search Console property, and the row is marked not applicable where a business does not post on that platform.

The six a person re-judges

Check Our standard, and who publishes it
Internal links Every page worth finding was linked from at least one other page at the last audit, and the pages that win work are reachable in a few clicks from the homepage — Google Search Central · Bing Webmaster Guidelines · Search Engine Journal, reporting Google's John Mueller
Thin content No empty or low-value pages stood at the last audit — Google Search Central
Practice-area hubs Each practice area has a main page, and its related articles link back to it — Google Search Central
Accessibility The site was reviewed against WCAG, and anything that would stop a screen-reader user completing an enquiry was fixed — U.S. Department of Justice · W3C Web Accessibility Initiative
Brave Search The site's pages are turning up in Brave, and its submit page has been used for anything recently changed — Brave Search
Markdown copies of pages Reported, never graded: whether key pages are also served as clean Markdown, so an assistant reading them gets the words without the page furniture — The llms.txt proposal · Google Search Central

Brave is on the list because it costs almost nothing to satisfy. Brave states that where Googlebot cannot crawl a page, its own crawler will not either — so the work already done for Google carries over, and there is no account to open, nothing to verify and no second console to keep. Its one lever is a submit page for asking it to re-fetch something you have changed. It matters more than its size suggests, because its index feeds answers beyond its own search box. Brave documents its crawler and that submit page.

The nineteen organizations behind the list

The checklist is written from published documentation rather than opinion, and no row claims more than its source claims. These are the organizations, and the tables above link the specific resource each row rests on:

Search engines — Google Search Central, Google Search Console Help, web.dev, Bing Webmaster Guidelines, Microsoft's Bing Webmaster blog and Brave Search.

AI assistants, in their own words — OpenAI, Anthropic, Perplexity, Apple and DuckDuckGo.

Standards bodies and specifications — schema.org, the Open Graph protocol, IETF RFC 9727, the W3C Web Accessibility Initiative and the llms.txt proposal.

Everyone else — the U.S. Department of Justice on web accessibility and the ADA, Cloudflare on how many sites are ready for agents, and Search Engine Journal reporting Google's John Mueller on click depth.

Two problems nothing warns you about

Both of the faults below leave the rest of a site looking fine: the pages load, the build passes, and no editor says a word. Finding them is the practical reason to write a checklist from the documentation rather than from habit.

A site can be crawlable and still uncitable. A noarchive, nocache or nosnippet directive leaves crawling completely intact — so every other check reads clean — while suppressing the site in AI answers. Two vendors document that independently, which is what lifts it from a plausible worry to a checkable fact. Microsoft states that "Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers," and that content tagged NOCACHE "may be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer." Bing's Webmaster blog states both. Apple names the same family of lever for its own answers: "Web publishers can opt out of their content being used in these broad world knowledge answers by applying the nosnippet meta tag to specific content." Apple documents that on its Applebot page. A CMS, a plugin or a well-meaning "stop AI scraping" setting can add any of these without anyone deciding to, so we check both places they can hide: the page's meta robots tag and the X-Robots-Tag response header.

The same logic put page titles and descriptions on the list. Nothing warns you when they are blank — no editor warning, no red mark, no failed build. The page goes live looking normal to whoever wrote it, and the only place the omission shows up is a search result nobody at the firm is looking at. On one site we reviewed, nearly every published article had both fields empty. Not one of those was a decision; it was the default, repeated post after post. Google says a good title link makes a result more relevant to searchers, and that "Snippets are primarily created from the page content itself" — so a blank description does not remove the page, it hands the wording to a machine. Google documents title links and how snippets are created. That is the shape of defect a monthly machine check earns its keep on, and the shape a human audit risks missing, because catching it means opening seventy-seven pages to read their <head>.

What we can responsibly claim

None of this is a ranking claim, and we do not sell it as one. Google is explicit that "Google does not guarantee that features that consume structured data will show up in search results," and the same restraint applies across the list. Google states that in its Organization documentation.

What a sourced, dated, re-running checklist does buy is narrower and more durable. You can see what was checked and what was not. You can see when, and you can see what each check actually means. You can see who says it matters, and quickly and easily open the source document yourself. And where a row is red, you can tell at a glance whether that is a fault on the site, an unfinished setup step, a fetch that failed on our side, or a measurement Google has not published enough visitor data for us to make.

The list is twenty-seven items today, and it grows when a source justifies a new row.

Ask for a checklist run on your site

Want to know which of these your site currently passes? Let us know. We can run the sixteen machine-checkable rows against your public URLs and send back each verdict with the URL we fetched, what came back, and the citation beside it.

Thank you for your time

If you have any questions or want to connect on anything that I wrote about above, please email me or book some time on my calendar. Any and all feedback is of course so appreciated.