# How we continuously monitor SEO + GEO health across 27 checks

**Author:** Cooper Veysey  
**Published:** 2026-08-28T15:42:55.696Z  
**Category:** Search

Our team maintains a 27-point SEO + GEO health checklist designed to catch the technical and structural issues that can affect whether search engines and AI assistants can find, read, understand and cite a site. Every check links to the published documentation that says it matters. As of August 2026, the checklist draws from thirty-four resources published by nineteen organizations.

Sixteen of the twenty-seven checks re-run automatically on the first of every month. The remaining eleven require human judgment, so they are reviewed manually, dated, and signed by the person who made the call. The result is not a one-time SEO audit. It is a continuously updated record of what is healthy, what needs attention, what has not yet been checked, and exactly why.

## Strong GEO starts with strong SEO

**Seventeen rows are about search engines, five are about AI assistants specifically, and five count for both.** GEO — generative engine optimization — is the part aimed at AI assistants: whether one can retrieve, read and cite your pages. SEO is the part aimed at the search results themselves. Most of what serves one serves the other, which is why the two live on one list rather than two.

The five assistant rows exist because those levers are documented only by the assistant vendors, and a Google-sourced checklist misses them entirely. A site can be wide open to Googlebot and invisible to ChatGPT, Claude and Perplexity at the same time, because those crawlers are separate and are controlled separately. When a row concerns assistants, it cites the assistants.

## Every row answers the same five questions

Open any row on the checklist and it tells you five things, in the same order, every time:

- **Standard** — what "good" means for this row, stated plainly enough that you can disagree with it.
- **Why it matters** — what it affects, claiming no more than the source claims.
- **Evidence** — what we actually saw: the URL fetched, and what came back.
- **Checked** — automatically once a month, or by a named person on a named date.
- **Source** — the published resource behind the standard, linked.

The verdict beside each row is one of seven, and the distinctions are the point. **Good** and **Needs attention** are the two you would expect. **Not yet audited** is for a row nobody has checked yet, in its own colour, because a checklist that hides its unchecked rows is a checklist that lies by omission. **Not yet connected** is for a check that is red only because a setup step is unfinished — a new site is not broken because nobody has connected Search Console. **Not enough visitor data yet** is for Core Web Vitals on a site Google has not published enough real-visitor numbers about — we will not grade its lab test as though it were the same measurement. **Could not check** is for a fetch that failed, which our engineering team is prompted to investigate. **Not applicable** is also an option.

## Sixteen checks that re-run themselves, eleven a person signs

Sixteen checks require nothing more than a public web address to execute. No plugin, no tracking script, no access to your CMS or your server. A machine reads your `robots.txt`, finds your sitemap wherever it actually lives, fetches pages that the sitemap lists, reads the homepage's `<head>` and response headers, and inspects the certificate. Then it writes down what it saw, with the URL beside it, and does the same thing again next month.

Running the same question against the same site on a schedule is what separates this from a point-in-time resource that turns stale over time. Sites drift. A plugin update adds a directive nobody chose, a redesign drops a canonical tag, a security setting starts blocking crawlers that were welcome last quarter. 

The eleven person-audited rows split into two kinds, and it is worth knowing which is which before you budget for them.

**Five are set up once and then stay done.** Registering in Bing Webmaster Tools, and verifying Instagram, TikTok, X and YouTube as platform properties in Search Console.

**Six are re-judged as the site changes**, because they are conclusions rather than readings: are there pages nothing links to, is any content "thin" enough to be worth removing, does each practice area have a main page with its supporting articles pointing back at it, has the site been reviewed against WCAG. 

## The sixteen automatic checks

Nine of these cover search engines, four cover AI assistants, and three count for both.

| Check | Our standard, and who publishes it |
| --- | --- |
| Sitemap | `sitemap.xml` exists, parses, and lists the site's pages — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview) |
| robots.txt | Crawlers are not blocked from the site — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/robots/intro) |
| AI assistant crawlers | The crawlers Claude, ChatGPT, Perplexity, Apple, DuckDuckGo and Copilot use to retrieve and cite your site are not blocked in `robots.txt` — [OpenAI](https://developers.openai.com/api/docs/bots) · [Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) · [Perplexity](https://docs.perplexity.ai/guides/bots) · [Apple](https://support.apple.com/en-us/119829) · [DuckDuckGo](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/) · [Google](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) · [Bing](https://www.bing.com/webmasters/help/webmasters-guidelines-30fba23a) |
| HTTPS redirect | Plain `http://` permanently redirects to `https://` — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/301-redirects) |
| One address | The www and bare forms of the domain collapse to one host with a permanent redirect — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls) |
| Canonical tags | The homepage declares a canonical URL on this domain — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/canonicalization) |
| Page titles and descriptions | Every page listed in the sitemap has a `<title>` and a `<meta name="description">`, and neither is left blank — [Google Search Central](https://developers.google.com/search/docs/appearance/title-link) · [Google Search Central](https://developers.google.com/search/docs/appearance/snippet) |
| Structured data | The homepage carries JSON-LD that parses and has its fields filled in — not a blank template — [Google Search Central](https://developers.google.com/search/docs/appearance/structured-data/local-business) |
| Organization schema and entity links | The homepage identifies the business as an Organization in JSON-LD, and every profile URL that block lists under `sameAs` still resolves — [Google Search Central](https://developers.google.com/search/docs/appearance/structured-data/organization) · [schema.org](https://schema.org/sameAs) |
| Link previews | The homepage carries `og:title` and `og:image`, so a link to it renders as a preview card — [The Open Graph protocol](https://ogp.me/) |
| AI citation directives | No `noarchive`, `nocache` or `nosnippet` directive is suppressing the site in AI answers — [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmasters-guidelines-30fba23a) · [Apple](https://support.apple.com/en-us/119829) |
| SSL certificate | The certificate is valid and not close to expiry — [Google Search Console Help](https://support.google.com/webmasters/answer/11396518) |
| Core Web Vitals | Google's measurements from real visitors to the homepage on a phone meet its published thresholds: LCP 2.5 seconds or less, INP 200 milliseconds or less, CLS 0.1 or less, at the 75th percentile. Where Google has not published enough real-visitor data for the site, the row says so rather than grading a lab test as though it were the same thing — [Google Search Central](https://developers.google.com/search/docs/appearance/page-experience) · [web.dev](https://web.dev/articles/vitals) |
| llms.txt | Reported, never graded: whether a guide for AI assistants exists at `/llms.txt` — [Google Search Central](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) · [OpenAI](https://developers.openai.com/api/docs/bots) |
| Agent capabilities | Reported, never graded: whether the site publishes a machine-callable capability surface — an API catalogue or an MCP server card — that an assistant can discover — [IETF RFC 9727](https://www.rfc-editor.org/rfc/rfc9727.html) · [Cloudflare](https://blog.cloudflare.com/agent-readiness/) |
| Search Console | Google Search Console is connected and answering queries — [Google Search Console Help](https://support.google.com/webmasters/answer/9128668) |

## The five you set up once

| Check | Our standard, and who publishes it |
| --- | --- |
| Bing Webmaster | The site is registered and verified in Bing Webmaster Tools — [Microsoft Bing Webmaster](https://blogs.bing.com/webmaster/June-2025/Start-Using-Bing-Webmaster-Tools-to-Improve-Your-Site-Visibility) |
| Instagram in Search | The firm's Instagram account is verified as a platform property in Search Console — [Google Search Console Help](https://support.google.com/webmasters/answer/17148418) |
| TikTok in Search | The firm's TikTok account is verified as a platform property in Search Console — [Google Search Console Help](https://support.google.com/webmasters/answer/17148418) |
| X in Search | The firm's X account is verified as a platform property in Search Console — [Google Search Console Help](https://support.google.com/webmasters/answer/17148418) |
| YouTube in Search | The firm's YouTube channel is verified as a platform property in Search Console — [Google Search Console Help](https://support.google.com/webmasters/answer/17148418) |

Google's platform properties let you see how content you publish on those accounts performs in Google Search, alongside News and Discover. The setup is verifying the account as its own Search Console property, and the row is marked not applicable where a business does not post on that platform. 

## The six a person re-judges

| Check | Our standard, and who publishes it |
| --- | --- |
| Internal links | Every page worth finding was linked from at least one other page at the last audit, and the pages that win work are reachable in a few clicks from the homepage — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/links-crawlable#internal-links) · [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmasters-guidelines-30fba23a) · [Search Engine Journal, reporting Google's John Mueller](https://www.searchenginejournal.com/google-click-depth-matters-seo-url-structure/256779/) |
| Thin content | No empty or low-value pages stood at the last audit — [Google Search Central](https://developers.google.com/search/docs/fundamentals/creating-helpful-content#content-and-quality-questions) |
| Practice-area hubs | Each practice area has a main page, and its related articles link back to it — [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/links-crawlable#internal-links) |
| Accessibility | The site was reviewed against WCAG, and anything that would stop a screen-reader user completing an enquiry was fixed — [U.S. Department of Justice](https://www.ada.gov/resources/web-guidance/) · [W3C Web Accessibility Initiative](https://www.w3.org/WAI/standards-guidelines/wcag/) |
| Brave Search | The site's pages are turning up in Brave, and its submit page has been used for anything recently changed — [Brave Search](https://search.brave.com/help/brave-search-crawler) |
| Markdown copies of pages | Reported, never graded: whether key pages are also served as clean Markdown, so an assistant reading them gets the words without the page furniture — [The llms.txt proposal](https://llmstxt.org/) · [Google Search Central](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) |

Brave is on the list because it costs almost nothing to satisfy. Brave states that where Googlebot cannot crawl a page, its own crawler will not either — so the work already done for Google carries over, and there is no account to open, nothing to verify and no second console to keep. Its one lever is a submit page for asking it to re-fetch something you have changed. It matters more than its size suggests, because its index feeds answers beyond its own search box. [Brave documents its crawler and that submit page](https://search.brave.com/help/brave-search-crawler).

## The nineteen organizations behind the list

The checklist is written from published documentation rather than opinion, and no row claims more than its source claims. These are the organizations, and the tables above link the specific resource each row rests on:

**Search engines** — [Google Search Central](https://developers.google.com/search/docs/fundamentals/seo-starter-guide), [Google Search Console Help](https://support.google.com/webmasters/answer/9128668), [web.dev](https://web.dev/articles/vitals), [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmasters-guidelines-30fba23a), [Microsoft's Bing Webmaster blog](https://blogs.bing.com/webmaster/June-2025/Start-Using-Bing-Webmaster-Tools-to-Improve-Your-Site-Visibility) and [Brave Search](https://search.brave.com/help/brave-search-crawler).

**AI assistants, in their own words** — [OpenAI](https://developers.openai.com/api/docs/bots), [Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), [Perplexity](https://docs.perplexity.ai/guides/bots), [Apple](https://support.apple.com/en-us/119829) and [DuckDuckGo](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/).

**Standards bodies and specifications** — [schema.org](https://schema.org/sameAs), [the Open Graph protocol](https://ogp.me/), [IETF RFC 9727](https://www.rfc-editor.org/rfc/rfc9727.html), [the W3C Web Accessibility Initiative](https://www.w3.org/WAI/standards-guidelines/wcag/) and [the llms.txt proposal](https://llmstxt.org/).

**Everyone else** — [the U.S. Department of Justice](https://www.ada.gov/resources/web-guidance/) on web accessibility and the ADA, [Cloudflare](https://blog.cloudflare.com/agent-readiness/) on how many sites are ready for agents, and [Search Engine Journal](https://www.searchenginejournal.com/google-click-depth-matters-seo-url-structure/256779/) reporting Google's John Mueller on click depth.

## Two problems nothing warns you about

Both of the faults below leave the rest of a site looking fine: the pages load, the build passes, and no editor says a word. Finding them is the practical reason to write a checklist from the documentation rather than from habit.

**A site can be crawlable and still uncitable.** A `noarchive`, `nocache` or `nosnippet` directive leaves crawling completely intact — so every other check reads clean — while suppressing the site in AI answers. Two vendors document that independently, which is what lifts it from a plausible worry to a checkable fact. Microsoft states that "Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers," and that content tagged NOCACHE "may be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer." [Bing's Webmaster blog states both](https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat). Apple names the same family of lever for its own answers: "Web publishers can opt out of their content being used in these broad world knowledge answers by applying the nosnippet meta tag to specific content." [Apple documents that on its Applebot page](https://support.apple.com/en-us/119829). A CMS, a plugin or a well-meaning "stop AI scraping" setting can add any of these without anyone deciding to, so we check both places they can hide: the page's meta robots tag and the `X-Robots-Tag` response header.

The same logic put page titles and descriptions on the list. Nothing warns you when they are blank — no editor warning, no red mark, no failed build. The page goes live looking normal to whoever wrote it, and the only place the omission shows up is a search result nobody at the firm is looking at. On one site we reviewed, nearly every published article had both fields empty. Not one of those was a decision; it was the default, repeated post after post. Google says a good title link makes a result more relevant to searchers, and that "Snippets are primarily created from the page content itself" — so a blank description does not remove the page, it hands the wording to a machine. [Google documents title links](https://developers.google.com/search/docs/appearance/title-link) and [how snippets are created](https://developers.google.com/search/docs/appearance/snippet). That is the shape of defect a monthly machine check earns its keep on, and the shape a human audit risks missing, because catching it means opening seventy-seven pages to read their `<head>`.

## What we can responsibly claim

None of this is a ranking claim, and we do not sell it as one. Google is explicit that "Google does not guarantee that features that consume structured data will show up in search results," and the same restraint applies across the list. [Google states that in its Organization documentation](https://developers.google.com/search/docs/appearance/structured-data/organization).

What a sourced, dated, re-running checklist does buy is narrower and more durable. You can see what was checked and what was not. You can see when, and you can see what each check actually means. You can see who says it matters, and quickly and easily open the source document yourself. And where a row is red, you can tell at a glance whether that is a fault on the site, an unfinished setup step, a fetch that failed on our side, or a measurement Google has not published enough visitor data for us to make.

The list is twenty-seven items today, and it grows when a source justifies a new row.

## Ask for a checklist run on your site

Want to know which of these your site currently passes? [Let us know](mailto:cooper@veyseysoftwaresolutions.com). We can run the sixteen machine-checkable rows against your public URLs and send back each verdict with the URL we fetched, what came back, and the citation beside it.

## Thank you for your time

If you have any questions or want to connect on anything that I wrote about above, please [email me](mailto:cooper@veyseysoftwaresolutions.com) or [book some time on my calendar](https://calendar.google.com/calendar/appointments/schedules/AcZssZ041TFFhKz0LzQydfOipmZ3l8K-22DadOWnsRMPnLkS8TjaDIreyZ8fFxAimnDxsCeimFSvnY_3). Any and all feedback is of course so appreciated.
