built to rank

How do I know if AI tools can read my website?

Also answers: What should my website have so AI can cite it?

Four checks, each takes under a minute, and three of them need nothing but a browser. Fetch the page with JavaScript disabled; read your robots.txt; search the HTML for noindex; check your server logs for OAI-SearchBot. If the text disappears with JavaScript off, assume it is not being read.

short answer · 49 words · inspected 10 / 03 / 2026

01section 01

The four checks, in order

Disable JavaScript and reload. If the body copy vanishes, the content exists only after client-side execution. Google documents that it renders JavaScript but that server-side or pre-rendered content is the reliable path, and other retrieval clients are less forgiving than Google, not more.

Open /robots.txt. Read what you are disallowing, and for whom. Google's robots.txt introduction is explicit that the file governs crawling, not indexing — a very common and expensive misunderstanding, because a page blocked from crawling can still appear, stripped of its content.

Search your rendered HTML for noindex. Check the X-Robots-Tag response header too; a header-level noindex is invisible in the markup and does the same thing.

Grep your access logs for the crawler names — OAI-SearchBot, ChatGPT-User, GPTBot — distinguished in OpenAI's crawler reference. Googlebot's variants are listed in the Google crawler overview.

sources

02section 02

The distinction that makes the difference

Blocking training is not blocking search. GPTBot gathers training data; OAI-SearchBot builds search results. A site that added a GPTBot disallow to stay out of model training may be entirely available to the surface that answers customer questions — and a site that disallowed everything removed itself from the surface it wanted to be recommended on.

Decide those two separately, in writing, and record which you chose. It is a two-line file and it is the single highest-leverage thing on this list. For what it is worth, we allow all of them deliberately: the answer library exists to be quoted, so blocking the engines that do the quoting would defeat the strategy.

03section 03

What should my website have so AI can cite it?

Four properties, and the first two are the floor rather than the strategy.

One: retrievable. Crawlable, indexable, and readable without JavaScript execution — the four checks above. Google's AI features documentation states there is no additional technical requirement beyond the ordinary ones, so there is no AI markup to add and no submission form.

Two: specific. An assistant assembling an answer quotes what is quotable. A price, a date, a named standard and a number are liftable; adjectives are not. This is the property that actually separates pages, and it is editorial rather than technical — nothing to cite, because no specification covers it.

Three: one question per page, answered completely. Google's helpful-content guidance is the standard, and it is about the page rather than a publishing cadence.

Four: attributable. A claim with a source attached travels; a claim without one is the assistant's risk to repeat. Cite exact pages rather than homepages, and date what you measured.

What we will not tell you is that doing all four earns citations. Across four assistant surfaces on 2026-10-01 our own site was cited zero times on all 68 questions in this bank, including the twelve already published. We are measuring that hypothesis in public on our own site, and at the last reading it has not paid off.

Think this is your site? Book the inspection.

$1,000, credited to your first month. You get the scored report either way.

inspected · chief engineer
PG 049 / 683 October 2026