Will AI engines cite your site?
SEO gets you listed. GEO gets you quoted.
willaicite is a free, deterministic audit that scores whether AI answer engines (ChatGPT, Perplexity, Google AI Overviews and Claude) can retrieve your page, read it without running your JavaScript, and quote it in their answers.
Search engines rank you in a list of links answer engines quote you willaicite measures whether they can
free · deterministic · results in about 30 seconds
Five checks SEO tools skip
SEO tools audit the path to a ranking. willaicite audits the path to a citation. These five checks are where the two paths differ.
-
The firewall block robots.txt hides
willaicite fetches your page twice, once as an ordinary browser and once identifying as an AI crawler, then compares what comes back. A firewall rule like Cloudflare’s one-click “block AI bots” toggle returns 403 to the crawler while robots.txt keeps saying allowed. SEO crawlers never make that second fetch, so they report clean while the engine silently drops your page.
check result meaning robots.txt, GPTBot allowed what every SEO tool reads fetch as browser 200 OK what your monitoring sees fetch as GPTBot 403 Forbidden what the engine actually gets -
Pages that need JavaScript arrive empty
Googlebot renders JavaScript; so does Lighthouse. GPTBot, ClaudeBot and PerplexityBot do not; they see whatever the server sends and nothing more. A page built in the browser can audit perfectly in SEO tools and still reach the answer engine as an empty shell. willaicite scores how much of your content survives in the raw HTML, reading the page exactly as those crawlers do.
browser · JS executedAI crawler · raw HTML<div id="root"></div> -
Citing sources lifts AI visibility
Old SEO advice says external links leak PageRank, so pages hoard authority and cite no one. The GEO study (Aggarwal et al., KDD 2024) measured the opposite effect on AI visibility: citing sources lifted it 24.9 percent, statistics 25.9 percent, and quotations 27.8 percent. Replications in 2025 and 2026, including a 252,000-trial controlled study, shrank those effects to a tie-break among already-retrieved pages. The direction held, and it still means optimizing for the ranking can cost you the quote.
citing sources+24.9%statistics+25.9%quotations+27.8%Visibility lift in generative answers (Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024), measured with the source already retrieved. Later replications find smaller, conditional effects; directional, not a guarantee. -
Which blocks cost you citations
SEO tools check whether Googlebot is allowed. willaicite runs that check for every AI crawler, and tells the two kinds apart: retrieval crawlers fetch pages to build answers, so blocking one means that engine can never cite you; training crawlers only collect model-training data, so blocking them costs no citations. The full registry documents every crawler willaicite checks.
crawler token tier consequence when blocked OAI-SearchBot retrieval blocked → ChatGPT can never cite you Claude-SearchBot retrieval blocked → Claude can never cite you PerplexityBot retrieval blocked → Perplexity can never cite you GPTBot · ClaudeBot · CCBot training blocking is a policy choice, not a citation bug -
Is there an answer to quote?
Engines quote a chunk word-for-word, so willaicite checks whether your page contains one: a self-contained answer near the top, headings phrased as the questions people actually ask, an FAQ an engine can lift whole. A page that opens with a sentence like
willaicite is a free audit that scores citation readiness
hands the engine its quote; a page that opens with a hero animation hands it nothing. You can rank first for a query and still contain nothing an engine can use.- ✓self-contained definition in the first screenful: “X is …”
- ✓headings phrased as questions users ask
- ✗no FAQ: nothing an engine can lift whole
The nine dimensions
Every audit scores the same nine dimensions. A dimension weighs more when it more often decides whether a citation happens, and the weights follow the 2025–2026 controlled studies rather than the 2024 headline numbers.
- 01 AI crawler access wt high Can AI crawlers fetch the page at all. Checked one crawler at a time, in robots.txt and at the firewall.
- 02 Renderability wt high How much of the content is in the raw HTML your server sends, before any JavaScript runs.
- 03 Answer-readiness wt high Is there a short, complete answer near the top that an engine can quote whole, under headings phrased as questions.
- 04 Topical focus & metadata wt high Do the title, headings, description and body all present the same one topic, so a retriever can match the page to a query.
- 05 Evidence density wt medium Statistics, quotations and cited sources: the material engines prefer to quote once a page is retrieved.
- 06 Structured data wt medium Schema.org markup that tells engines who wrote the page and when. Matters most for Google AI Overviews.
- 07 Freshness wt high A date a reader can see and a date a machine can read. A recent timestamp is one of the few factors that consistently lifted citation odds in controlled 2026 testing.
- 08 Entity & E-E-A-T wt medium A named author and organization an engine can credit the page to.
- 09 SEO foundation wt low A canonical URL that points at itself, plus unique titles and descriptions, so a citation lands on the right page.
Frequently asked questions
How is GEO different from SEO?
SEO asks where you rank in a list of links. GEO asks whether the answer quotes you. To quote you, an answer engine must fetch your page with its own crawler, read it without running JavaScript, and find a passage worth quoting word-for-word. No ranking audit tests those three steps.
Does blocking GPTBot stop ChatGPT from citing my page?
No. GPTBot is OpenAI’s training crawler; ChatGPT search retrieves through OAI-SearchBot and ChatGPT-User. Blocking GPTBot is a legitimate policy choice about model training. Blocking OAI-SearchBot removes you from ChatGPT’s citations entirely. willaicite reports each token separately so you can tell the two decisions apart.
Why does my JavaScript-rendered site score low?
AI crawlers read the raw HTML your server sends; they do not run JavaScript. If your content only appears after a framework boots in the browser, those crawlers see an empty shell. Server-render or prerender everything you want an engine to quote.
How fresh does my content need to be?
Answer engines cite pages updated within the last 90 days far more readily. A 252,000-trial controlled study in 2026 found a recent timestamp is one of the few content factors that consistently lifts citation odds across engines. Keep a visible date and a machine-readable one, update them only with genuine edits; willaicite reads both and flags undatable pages.
Does willaicite check SEO too?
No. willaicite audits the citation path, which is GEO. Several signals it checks (crawler access, structured data, visible dates, clear headings, authorship) are SEO fundamentals too, so fixes often help both. It does not measure rankings, keywords, backlinks or page speed; SEO tools already cover that ground.
What does an audit cost, and what happens to my data?
Nothing; willaicite is free. An audit makes at most ten requests to your site and typically finishes in under 30 seconds. It is deterministic: no LLM calls, no sampling, and the same page gets the same score.
Method
How the score is made
willaicite is deterministic: no LLM calls, no sampling, and every point is tied to evidence you can check in the report. An audit makes at most ten polite requests (your page twice, robots.txt, the sitemap and a few supporting files) and honors robots.txt, while telling you what robots.txt won’t.
Audit your site