Get started
SEOSEOAI agentsQuoxSEO

Can AI read your website? We checked ours and found 121 blank pages

Adam Cowles2026-09-23T22:00:00.000Z5 min read
A page rendering fully in one layer while a second, thinner layer beneath it stays almost empty, cyan and violet over near-black

Ask an AI search engine about your own site and there is a decent chance it is reading less of it than you think. ChatGPT, Claude, Perplexity and Google's AI Overviews mostly do not run JavaScript when they fetch a page. They read the raw HTML the server hands back on the first request.

If your content only appears after a script runs and finishes fetching, a browser sees a full page and a crawler sees almost nothing. We build a tool that checks exactly this, so this week we pointed it at ourselves.

What AI crawlers actually see

QuoxSEO, our self-hosted SEO warehouse, diffs the raw HTML a page returns against the fully rendered DOM a browser builds, for every URL on a site. That diff is the practical core of what the industry calls generative engine optimisation, or GEO: not guessing what a model will say about your site, but checking whether the model can even read it in the first place.

Run against quox.ai, it flagged 121 pages returning close to nothing to a non-JS reader: our own blog's tag and category archive pages, things like /blog/tag/access-control.

121 pages, serving almost nothing

The cause was ordinary, which is exactly why it is worth writing down. Those archive pages are dynamic routes, and our build's static prerender step, which bakes real HTML ahead of time for most of the site, skips routes it cannot enumerate in advance. On top of that, the pages fetched their post list from our own API after the component mounted, inside a useEffect.

A browser runs the script, waits for the fetch, and renders a full grid of posts. Anything that only reads the first HTTP response saw a heading, a loading state, and nothing else. /blog/tag/access-control was baking zero raw words of body content. Multiply that by every tag and category on the blog and it was 121 pages, not one.

The fix, not the redesign

We changed how the pages get their first content, not what they look like. Two changes:

First, both page types now compute their initial post list synchronously, from a static index already built at deploy time, instead of waiting on a fetch. The live API call still runs straight after, so the numbers stay current for a human browsing the page, but the first response already carries real content.

Second, the prerenderer now enumerates every tag and category slug in that index and bakes a page for each one. All 121 taxonomy pages bake with their actual posts now. /blog/tag/access-control went from 0 raw words to 2,091.

Then we added a backstop: the build now counts the visible words in every baked page and fails outright if any of them drops under 50. It is a deliberately blunt trip-wire, built so this exact failure cannot happen again quietly, the next time someone adds a page that looks fine in a browser and ships empty everywhere else.

Check whether AI can read your website

You do not need our tool for the basic version of this check. Fetch a page the way a non-JS crawler would and read what comes back, stripping the tags out to see the actual words:

curl -s https://yoursite.com/some-page | sed 's/<[^>]*>/ /g'

If that returns a heading and little else, and the same page looks complete in a browser, the gap is real. Check a page type you have many of (tag pages, category pages, paginated listings, anything that loads its content client-side) rather than just your homepage. That is where this pattern hides, because homepages are usually the one page everyone remembers to prerender.

What this does and does not prove

This is a content-accessibility check, not a ranking promise. It tells you whether an AI crawler or a non-JS fetch can read the words on a page. It does not tell you whether ChatGPT will cite you, or where Google will place you, and we are not claiming it does.

Whether a model chooses to use content it can read is a separate, harder problem, and we would rather say that plainly than let AI SEO become another word for guesswork dressed up as certainty.

What we can say is narrower and checkable: point a tool, ours or a five-line curl command, at your own site before you assume an AI crawler sees what you see. A build gate that fails loudly beats a claim nobody checked, which is the same pattern behind how we try to run evidence generally: trust the check, not the assumption.

We found our own gap by running our own product against our own site. That is the whole method, and it works because it is boring enough to actually do.

We publish the report itself: our own witnessed QuoxSEO crawl, verifiable start to finish.