The firewall block is gone, and on the fullest available measure Siteimprove is cited in AI answers nearly three times as often as its two closest competitors combined. The exposure that remains is subtler: most of those citations are for generic SEO content, and the machines reading the site are being told an out-of-date story about what it is.
The headline finding from May is closed. Bingbot, MSNBot and BingPreview now return 200 on both the homepage and /robots.txt, confirmed on two independent passes. The WAF rule that made ChatGPT and Microsoft Copilot structurally unable to see siteimprove.com has been removed.
And the citation picture is better than May implied. On the fullest measure available — citations across every AI surface — Siteimprove is cited 3,419 times across 325 pages. That is nearly three times its two closest competitors, Level Access and Deque, combined. The "structurally absent" framing this audit inherited from May does not survive contact with the peer data.
What survives is narrower, and more fixable. Roughly 85% of that citation volume comes from one place — Google's AI Overviews pulling keyword-level content, mostly generic SEO-education pages. That is the content least tied to what Siteimprove sells and most exposed to being answered without a click. Meanwhile the site's own metadata and server configuration feed machines an out-of-date, in places factually wrong, account of the company. The constraint was never access. It is what the machines conclude once they arrive.
Total AI citations, Siteimprove against its direct competitors and two adjacent SEO-content brands, measured the same day (Ahrefs, 1 September 2026). The comparison is the point: a citation count means nothing until you know what a competitor's looks like.
| Domain | What it is | Total AI citations | ChatGPT | Copilot | Relative total |
|---|---|---|---|---|---|
| siteimprove.com | Accessibility / governance | 3,419 | 32 | 2 | |
| levelaccess.com | Accessibility (direct) | 722 | 4 | 4 | |
| deque.com | Accessibility (direct) | 430 | 12 | 1 | |
| semrush.com | SEO tooling (adjacent) | 89,194 | 6,356 | 3,271 | |
| ahrefs.com | SEO tooling (adjacent) | 28,040 | 5,471 | 1,328 |
Two things read off this at once. Among the accessibility vendors it actually competes with, Siteimprove is the most-cited by a wide margin — nearly 3× Level Access and Deque combined. That is the opposite of the "structurally absent" story. And ChatGPT and Copilot are thin for Siteimprove (32 and 2) — but thin for every accessibility vendor here, while the SEO-content brands post thousands. That pattern is a property of the category and of content volume, not a lingering symptom of the old block. The right conclusion is not "the door opened and nobody came" — it is that the Bing-fed engines are an open category race that no accessibility vendor has yet won.
The 3,419 total is not evenly spread. It is dominated by one Google surface pulling one kind of content.
| AI surface | Citations | Share of total |
|---|---|---|
| Google AI Overviews (keyword-level) | 2,903 | 85% |
| Google AI Mode | 157 | 5% |
| Perplexity | 150 | 4% |
| Google AI Overviews (page-level) | 100 | 3% |
| Gemini | 75 | 2% |
| ChatGPT / Copilot / Grok | 34 | 1% |
Around 90% of Siteimprove's AI-citation footprint is Google AI surfaces citing its glossary and blog — the generic SEO-education content in the next section. That is simultaneously the good news (real, substantial presence) and the exposure (it rests on the content most likely to be answered without a click, and least connected to what the company sells).
An AI engine arriving at /platform/seo/aeo-visibility/ — the page that sells answer-engine optimisation — finds 525 words, no product schema, no FAQ markup, and no entity anchor. It is the least-described page in the audit for a machine trying to identify what it is, and it is the one the category argument depends on.
FAQ schema is already live on /platform/seo/, /toolkit/accessibility-checker/ and /toolkit/color-contrast-checker/. The template supports it. Its absence on the AEO page and both flagship checkers is an unfilled CMS field, not an engineering gap.
The citations in the previous section rest on Siteimprove's organic content, and that content sits in a different world from the product. Ranked by shared keywords, its closest organic competitors are SEO publishers, not accessibility vendors:
| Nearest organic competitors | What they are | Their monthly traffic |
|---|---|---|
| semrush.com | SEO tooling | 3,685,655 |
| ahrefs.com | SEO tooling | 4,105,955 |
| hubspot.com | Marketing platform | 1,910,471 |
| moz.com | SEO publisher | 560,505 |
| neilpatel.com | SEO publisher | 346,882 |
| conductor.com | SEO platform | 122,087 |
| searchengineland.com | SEO trade press | 105,656 |
One caveat on this metric, stated plainly: keyword-overlap competitors are whoever publishes similar-ranking content, so a site whose traffic is SEO-education articles will surface SEO publishers by construction — accessibility vendors publish little such content and would not appear regardless. So this is not proof of a wrong strategy on its own. What it does establish is where the traffic is: the top 25 pages by traffic are almost entirely blog and glossary articles on generic SEO (seo content strategy, seo tips, seo audit), and the homepage ranks tenth on its own site.
That is the fragility. The 85% of AI citations coming from Google AI Overviews (previous section) are drawn from exactly this content — informational SEO-education queries, which is the class AI Overviews increasingly answers without a click. Siteimprove's largest AI-citation channel is built on the traffic most exposed to AI substitution, and least connected to accessibility governance. The presence is real today; the question is how durable it is as the same engines that cite this content start satisfying the query themselves.
Meanwhile the pages that carry commercial intent are close to invisible: the AEO product page draws 2 ranking keywords, the seven-page competitor comparison hub draws 14, and the eight free tools draw 201 — against 6,435 site-wide.
200 on / and /robots.txt. The single most consequential finding of the May audit is resolved. Yandex and Baidu remain blocked — commercially marginal for this market./platform/, which now carries SoftwareApplication, WebPage and VideoObject markup, and FAQ schema on three pages. The AEO product page, the homepage and all seven comparison pages remain on template-only markup..txt path and every /.well-known/ path returns 200 with an HTML error page instead of a 404 — including llms.txt, security.txt, and paths that were never meant to exist. Ordinary page URLs 404 correctly. See the panel below for what this is already costing./pricing/ redirects to the homepage. Third-party pricing claims continue to set the number buyers see.noindex. The AEO Checker is the one tool correctly listed.An independent agent-readiness scanner assessed siteimprove.com about an hour after this audit's measurement window and scored it 61 out of 100. The score is not the interesting part. What it credited them with is.
| What the scanner recorded | What is actually served there |
|---|---|
| OAuth 2.0 support — Passed. "OpenID Connect discovery endpoint found" | A 795-byte HTML page reading "Resource Not Found", returned with a 200 OK status |
| Agent instructions — Partial. "Agent instruction file at /.well-known/agent-skills/" | The same 795-byte error page |
| MCP server — Partial. "MCP endpoint found… but response is not valid JSON" | The same 795-byte error page. It is not valid JSON because it is an error page |
Siteimprove publishes none of those three things. The scanner believed it did because the server answers 200 OK to every request in that namespace — including one I invented to test it. The score is inflated by credit for capabilities that do not exist, and the same misconfiguration will mislead every automated evaluator that probes there, in either direction.
A company selling AI-visibility measurement is being systematically mis-measured by its own web server — and this is the first documented instance of it happening to a real evaluator. The remedy is a server configuration change, not a content project.
The same scan independently confirmed several findings here from different infrastructure — an identical sitemap count, identical homepage heading structure, and AI crawlers reachable — which is useful corroboration. It also contradicted itself twice and passed the site's 404 handling, having tested only the one URL shape that works.
The homepage carries an Organization record in structured data. It is the single most machine-readable statement of what this company is, and it is what Google's Knowledge Graph and any parsing agent reads first. It currently says:
"Siteimprove's all-in-one platform amplifies digital marketing, so you can maximize your reach and provide and an exceptional, inclusive digital experience."
That is the pre-rebrand positioning. Not Agentic Content Intelligence, not AEO, not the unified Siteimprove.ai platform. The May audit found the brand's stale identity lodged in AI training data and rightly called that slow to shift. This is the same stale identity sitting in a field the company owns outright and can change today.
It also contains a grammatical error — "provide and an exceptional" — shipped in structured data on the homepage of a content-quality vendor. Rewriting one string is the highest ratio of positioning impact to effort anywhere in this report.
Both free checkers were taken apart at the API level. They fail in opposite directions, and both failures are the kind a competitor raises in a sales call.
| Tool | What the page promises | What the API actually returns |
|---|---|---|
| AEO Checker | Four dimensions — answer readiness, trust signals, topic clarity, technical readiness — naming FAQ and HowTo schema, Speakable markup, Person schema, author attribution, H1 structure, content depth and site speed. | One thing: whether eight AI crawlers can reach the URL, plus whether a sitemap exists. No schema parsing, no content analysis, no score. Of roughly fifteen named signals, it measures one. |
| SEO Checker | Core Web Vitals, structured data, accessibility and internal linking, "crawled at scale". | Genuinely substantial — 30 real on-page checks plus a live Lighthouse run. But the Lighthouse build is version 9.6.8, which cannot measure INP, the Core Web Vital that replaced FID in March 2024. It still emulates a 2016 handset and still reports a PWA score Google retired. |
Enter a domain that does not exist and the AEO Checker reports status: blocked, sitemap: found, and robots.txt permits every AI engine. Its own diagnostics record the truth — ENOTFOUND, fetch failed — and the user-facing result discards it. A prospect who mistypes their domain is told, on a vendor's own tooling, three things that are not true.
The AEO Checker is honest engineering with dishonest packaging: it checks the one thing that must be true before citation is possible, and checks it well — it correctly identified the real crawler blocks on nytimes.com, levelaccess.com and silktide.com. The gap is the label on the box. The SEO Checker is the reverse: a better tool than its marketing suggests, quietly running on a measurement stack four years out of date.
Make the error endpoint return a real 404, stop rewriting .txt and /.well-known/ requests into it, then publish a real llms.txt. In the same sitting, rewrite the homepage Organization description to the current positioning and add logo, address and contactPoint. Neither task is a project, and together they close the two findings where machines are actively forming wrong beliefs about the company.
It carries the category argument and currently carries the weakest signal: 525 words, no product schema, no FAQ markup, no entity anchor, 71 visits a month. Ship SoftwareApplication, Organization and FAQPage schema, and expand the page to answer the questions buyers actually ask. The markup pattern already exists elsewhere on the site.
Half a day of development · one week of editorialCorrect the false "blocked" verdict for unresolvable domains, upgrade the Lighthouse runtime so Core Web Vitals means the current three, and bring the AEO Checker's copy in line with what it measures — or extend it to measure what the copy claims. The schema-detection capability already exists in the adjacent SEO Checker.
One to two sprints · product and engineeringThis is a leadership question, not an SEO task. Siteimprove's traffic — and ~85% of its AI citations — come from generic SEO-education content, the same ground Ahrefs and Semrush dominate and the query class generative engines are absorbing fastest. Either commit to defending that content with genuine differentiation, or rebalance toward accessibility governance and answer-engine topics where the company has an actual claim and where its AI citations would describe what it sells.
One quarter · CMO decision, then content
The 4% citation-share figure from the May audit was not reproduced, and this pass does not claim to update it. That number was share-of-voice from a five-engine prompt sweep; this workspace has no Brand Radar entitlement, so SoV could not be recomputed on the same basis. The citation counts here are a different metric — but they are compared against direct competitors measured the same way on the same day, which is what makes the "leads its peers" read defensible. Counts are modelled estimates, not a census of every AI answer.
The unblock could not be dated. The Bing block was open on 28 May and closed by 1 September; where in that window it came off is unknown. So no claim is made about how much citation recovery "should" have happened by now — the citation figures are a baseline to re-measure, not a verdict on the unblock's payoff. A fixed 60-day re-measure of ChatGPT and Copilot counts is the clean test.
Brand memory was not re-scanned. No token-level perception test ran this pass, so the May finding that the model still describes a pre-2026 product is neither confirmed nor refuted here.
All requests came from one IP in Germany. Firewalls can vary by region and network, so the Bing unblock should be confirmed from a second region before it is treated as universal — though every Bing and MSN variant returning 200 on two passes, plus the independent scan reaching the site from separate infrastructure, makes a local false positive unlikely.
No security probing was performed. Testing for exposed files or admin paths on a domain we do not own requires written authorisation, which was not in place. That layer is untested, not passed.