The State of AI Disclosure 2026
Eight days after Article 50 of the EU AI Act began to apply, we swept our locked cohort of 1,142 detector-flagged, EU-facing sites — 1,088 scanned, 54 excluded by robots.txt and counted, each rendered in a real browser, homepage only, read-only — on a staffed Monday, early afternoon in EU time. Corrected for the detector's measured error, roughly 11% of the 12,285-site frame shows a real chat launcher (the frame and its correction were measured 27–28 July; the surfaces below, on the scan date). Most of what happens at the first interaction is invisible to our automated visitor: on 78% of confirmed widgets it could not read the first message at all, and on the 174 surfaces it could read, a disclosure that the chat is automated was detected on 19 — about one in ten. This page is a single post-application snapshot: it publishes no before/after delta and no compliance verdicts, and every exclusion is counted below.
≈11% of 12,285 EU-facing candidate sites show a real chat launcher once measured detection error is corrected (95% CI 9.0–14.3); the raw detector said 9.4%. Frame and correction measured 27–28 July 2026 — this figure reads no sweep data and is not dated to the scan day.
How to read these numbers
Every figure on this page is an aggregate. No individual site is named, and none will be: the point is the state of the ecosystem, not a wall of shame. We hold that a raw detector count without its precision and recall is not a finding — and by that standard exactly one rate here is published corrected for its detector's measured error, with the error rates themselves in the open: the prevalence headline. The other numbers below are raw, and are labelled as such where they appear. The widget detector's error was measured, but it is not corrected out of the coverage figure, whose denominator it contaminates in a direction we name in that section. The disclosure detector — the one deciding whether a first message discloses AI — has no measured precision or recall at all: no judging pass in this study ever read disclosure wording, so the three disclosure-state counts are raw detector output whose error rate is unknown rather than known to be small. Where we report disclosure, a finding of “not detected” means our scanner, loading the site the way this study's automated visitor does — which is not the way a human visitor does, and this study ran no human arm, so it cannot tell you how far apart the two are — could not observe a disclosure at the widget's first interaction: a statement about what was observable, never a legal conclusion about any organisation. We report three outcomes only: detected, not detected, and could not verify — and we publish the could-not-verify rate instead of hiding it. Why we never say “compliant” applies to this study exactly as it applies to a report.
How many EU-facing sites run a chat widget?
The corrected figure above comes from auditing our own detector, not from trusting it. Both admission rules were checked against what a blind judge could see in a real-browser screenshot of the rendered page — that capture is the ground truth here, and it is not the same thing as what a site actually runs, nor what a human visitor would find — and the raw detector count is adjusted by the measured rates below. The two errors run in opposite directions and do not cancel; the correction and its confidence interval carry both.
One defect in that interval's history is disclosed here rather than left for a reader to find. The bootstrap behind the interval this study first drafted drew the same sampling variance twice, so the interval it produced — 7.95–15.42 — was labelled 95 % but sized closer to 99 %. The estimator was frozen under precommitment until the second sweep completed, and the one-line fix was applied after it, on 10 August 2026; the correctly sized 95 % interval is the one printed above, and the point estimate did not move. Note which way the fix cut, because it ran against this page rather than for it: the load-bearing claim here is that the corrected figure is not distinguishable from the raw detector's, and the narrower interval makes that claim harder to sustain, not easier. The raw figure remains inside the corrected interval — the claim survives the fix, by less room than the draft interval suggested. Both intervals are stated here so the correction is checkable rather than silent.
| Correction component | Measured on | Rate (95% CI) |
|---|---|---|
| Generic-launcher admits showing a real launcher | 49 decidable | 47% (34–61) |
| Vendor-fingerprint admits showing a real launcher | 52 decidable | 73% (60–83) |
| “No widget” sites showing a launcher anyway (recall gap) | 322 decidable | 6.8% (4.6–10.1) |
| Inert-SDK sites showing a launcher anyway | 41 decidable | 12% (5–26) |
By widget vendor
Vendor shares are of sweep-confirmed widgets. One caveat travels with this table: “the vendor's SDK is present” and “a visitor can chat” are different claims. A substantial share of the sites whose vendor SDK we fingerprinted showed no chat our scanner could reach. We cannot tell you why, and we will not guess: that gap is a single undecomposed quantity mixing our own detector's false positives, inert or disabled SDKs, staffed hours, login gates and several days of site churn between the two measurements. This study forbids itself from republishing it as any one of those causes — in particular as an after-hours or weekend penalty — and its size is not distinguishable from the detector precision published above.
Two limits bear on this table specifically. First, a consequence of the frame, pre-registered before the data existed: candidates are taken from the top of a popularity ranking, which over-represents enterprise sites relative to the long tail, so SMB-leaning vendors are under-represented here against their real install base. The unknown / custom cell is precisely the one a rank-truncated frame inflates, because the vendors that concentrate below the rank cut never enter the study at all. Second, the SDK-versus-reachable caveat above is scoped to the fingerprinted rows, while the detector contamination sits almost entirely in the other class: the sites matching no vendor fingerprint are the ones our admission audit found least likely to be chat widgets in the first place. “Unknown or custom” is therefore a statement about what our fingerprints matched — it is not a measurement of how much of the ecosystem is custom-built, and correcting it for measured precision would move it down. Vendor cells below the minimum-cell threshold are grouped as “other / n too small” rather than published individually; the country section's note on how the pre-registered threshold is applied covers this table too.
| Vendor | Sites | Share of confirmed |
|---|---|---|
| unknown / custom | 570 | 72% |
| zendesk | 60 | 8% |
| livechat | 33 | 4% |
| other (13 vendors, each n < 30) | 131 | 16% |
Can an automated visitor reach the first message at all?
On 78.1% (95% CI 75.1–80.8) of the 794 sites where the sweep confirmed a chat widget, our automated visitor could not read the first-interaction surface. Raw share on the detector-confirmed denominator, scanned Monday 10 August 2026 (graded 11:01–12:12 UTC).
This number is usually buried as a denominator footnote; we measured it on purpose and publish it as a headline. For every confirmed widget the sweep tried, read-only, to dismiss the consent layer, click the launcher and read the first message, and recorded where that failed. Consent walls, launchers our scanner could not click, panels that never open, panels whose text could not be read: each failure cause is counted below, including the ones that are our instrument's limits rather than the site's.
Three qualifications on that accounting, none of them buried. The consent step is attempted but its outcome is not recorded per site, so neither a reader nor an author can audit which unreadable rows really died on a cookie wall — only the ones that failed loudly enough to be bucketed appear as consent failures, on a scanner whose own source calls cookie walls the primary blocker to reading a widget's first message. The bucketing itself carried a defect in this study's first sweep: rows were assigned by matching error-message wording, so panels that opened but whose text could not be read under unexpected wording fell into “other”, and the drafted funnel failed simple subtraction. The aggregation code was frozen under precommitment until the second sweep completed and was corrected after it, on 10 August 2026; buckets now key on the recorded interaction outcome, and the table below is generated by the corrected code. And each failure cause below carries an attribution — instrument-side, site-side, or mixed — because one cause this study first drafted was labelled as a fact about the site when the code that emitted it reads it as a fact about us: rows where a vendor SDK was fingerprinted but our launcher selector no longer matched are a stale recipe of ours, not an absent widget. They are labelled that way below, and the interaction-stage instrument-side subtotal is printed as its own row. Read that subtotal as a floor on our own share of this headline, not the total: the blind audit below shows the largest failure bucket is dominated by a different error of ours — admissions that were never chat widgets — which interaction-stage tags cannot see, and which is why that bucket is tagged mixed rather than site-side.
One known contamination is stated rather than hidden: our admission audit found that a share of “confirmed” widgets are not chat widgets at all, and those sites land in the “panel did not open” row — so the denominator here is detector-confirmed, not verified widget presence. That is not a small caveat on a large number. The interval printed with this headline is sampling error only, and sampling error is not the dominant uncertainty here: the denominator contamination is larger, runs in one direction, and is not inside that interval. Correcting for it would move this figure down. The correction pre-registered for exactly this purpose — a blind re-judging of the failure-bucket screenshots this study already holds — has now been run against this sweep's rows, and its result is published here rather than left as an open action item: the sweep’s own above-fold screenshots of all 361 “panel did not open” sites were re-judged blind (judges saw the image only — no site, vendor, or error metadata; 20 planted duplicates agreed 19/20 on the launcher question). 260 were decidable: 67.7% (95% CI 61.8–73.1) showed no chat launcher visible in the capture — accessibility widgets, carousel arrows, trust badges and contact links our detector had admitted as chat — and 32.3% (26.9–38.2) showed a real launcher whose panel nonetheless did not open. Two bounds travel with that rate: a launcher below the fold or rendered after the screenshot reads as absent, so the no-launcher share is biased upward by an unmeasured amount; and the 101 undecidable screenshots are dominated by consent layers covering the corners where launchers sit — an exclusion that is not neutral. Both are stated in the method log record beside the judgements. The headline above is still printed raw, on the detector-confirmed denominator, exactly as pre-registered — re-basing it on the audited denominator is an analysis change that goes through adversarial review before it is published, not a hand adjustment — but the size and direction of its largest known error term are now measured and on the page, not disclosed and unresolved.
Whatever an automated visitor cannot reach, this automated check cannot verify. That sentence is bounded to this instrument deliberately, because the unbounded version — that no automated check could do better — is one this study falsified in its own workshop. A single fix to how our scanner clicks launchers moved the readable share on the same subsample of sites by around eight points — most of that movement was in our code rather than in the web, though the log is explicit that this comparison ran across an instrument change and cannot fully separate the two (method log §13a: 9 readable gained against 2 lost, 137 sites). That cross-code movement is also this study's only noise estimate — there is no test-retest floor of any kind, as the provenance table below states. Our own aggregation code refuses to publish a per-vendor readable column for this reason, in its own words, because it “would rank OUR recipes while reading as a ranking of vendors”. Aggregating that quantity across vendors does not stop it being a property of our recipes, and the headline above is its aggregate. Read it as the ceiling on what this scanner could see — an indication of the difficulty an automated check faces, never a measurement of the limit of automated checking.
| Stage | Sites | Share of confirmed widgets |
|---|---|---|
| Widget confirmed by the sweep | 794 | 100% |
| Chat panel opened | 259 | 33% (29–36) |
| First-interaction surface readable | 174 | 22% (19–25) |
| Unreadable: Panel did not open after the click — mixed | 361 | 45% (42–49) |
| Unreadable: Panel opened but its text was unreadable — instrument-side | 85 | 11% (9–13) |
| Unreadable: Launcher covered by another overlay — site-side | 54 | 7% (5–9) |
| Unreadable: Vendor recipe stale (SDK fingerprinted; our launcher selector did not match) — instrument-side | 42 | 5% (4–7) |
| Unreadable: Panel selector matched only the closed launcher — mixed | 34 | 4% (3–6) |
| Unreadable: Launcher found but not clickable — site-side | 25 | 3% (2–5) |
| Unreadable: Launcher covered by a consent/cookie layer — site-side | 14 | 2% (1–3) |
| Unreadable: Other — unattributed | 5 | 1% (0–1) |
| Unreadable subtotal attributable to the instrument at the interaction stage (rows tagged instrument-side above; mixed rows are not counted). A lower bound: it excludes the admission contamination the blind audit above measures inside the panel-did-not-open row | 127 | 16% (14–19) |
By country
Country is the site's ccTLD. The rate shown is the share of confirmed widgets whose first message was readable. Cells under 30 are reported as “n too small” per the pre-registered minimum-cell rule — applied here to confirmed widgets, which is the denominator of the published rate. The rule as pre-registered set that threshold on cohort sites; this note is where that change is recorded, and the vendor table above applies the same threshold.
This table ranks our scanner at least as much as it ranks countries. The readable rate is far higher on sites where we fingerprinted a known vendor than on sites admitted by the generic rule, for the mundane reason that a known vendor is what gives our scanner a panel recipe to follow. That split is large: readable rate runs far lower on no-vendor admits than on vendor-fingerprint admits. The per-country share of unknown / custom widgets is not published in the table below, so a reader cannot check that correlation against what is printed here — it is asserted from the underlying rows, not shown. The table is published because it was pre-registered, and cutting a pre-registered table after seeing which way it fell would be worse than publishing it with this warning attached — but read the ordering as a map of our recipe coverage per country before reading it as anything about the countries. Rank depth was tested as an alternative explanation and does not account for the spread.
| ccTLD | Confirmed widgets | Surface readable (95% CI) |
|---|---|---|
| .de | 202 | 11% (8–17) |
| .fr | 99 | 20% (13–29) |
| .it | 62 | 18% (10–29) |
| .pl | 53 | 38% (26–51) |
| .nl | 47 | 15% (7–28) |
| .se | 35 | 29% (16–45) |
| .cz | 33 | 30% (17–47) |
| all others (each n < 30) | 263 | 28% (23–33) |
Disclosure at first interaction — on the readable subset only
Read the scope before the number. This table covers only the sites whose first-interaction surface we could actually read — the readable share shown above — and those sites are not a random sample: they skew toward simpler widgets and weaker consent walls. It is reported as the state of the readable web, not extrapolated to the rest, and that selection is exactly why the coverage number above is a headline finding rather than a footnote.
What this table is not: evidence about Article 50 as a duty. Article 50(1) obliges a provider to inform a person that they are interacting with an AI system, and the rule pack's failing clause accordingly requires that the chat was confirmed to be automated. Confirming automation means probing a chat for an automated reply, and this study never does that to a site it does not own — the probe is hard-disabled for every third-party site, as a safety rule. The consequence is structural and belongs on the page rather than in a log: no site in this study could have been graded as failing Article 50(1), and none was. Every disclosure finding here is at most a hedged warning, which is what the rule pack itself says in its own words — automation could not be confirmed by the probe, so the finding stays hedged rather than a failure. Two further blind spots run the same way. The instrument cannot see Article 50(1)'s carve-out for cases where it is obvious to a reasonably well-informed person that they are dealing with an AI: disclosure is read from the first-interaction text alone, so a widget whose launcher plainly reads “AI assistant” but whose greeting does not repeat it is counted below as no disclosure detected — an error running against the sites doing the most visible thing. And Article 50(1) addresses providers, while the duties falling on the site that deploys a widget are 50(2) and 50(4). Nothing on this page allocates a duty to anyone.
How the three cells are drawn. “Could not verify” here means the wording was inconclusive on a surface we did read — a different thing from the coverage figure above, which counts surfaces we could not read at all. The split between “could not verify” and “no disclosure detected” turns partly on a minimum character count for the first-interaction text — 120 characters — which lives in the detector code rather than in the versioned rule pack, so a sealed report citing a pack version does not pin it. “No disclosure detected” is also a union of two different pack outcomes, one of them covering chats that present as human-staffed — where, if humans genuinely staff the chat, no disclosure is required at all. This page cannot tell you how that cell divides, because the aggregation query did not retain the field separating them. And, as stated at the top: these three counts are raw detector output. The disclosure detector's precision and recall were never measured, so unlike the prevalence headline these cells carry no correction and no published error rate.
| Outcome on the readable first-interaction surface | Sites | Share of readable |
|---|---|---|
| Disclosure detected | 19 | 11% (7–16) |
| No disclosure detected | 73 | 42% (35–49) |
| Could not verify (wording inconclusive) | 82 | 47% (40–55) |
What we could not measure — published, not buried
There is no before/after comparison on this page: the baseline sweep taken on 1 August was retired by our own decision on 2 August, so this study holds no pre-application measurement of these surfaces and publishes none. There is no test-retest noise floor either — the pre-registered same-day retest never ran on baseline day, and its registered substitute never ran on run-2 day; the only noise estimate we hold is a cross-code bound (≈4 points on widget presence, ≈8 on readable surfaces, 137 sites, measured Monday 27 July across an instrument change), and nothing on this page calls that a noise floor. Of 794 confirmed widgets, 259 opened a panel and 174 yielded readable first-interaction text; the disclosure counts cover those 174 alone. Of the 361 sites whose panel never opened, a blind re-judging of the captured screenshots found no chat launcher visible in the capture on two-thirds of the decidable cases — consistent with the admission false-positive classes our own audits measured, and counted in the coverage section rather than buried. The sweep egressed from the United States (measured on scan day: US, Miami PoP); a widget that geofences or staffs differently for EU visitors would look different from inside the EU, and we did not measure that. The scan ran on a Monday where the registered plan said Saturday — a deviation recorded as one, with its reason, in the method log; that log, like the rest of the bundle, is not yet published.
Methodology
The sampling frame, the exclusion rules and the judgement protocol were pre-registered before the data existed, and are versioned. Several pre-registered requirements were nonetheless not met; those are named in the section above rather than dropped. The method log itself is not currently published — see the note under the instrument table. In brief: candidates are the highest-ranked pay-level domains in the Tranco list (XN23N (2026-07-27)) whose public suffix is an EU-27 ccTLD or .eu. That is a rank-truncated slice of that list, not every such domain: the frame stops at a fixed rank cut, so sites and vendors concentrated below the cut are outside this study entirely. Each candidate was rendered in a real browser (widgets inject after JavaScript, so a static check would systematically miss client-rendered vendors); sites whose robots.txt disallowed our scanner — or whose robots.txt could not be reached — were excluded and counted. The chat panel is opened read-only: the scanner never types into, sends, or otherwise writes to any site's chat.
The discovery funnel — the stage before the sweep table below, at which the prevalence denominator is built — is: 16,000 candidates taken at the rank cut (of 128,124 eligible pay-level domains), 13,441 rendered, minus 708 that redirected off the frame and 448 duplicate final sites, giving the 12,285-site frame the prevalence headline is computed over. Two of its exclusions are named here because they run against this study's own interest, and both are discovery-stage counts — not the single-digit rows of the sweep funnel below. Candidates that answered the discovery pass with a bot challenge are excluded outright rather than counted as having no widget — so the sites most likely to be running defended commercial deployments are absent from the denominator, not scored zero inside it. And the 708 candidates that redirected off the frame were dropped after rendering; as a group they carried a higher widget rate than the sites retained. Both nudge the prevalence denominator the same way, and neither is corrected for. The arithmetic effect is small. We state the counts because a funnel that discards a large share of its candidates before the first published row should not report only the part that survived.
| Stage | Sites |
|---|---|
| Submitted to the sweep | 1142 |
| Excluded: robots.txt disallows scanning | 4 |
| Excluded: robots.txt unreachable (treated as no) | 50 |
| Excluded: invalid / unresolvable | 0 |
| Capture blocked (bot protection / redirect) | 1 |
| Scan did not complete (cause not attributed; excluded) | 22 |
| Scanned to completion | 1065 |
| Captured outside pre-registered hour band (excluded) | 1 |
| Graded inside the pre-registered band | 1064 |
| Widget confirmed at sweep time | 794 |
| First-interaction surface readable | 174 |
Instrument & provenance
What you can and cannot check today. The repository snapshot, the method log and the underlying data are not published. The commit, cohort hash and timestamp anchors above are therefore checkable only by someone who already holds the bundle; no dataset download, tarball or DOI exists, and nothing on this page can be independently recomputed from a public source today. Publishing the bundle is an open action item, not a completed one, and the claims above should be read with that discount applied. Even with the bundle in hand, the sweep aggregates are not recomputable by an outside reader: the aggregation script reads from a private database that does not ship with the snapshot. One practical note for whoever does get the bundle — verify the manifest against the commit named above, not against a later tree: documentation inside the bundle has been edited since the anchor was taken, which is expected and is described in the precommitment note.
Limitations
EU-facing is a ccTLD proxy: candidates are the highest-ranked pay-level domains whose public suffix is an EU-27 ccTLD or .eu — a rank-truncated slice that misses .com EU businesses and everything below the rank cut.
The sweep ran from a United States vantage (measured on scan day: US, Miami PoP). Geofenced or EU-differentiated widgets may present differently from inside the EU; the baseline run’s egress was never recorded and never can be.
Homepage only, one page per site: a widget that lives on a support or checkout page is out of scope by design.
The prevalence correction’s precision and recall inputs are measured on judged samples with stated confidence intervals; the recall term dominates the corrected interval’s width.
Ground truth is a blind judge reading an above-fold real-browser screenshot — not what a site runs, and not what a human visitor exploring the page would find. A launcher below the fold or rendered late reads as absent.
“Unclear” judgements are excluded from every measured rate, and the exclusion is not neutral: unclears are dominated by consent walls that hide exactly the corners where launchers sit.
The disclosure lexicon covers 23 of the 24 official EU languages plus Catalan — Irish is absent — and a greeting in an uncovered language publishes as “no disclosure detected”, not as “could not assess”. The pre-registered language-fallback outcome was never implemented.
No human arm: this study cannot say how far an automated visitor’s view sits from a human visitor’s.
Single post-application snapshot: the retired 1 August baseline is not a publishable comparator, so there is no before/after delta, and there is no test-retest noise floor of any kind.
The automation probe is hard-disabled on third-party sites (it would write into a stranger’s support queue), so no site here could have been graded as failing Article 50(1); disclosure findings are hedged observations, never legal conclusions.
Cite as: DisclosureProof, “The State of AI Disclosure 2026”. Scan dates: Monday 10 August 2026. The aggregate tables on this page are licensed CC BY 4.0 — reuse with attribution and a link. That licence covers what is printed here and nothing more: no underlying dataset, repository snapshot or per-site record is published, and per-site records never will be. Questions, corrections, or press: hello@disclosureproof.com.