What has shipped
DisclosureProof is maintained by hand. Every scan is graded against a dated rule pack, and the version string is printed on the report, so a finding can always be tied to the rules as they stood when it ran. This page tracks the notable changes, newest first.
The State of AI Disclosure 2026 — our first published study
Eight days after Article 50 began to apply, we swept a locked cohort of 1,142 detector-flagged, EU-facing sites and published the aggregates: how many of the most-visited EU sites show a real chat launcher once our detector's measured error is corrected out, how often an automated visitor can read a widget's first message at all, and what share of the readable first messages disclose the AI. Every number travels with its methodology — the exclusion funnel, the pre-registered timestamp anchors, what went wrong on the way (a retired baseline, a noise floor never obtained, both stated rather than smoothed over), and our own instrument's measured errors, including a blind re-audit of our largest failure bucket that found two-thirds of it was our detector admitting things that were never chat widgets. Aggregates only: no site is named, and none will be. Read the study →
rule-pack 2026.08.2
A Hungarian term the review killed had shipped anyway — now it is out
When the disclosure lexicon was extended to 24 languages in July, every proposed term went through a cross-language collision review, and one Hungarian entry was recorded as killed: hyphens in our matcher also match spaces, so the term for an AI assistant, MI-asszisztens, really matched mi asszisztens — and mi is the everyday Hungarian word for "we", while asszisztens is a human job title. A clinic writing "a mi asszisztens csapatunk" ("our assistant team") about its own staff would have been graded as disclosing an AI. The kill was recorded; the term shipped anyway. We found the slip this week while preparing the lexicon for open-source release — the pre-publication review checked the shipped set against the review record, and they disagreed. The term is now removed, with the sentence above pinned as a regression test. Language coverage is unchanged: Hungarian keeps its verified terms, and the idiomatic greetings this entry was meant to catch nearly always carry mesterséges intelligencia or chatbot anyway.
rule-pack 2026.08
Irish joins the disclosure lexicon, and the scanner now says when it cannot read the language
The disclosure lexicon covered 23 of the 24 official EU languages, plus Catalan; Irish was the gap. Its terms were verified against the national terminology database and real Gaeilge usage rather than machine-translated — intleacht shaorga, bota comhrá, and their declined forms — with the same cross-language collision review every other language went through (the Irish abbreviation for artificial intelligence can never be matched: it is spelled exactly like the English word "is"). Irish cookie-consent buttons are now dismissed too, so Irish-language sites' widgets get opened and read at all. And a new outcome closes an old gap: when a page's own markup declares only languages the lexicon does not cover, an unmatched greeting is now reported as could not assess (language) instead of "no disclosure detected" — the latter was a claim about text the scanner cannot actually read, and our own study review called it out.
rule-pack 2026.07.6
The chatbot probe stopped mistaking your own message for a reply
When the scanner types a question into a live chat widget to see whether it answers like a machine, it now waits for the widget to echo that message back before it starts timing anything. Nearly every chat widget renders what you just typed within a fraction of a second, and the probe had been reading that echo as the assistant's reply, which made almost any widget it could open look automated. Confirmed automation is the one signal that can turn a hedged finding into a plain fail, so the cost of getting it wrong is a fail on a site that never earned one. It is now measured from the echo forward, and a candidate reply that merely repeats our own words is discarded.
Three other reads were tightened for the same reason. A page stating that its content is not AI-generated used to be counted as carrying an AI label, because the phrase matched either way. Article 50(4) now gained a clause for the case where a publisher declares AI involvement themselves: that used to fall through to a report line saying no signal was found, which was untrue on the page it was printed on. Image metadata reading the Spanish word for image, or the Italian for "from the", was being matched against the names of AI image tools. And a chat greeting saying "I am not a human, I am an AI assistant" was being read as no disclosure at all, because the denial of being human was taken to cancel the disclosure that followed it. French, Spanish and Italian gained the compound phrasings those languages actually use.
rule-pack 2026.07.5
Article 50(4) reports the evidence, not a guess
Reports, and the Evidence Pack PDF, now show the evidence chain behind an AI-content finding: which signals the site itself emitted, what kind each one is, and where we found it. A generator tag, provenance on published images, and an AI policy page describing how content is produced are three independent kinds of evidence; the same tag repeated across five hundred pages is one. Strength is measured that way, by independent kinds rather than by count, so a single template cannot inflate a site into looking heavily flagged. Every row is something you can go and check yourself. We still do not run automated AI-text detection and still publish no likelihood score: a percentage would read as a measurement of the writing, and no such measurement is reliable enough to seal into evidence.
rule-pack 2026.07.4
Article 50(4) now reaches a clear fail, and stops reporting what it cannot see
This check now has three distinct answers instead of one. When a page's own markup declares its content machine-generated and the reader is shown no disclosure, that is a fail: the publisher has stated it, we simply compared that against what a visitor sees. When a site declares a machine writer site-wide but nothing ties it to a given page, that is flagged for review, because a generator tag does not establish which pages the tool produced. And when there is no machine-readable signal at all, the check says plainly that it could not determine who wrote the text, instead of reporting every unlabelled article as a finding. On an ordinary site with no AI involvement the old behaviour flagged nearly every page, which is noise dressed up as a finding. We do not run statistical AI-text detection, and we do not intend to: it is not accurate enough to seal into evidence, and a wrong answer would be an accusation carrying our signature. The label vocabulary also grew to cover the natural phrasings publishers actually write, in all six languages, including text using typographic apostrophes.
Crawls read your whole sitemap, and respect more of your robots.txt
Large sitemaps are no longer read only in part, and a crawl now records whether it saw all of your sitemap so a coverage claim can never overstate itself. Robots directives using wildcards or query patterns, such as disallowing tracking parameters, are now honoured rather than ignored. Page ordering was rebuilt too: contact, pricing, about and policy pages are picked first, and article pages follow, instead of a headline that happened to contain the word "pricing" outranking your actual pricing page.
Anyone can now check an evidence pack
A new verification page takes an evidence manifest and answers two questions separately: whether the evidence is unaltered, which you can prove yourself with a plain SHA-256 and no help from us, and whether we sealed it, which is a shared-secret signature only we can confirm. The page says plainly which of the two relies on trusting us. Verification keeps working after a scan has aged out of retention, because the signature check is pure arithmetic over the bytes you hold.
Reports are organised by duty
Every report now opens with a summary table: one row per Article 50 duty, the checks it covers, and where they landed. The detail below is grouped the same way instead of running as one flat list, so you can answer a question about synthetic media without reading the chatbot findings first. Printing a report from the browser now produces a usable document.
A real dashboard on every plan
The dashboard gained a sidebar and several new sections for all accounts, free included: open findings gathered across every site you own, an Article 50 readiness view explaining each duty and where your latest scans stand against it, and a disclosure toolkit with wording in the six languages the scanner reads, checked against the scanner's own lexicon so the text it gives you is text it recognises.
Better reading of cookie walls
Consent banners are the main thing standing between a scan and a chat widget's first message. The scanner now recognises fourteen more consent platforms, reads accept buttons in ten languages rather than English only, and can see banners that render inside shadow DOM. It still refuses to click anything that only accepts necessary cookies, since dismissing a wall without granting consent would hide the widget just as effectively.
rule-pack 2026.07.2
Official EU icons and structural C2PA
Added detection for the European Commission's official AI-content icons (the Annex I set from the 10 June 2026 Code of Practice) and structural inspection of C2PA Content Credentials, not just a mention of them. The guide library grew to twelve chat-widget vendors and eleven EU member states, each researched to its enacted-versus-proposed status.
Multi-page coverage
Scans can look past the homepage to the pages where Article 50 duties tend to live: contact, pricing, article, and policy paths first, then report exactly which URLs were checked, so the coverage claim in an evidence pack is honest.
Owner verification
Prove you control a domain with a DNS TXT record or a meta tag. A verified owner can scan their own site more deeply and keep evidence longer.
Sealed, tamper-evident evidence
Every scan now hashes its findings and captured artifacts into a manifest and signs it (HMAC-SHA256). A report can be re-verified later, and any change to the underlying evidence breaks the seal.
rule-pack 2026.07
The six core Article 50 checks
Chatbot disclosure at first interaction (matched against a six-language lexicon), machine-readable media marking, AI-content labels on article-like pages, and an AI-policy-page check went live, each a deterministic check graded against the versioned pack and reported as detected, not detected, or needs verification.
DisclosureProof launched
The free homepage scan, the plain-language Article 50 guide, and the per-vendor and per-country reference library went live.