EU AI Act · Article 50 in force · marking deadline 2 December 2026
EU AI Act · Article 50 · procedure

How to audit an AI chatbot disclosure on desktop and mobile

Seven stepsNo credentials neededDesktop and mobile as separate tests

Most Article 50 findings are not discovered by auditors. They are discovered by whoever complains first, because the check takes ninety seconds and needs no access to anything. This page is the procedure run properly: what to test, in what order, on which devices, and what to write down — so the first person to run it on your site is you.

The whole audit is external. You need no credentials, no vendor dashboard, and no cooperation from the team that owns the widget. That is the point: it reproduces exactly what an authority or a competitor can see.

On this page Before you start: define the scopeThe procedureWhile you are in there: the other three dutiesThree ways a manual audit goes wrongWhen to automate it Common questions

Before you start: define the scope#

An audit that tests the homepage on one laptop will pass on sites that would fail in front of a regulator. Three dimensions decide whether your audit is real.

Devices. Desktop and mobile are separate tests, not one test viewed twice. Compact layouts routinely drop header labels and pre-chat notices that the desktop layout keeps. Use a real phone where you can; a browser's device-emulation mode is a reasonable substitute but will occasionally render a layout no real device produces.

Entry points. Widgets are often configured per page or per audience. Test at minimum the homepage, one product or pricing page, one support or contact page, and one deep content page. If you run different widget settings by geography or by logged-in state, each combination is a separate test.

State. Always a fresh private window. A returning visitor's widget may skip the greeting entirely — which means your everyday experience of your own site is the one state that tells you nothing.

The procedure#

Step 1. Load the site as a first-time EU visitor would#

Private window, desktop, no extensions. Note the date and time. If a cookie or consent banner appears, deal with it the way a visitor would and note whether it covered the chat launcher — consent layers sitting over the corner where launchers live is a real pattern, and it matters because it changes what a visitor can see.

Step 2. Look at the launcher before you click it#

Does the bubble itself disclose anything? "Chat with our AI assistant" on the launcher discloses before the first interaction, which is the safest side of the line. Record what it says, or that it says nothing. Screenshot.

Step 3. Open the panel and read it before you type a word#

This is the test that matters. Open the widget and do not type anything. Read what is on screen: the greeting, the header, any pre-chat notice.

Grade what you see into one of three states, and resist the urge to be generous:

Disclosed
it states the counterpart is an AI system.
Not disclosed
nothing about AI or automation appears.
Inconclusive
something gestures at it without saying it. "Virtual assistant." "Automated helper." "Powered by AI" in the chrome. A header label with no line in the conversation.

The third bucket is the one people misgrade, and it is the biggest. On the first-interaction surfaces our 2026 sweep could read, 47% fell into it against 11% clearly disclosed. Fifteen worked examples, sorted, if you want a reference to grade against.

Step 4. Ask directly#

Type: "Am I talking to a real person?" Record the answer verbatim.

Note carefully what this test is and is not. Article 50(1) requires disclosure at first interaction, so an assistant that admits it only when asked has still not met the duty — a pass here does not rescue a fail at step 3. But a bot that denies being AI, or deflects, is a materially worse problem than one that stayed quiet, and it is common precisely because nobody configures it deliberately: it falls out of a system prompt written for warmth.

Step 5. Do it again on a phone#

Steps 1 through 4 again, on a phone. Do not assume. This is where the majority of real regressions live, and where a desktop-only evidence file quietly stops matching reality.

Step 6. Where else does the widget appear?#

The pages from your scope list. You are looking for per-page configuration drift — a widget that discloses on the homepage and not on the support page it was actually configured for.

Step 7. Write it down while it is still on screen#

For each combination of page and device: the date, the URL, the viewport, what the launcher said, what the first message said, your grade, the answer to the direct question, and a screenshot with the timestamp visible.

A finding you remember is not a finding you can produce later. And record the failures, not just the passes — a dated record showing a gap and a later record showing it fixed is a remediation timeline, which is worth considerably more than an unbroken run of clean screenshots that starts after you fixed everything.

While you are in there: the other three duties#

The chat widget is one of four Article 50 duties, and the audit is cheap to extend:

Machine-readable marking (50(2))
Download a few published images and inspect them for C2PA Content Credentials or IPTC provenance fields. Optimisation pipelines and CDNs strip metadata routinely, so content that left the generator marked often arrives unmarked. See C2PA vs watermarking vs metadata.
Visible labels (50(4))
Open your most recent AI-assisted article. Is there a label a reader would notice, near the headline rather than after the body?
AI-use policy page
Does one exist and is it linked? Not a substitute for the in-chat disclosure — see does a privacy policy count? — but worth having.

The printable one-page checklist covers all four if you would rather work from paper.

Three ways a manual audit goes wrong#

Auditing the settings screen instead of the site. The single most common error. The configuration is a claim about the site; the site is the evidence. Check the rendered page.

Auditing while logged in, on the office network, in the browser you use every day. Cached state, a returning-visitor path, an internal IP that gets different widget behaviour. Private window, and ideally a connection that is not your office.

Auditing once. A widget update, a theme change, a consent-banner tweak, or a vendor's own release can silently remove a disclosure that was there last quarter, and nothing notifies you. That is a different problem from getting it right the first time — why this is not a one-time checkbox.

When to automate it#

Do the manual audit once, however you plan to proceed. It teaches you what your own site does and gives you a reference to grade against.

Automate when the matrix stops fitting in a morning: several sites, several entry points, two viewports on a repeating cadence. The records have to exist a year from now. That is what our scanner does on every run — real browser, widget opened, first message read before anyone types, mobile checked separately, media sampled, findings graded detected / not detected / could not verify, and the whole capture sealed into a dated record. The free homepage scan runs the whole thing with no signup, and the sample report shows exactly what comes back.

Run the automated version in 90 seconds. Real browser, widget opened, first message read before anyone types. Mobile checked separately, media sampled, and the whole capture sealed into a dated record. Run the free scan →

Common questions

Do I need vendor access or credentials to audit a chatbot disclosure?

No. The whole audit is external and needs no credentials, no vendor dashboard and no cooperation from the team that owns the widget. That is the point — it reproduces exactly what an authority or a competitor can see, which is how most Article 50 questions actually start.

Why test desktop and mobile separately?

Because compact layouts routinely drop header labels and pre-chat notices that the desktop layout keeps, and the duty follows the visitor. Mobile is where most real regressions live, and it is where a desktop-only evidence file quietly stops matching reality. Use a real phone where you can; device emulation is a reasonable substitute but will occasionally render a layout no real device produces.

What should I grade an ambiguous greeting as?

Record it as inconclusive rather than as disclosed. Wordings such as "virtual assistant", "automated helper", or "powered by AI" in the chrome gesture at automation without stating it. That bucket is the biggest one in practice — 47% of the readable first-interaction surfaces in our 2026 sweep, against 11% clearly disclosed — and grading generously is how teams miss the finding that matters.

Should I record failures as well as passes?

Yes. A dated record showing a gap on one date and the same check clean on a later date is a remediation timeline, which is worth considerably more than an unbroken run of clean screenshots that starts after you fixed everything. Keeping only the clean records is a common instinct and a bad one.

How often should the audit be repeated?

Quarterly as a floor and monthly as a sensible default, plus an explicit check either side of any redesign, widget migration or CMS change. Widgets update themselves, greeting fields get optimised for conversion, and consent banners move — none of which raises an error or fails a test.

Sources and further reading

Last updated September 2026. Informational only, not legal advice: this page describes what the text of the EU AI Act says and what an external check can observe, not whether any particular site complies. Corrections welcome at hello@disclosureproof.com.