State of AI Disclosure 2026 — method log
This is the working log, not a write-up. It is the record kept while the study was designed, run, broken and re-run — pre-registrations that were not met, a baseline that was retired, a noise floor never obtained, and instrument bugs found after the numbers existed. Read it against the finished study: The State of AI Disclosure 2026 →
This is the working method log for The State of AI Disclosure 2026, published in full. It is not a write-up. It is the log as it was kept while the study was designed, run, broken, argued with and re-run — including the pre-registrations that were not met, the baseline that was retired, the noise floor that was never obtained, and the instrument bugs found by adversarial review after the numbers already existed.
It is published because a study whose method is described only in summary asks to be taken on trust. Reading this should make it easier to find fault with the study, not harder. Where the log contradicts the research page, the log is the older document and the research page's published figures govern; where the log records a failure the research page does not mention, the failure is real and the omission is ours.
Two redactions, and only two. Neither removes a finding.
- No third-party site is named anywhere in this document. The research page states that no per-site record is published and that per-site records never will be, and we repeated that commitment in writing to a journalist who asked for this log. Sites that the internal copy names — in frame-composition adjudications, in one shadow-DOM panel finding, and in one throttled scan row — are described here by what made them methodologically interesting, never by identity. Company names attached to those domains went with them. Infrastructure that is not a studied site (the Tranco list, the OpenTimestamps calendars, the egress-probe endpoints) is named normally.
- The founder's personal name and local machine paths are replaced. Decisions recorded against a personal name read as "the founder"; a home directory reads as
~. No research fact is carried in either.
Everything else is verbatim, including the parts that are unflattering. Section numbers, cross-references and dates are unchanged, so §-citations from the research page resolve here.
Licence: this page, like the aggregate tables on the research page, is CC BY 4.0. It still publishes no dataset, no repository snapshot and no per-site record.
Questions and corrections: hello@disclosureproof.com.
Sections
- 1. Source list
- 2. EU-facing definition — the frame's central judgement call
- 3. Candidate selection
- 4. Discovery pass (rendered, never static)
- 5. Pre-registered analysis rules (declared before any data existed)
- 6. Instrument freeze (extends STUDY_PIPELINE.md §8b)
- 6b. Discovery instrument notes (2026-07-27 run)
- 7. Counts — FINAL, cohort locked (2026-07-27; assembly: assemble-cohort.mjs, lock: cohort-lock.mjs)
- 7-old. First-pass counts (10,000-candidate checkpoint, superseded)
- 7a. non_eu_operator adjudication (2026-07-27; full table in tools/study/adjudication-and-review.json)
- 7b. Frame-as-built review (2026-07-27; findings in adjudication-and-review.json)
- 8. Pre-flight adversarial review (2026-07-27)
- 9. Harness (2026-07-27) — built, reviewed, deployed, verified
- 10. Pre-baseline gate results (2026-07-27) — THE AGREEMENT GATE FAILED
- 11. Cohort-quality audit (2026-07-27) — first pass, and why it is a LOWER BOUND
- 12. Cohort-quality audit — FINAL (2026-07-27, production-consent capture)
- 13. Post-fix re-measurement (2026-07-27, Worker 4af9624f)
- 14. fingerprint_strong precision — the §12b gap, now closed (2026-07-27)
- 15. DECISION (2026-07-28, the founder) — reframe adopted; baseline this week
- 16. Run records — preflight DONE 2026-07-28; sweeps to be filled on the day
- 17. AMENDMENT (2026-07-30, the founder): baseline slips one day to Friday 31 July
- 18. AMENDMENT (2026-07-31, the founder): baseline moves to Saturday 1 August
- 19. The baseline as executed (2026-08-01)
- 20. The substitute noise floor (decided 2026-08-02, before run 2) — NOT OBTAINED (2026-08-10). §19f's rule revives unchanged — see §26
- 21. Corrections from the Phase 5 adversarial review (2026-08-02)
- 22. §6 RECORDED DEPLOY (2026-08-02) — the freeze window ended before run 2
- 23. The sweep egress country, measured at last (2026-08-03)
- 24. §6 RECORDED DEPLOYS (2026-08-06) — the launch-week fix wave, and a bookkeeping failure stated plainly
- 25. Run 2 missed 8 August and moves to Monday 10 August — a §5 weekday deviation, stated as one
- 26. run2-noise did not run; the floor is dropped (founder decision, 2026-08-10)
- 27. The §15b blind re-judging, run at last — two-thirds of the largest failure bucket shows no launcher at all (2026-08-10)
Written as the work happens (per STUDY_PIPELINE.md §4), not reconstructed afterwards. Every judgement call goes here with its reasoning. This file is the source for the published methodology section.
Status: frame built, pre-flight review applied, rendered discovery running (2026-07-27, overnight).
1. Source list#
- List: Tranco list XN23N, top 1,000,000, pay-level domains (
filterPLD: on), providers: crux, farsight, majestic, radar, umbrella. Published daily; each daily list aggregates a 30-day window — XN23N covers 2026-06-27 → 2026-07-26 (from the list's own configuration metadata). The 30-day averaging is where Tranco's stability and manipulation-resistance come from; cite the window, not just "daily", or a critic will claim the property was configured away. - Retrieval:
https://tranco-list.eu/download/XN23N/1000000, fetched 2026-07-26. Tranco lists are permanently retrievable by ID, which is why the ID is the citation. - Raw CSV kept locally at
tools/study/raw/tranco-XN23N.csv(not committed; 1M rows).tools/study/candidates.csvandtools/study/tld-counts.jsonare committed and are byte-reproducible from the raw list viatools/study/build-frame.mjs.
Why Tranco: reproducible by construction, standard in web-measurement research, manipulation-resistant relative to any single-provider rank. Hand-picking sites is ruled out by STUDY_PIPELINE.md §4.
2. EU-facing definition — the frame's central judgement call#
Chosen: a pay-level domain whose public suffix is an EU-27 ccTLD or .eu.
The full TLD set: at be bg hr cy cz dk ee fi fr de gr hu ie it lv lt lu mt nl pl pt ro sk si es se eu.
Why this definition and not the alternatives (STUDY_PIPELINE.md §4 lists three):
- It is decidable from the source list alone. No extra fetches, no language detector, no imprint heuristic. The frame is exactly reproducible from (list ID, TLD set) — two constants in a committed script. Both alternative definitions (EU language, EU legal surface) require a crawl with its own failure modes and tunable thresholds, and their frame cannot be re-derived without re-crawling.
- Registering a member-state ccTLD is a strong (not perfect) signal of addressing that market. Imperfect both ways: domain hacks ride EU TLDs from outside the EU (see the adjudication rule in §3), and EU businesses operate on .com/gTLDs.
- The stated limitation: this frame is a large, well-defined subset of EU-facing sites. Published claims are phrased about "sites under EU country domains", never "all EU sites". The run-1 vs run-2 delta compares the same cohort with itself, so frame-coverage limitations do not touch the delta.
Deliberate exclusions from the TLD set (recorded, negligible in the top 1M): IDN variants (.ελ/.ευ), French outermost-region ccTLDs (.re .gp .mq .yt), and EEA-but-not-EU (.no .is .li) — the AI Act is not incorporated into the EEA Agreement as of the study date.
3. Candidate selection#
The 10,000 highest-Tranco-ranked domains in the EU TLD set.
- Population: 128,124 of the 1,000,000 Tranco domains fall in the EU TLD set (per-TLD counts in
tools/study/tld-counts.json; top: .de 28,245 · .nl 18,868 · .fr 11,013 · .it 10,050 · .pl 8,871). - The 10,000 candidates span Tranco global ranks 45–104,186.
- Why 10,000 and not the originally sketched 5,000 (pre-flight review, power arithmetic): at an expected 10–15% widget-detection rate, 5,000 candidates give a ~500-site cohort, whose before/after delta carries a ±~7pp 95% CI — likely indistinguishable from zero. 10,000 candidates target a 1,000+ cohort. Discovery is a few hours of local compute; sweep compute is explicitly not the constraint (STUDY_PIPELINE.md §2).
- Top-N by rank, not a random sample: zero sampling choices to defend, fully reproducible. Known consequence to state with the vendor split: a popularity-ranked frame over-represents enterprise sites relative to the long tail, so SMB-leaning vendors are under-represented versus their raw install base.
3a. Rank extension to 16,000 (pre-registered 2026-07-27, before extension output existed)#
The realized cohort rate was 6.44 per 100 candidates (644/10,000) — §3's power target of 1,000+ was missed. Fix: the identical rank-ordered rule continued down the identical frozen list — candidates 10,001–16,000 from XN23N, same TLD set, same hygiene rules, same instrument (plus the §6b timeout salvage, which applies to old and new candidates alike). Decided, committed, and recorded here BEFORE the extension run produced any output and before any disclosure data existed. At the realized rate the extension targets a ~1,000-site cohort.
Frame-hygiene rules (pre-registered BEFORE any discovery output was read)#
Applied at cohort assembly from recorded fields; discovery itself excludes nothing.
- redirected_off_frame: a candidate whose homepage finally lands on a pay-level domain outside the EU TLD set is excluded and the count published. Evidence: the full redirect chain is recorded per candidate. (Guard's example: a top-50-ranked URL-shortener domain on an EU ccTLD that redirects to its parent site outside the frame.)
- duplicate_final_site: candidates landing on the same final pay-level domain collapse to the best-ranked one; the rest are counted as duplicates.
- non_eu_operator (adjudicated): a cohort site that serves directly on an EU TLD but is operated by a non-EU entity with a lexical TLD choice (domain hack — e.g. a link-in-bio service on
.ee, candidate #7; a background-removal tool on.bg) is excluded, adjudicated from the site's own imprint/privacy page, with the adjudication list and count published. Recordedhtml_langsupports this without a re-crawl. - Minimum cell size: per-country and per-vendor splits are published only for cells with n ≥ 30 cohort sites; smaller cells are reported as "n too small", never as a percentage.
4. Discovery pass (rendered, never static)#
Static pre-filtering is prohibited (STUDY_PIPELINE.md §4): Intercom, Tidio and others inject post-JS; a static filter would bias the vendor split toward server-rendered embeds. Every candidate gets a real headless-Chrome render.
Mechanics (tools/study/discover.mjs), mirroring production capture so a site recruited here is detectable by the real sweep:
robots.txtchecked first with the production matcher (src/lib/robots.js, RFC 9309 wildcards,DisclosureProofBotproduct token). Disallowed → no render, recorded asrobots_disallowed(stated exclusion, hard rule 4).- Navigation:
https://<domain>/(onehttp://retry on TLS/connection failure, recorded),networkidle2, 20s timeout — production values (src/lib/capture.js). - UA presented: the browser's own UA with
HeadlessChrome→Chrome— same as production capture as of 2026-07-27. Reason (verified live 2026-07-27): tawk.to refuses to boot its widget for aHeadlessChromeUA — SDK loads,Tawk_APIstays a stub, no iframe is ever created — so the single most widely deployed chat product read as "no widget" everywhere. Instrument decision, stated openly: robots.txt and every non-browser fetch still identify as DisclosureProofBot. - Synthetic read-only interaction before classification (mouse move + 250px scroll + back): delay-JS-until-interaction optimizers (WP Rocket et al.) otherwise never load the widget SDK at all.
- Consent walls: production CMP dismissal (
dismissConsentBanner), which as of 2026-07-27 also (a) covers 23 of the 24 official EU languages in its accept-phrase list (the previous list missed cs/sk/hu/ro/bg/el/hr/sl/lt/lv/et/mt — a differential per-country recruitment bias, fixed pre-run), and (b) runs inside cross-origin CMP iframes (Sourcepoint, consentmanager.net). After a successful consent click the settle is 6s (GTM-gated vendor scripts load only post-consent); otherwise 2.5s. > CORRECTED 2026-08-02 (§21a). This bullet read "covers all 24 EU official > languages". It does not.CONSENT_ACCEPT_PHRASES(src/lib/interaction.js) > carries 23 languages — en de fr nl es it pl pt sv da fi cs sk hu ro bg el hr sl > lt lv et mt — and Irish (ga) is absent. The count "24" was reached by > adding the 12 languages named above to the 11 already present; 11 + 12 = 23, and > the arithmetic was never checked against the official list. Exposure is small (an > Irish-language CMP is rare even on.ie) but the claim was false, and it is the > same off-by-one as §5's lexicon claim, in a different list. - Bot walls: challenge interstitials are classified (
bot_challenge) from strong DOM/host markers, or weak markers combined with a 401/403/429/503 — and reported as their own exclusion class, never folded into "no widget". Weak-only signals on a 200 page (a form captcha) do not exclude. - Signals: observed network hosts, vendor JS globals, rendered DOM, generic-launcher visibility (production
GENERIC_LAUNCHER), the matched vendor's own recipe launcher selector, and a read-only per-vendor API liveness probe for selector-less vendors (tawk.to:Tawk_API.maximizepresent and not hidden; Freshworks:fcWidget.isLoaded()). Classification is productionfingerprintVendor(imported, not copied) + thedetectWidgetpresence rule including SDK-present-but-inert. - Recorded per candidate: rank, robots result, HTTP status, redirect chain, final URL + final pay-level domain,
html_lang, challenge signals, widget present/vendor/confidence/signals, launcher flags, inert flag, consent outcome, error class. Append-only NDJSON, resumable; ametarow records the run's vantage (public IP + geolocation), raw and presented UA, viewport, and all timing constants.
Vantage (pre-registered limitation): discovery ran from a US residential IP (recorded in the meta row); the sweeps run from Cloudflare Browser Rendering egress (country to be probed and recorded before each run). Geo-dependent redirects, consent flows, or widget configs may therefore differ from what an EU visitor sees; the redirect chain and consent outcome are recorded per site so vantage-sensitive exclusions are auditable. Both runs use the same instrument, so the delta is vantage-consistent even where the level is not EU-exact.
UNMET, recorded 2026-08-02 (§21d/3). The clause "country to be probed and recorded before each run" is a pre-registration requirement, and it was never performed — not before the baseline and not at any other point.
egress,egress_countryand every equivalent return nothing in this log, inPREFLIGHT-2026-07-29.md, inPREFLIGHT-2026-07-30.mdor indata/study-results.json; the only hit anywhere is this pre-registration itself. §16c's run record logs the enqueue window, the band reconciliation, the Worker version and the drift check, and no egress country. So the sweep half of this paragraph's vantage disclosure is undocumented: the study cannot say from which country Browser Rendering egressed on 1 August, and the run-1 rows carry no field that would let anyone reconstruct it. It is now a run-2 checklist item (§21g), but that recovers it for 8 August only — the baseline's egress country is not recoverable after the fact.Two further inaccuracies in the sentences above, recorded rather than edited away. (1) "discovery ran from a US residential IP" is not accurate. The eight
metarows indiscovery-run1.ndjsonrecord three networks:73.132.238.210(AS7922 Comcast, 5 runs), two rotating mobile-CGNAT addresses172.58.240.17/172.58.243.71(AS21928 T-Mobile), and108.26.125.166(AS701 Verizon Business). Counted by the rows following each meta row: Comcast 46.9%, T-Mobile 45.1%, Verizon Business 8.1% of 29,677 discovery rows — so a majority (53.1%) of discovery fetches came from something other than the residential cable line the sentence describes. The country ("US") is right for all three; "residential" is right for fewer than half. No bias is measured — per-segment exclusion differences are confounded with rank, and within-vantage spread matches between-vantage spread — but the description is wrong. (2) "the redirect chain and consent outcome are recorded per site" is true of discovery (discover.mjs:155-158) and false of the sweep:capture.js:43-54readspage.url()only to re-run the SSRF check and then discards it, andconsentDismissedis assigned insrc/and read nowhere — it never reaches the facts object,facts_jsonor D1. There is therefore no per-site consent outcome to audit for any of the 619 unreadable run-1 rows, on a scanner whose own source calls cookie walls the primary blocker to reading a widget's first message.
What discovery never does: open a chat panel's composer, type, or send anything (hard rule 1). The only page interactions are the consent-accept click (documented scanner behaviour) and the synthetic scroll.
5. Pre-registered analysis rules (declared before any data existed)#
- Headline denominator: cohort sites where the sweep confirms the widget and could read its first-interaction surface. Numerator: graded "not detected". Every other terminal state is published as its own named rate (widget absent at sweep, panel unopenable, could-not-verify, robots-disallowed, nav error, bot-challenge) — never folded into the headline, never silently dropped.
- Language coverage: RESOLVED 2026-07-27 (the founder's call) — the disclosure lexicon was extended pre-baseline from 6 to 24 languages in pack
2026.07.9, which is 23 of the 24 official EU languages (Irish is not covered) plus Catalan: terms proposed per-language and adversarially reviewed for cross-language collisions (proposals, kill list and reasoning archived intools/study/lexicon-proposals.json; 489 tests green, 56 of them new per-language suites). The review also found and fixed two pre-existing false-positive classes in production: strong terms colliding with ordinary Portuguese/Romanian copy ("o assistente ia transferir…", "un agent ia legătura"), and weak-term corroboration via bare acronyms ("ai"+"ia" on Romanian, "robot"+"ki" on Slovenian — "ki" is the Slovenian relative pronoun); bare acronyms no longer corroborate. The "could not assess (language)" rule still applies to whatever remains uncovered (non-EU languages, undeclared lang markup). > CORRECTED 2026-08-02 (§21a). This bullet read "extended pre-baseline from 6 to > 24 EU languages", and the claim was repeated on the published draft as "the > disclosure lexicon covers the 24 official EU languages in rule pack 2026.07.9". > Both are false, and a reader disproves them by counting keys.TERMS_BY_LANG> (src/lib/lexicon.js) has 24 keys — bg ca cs da de el en es et fi fr hr hu it lt > lv mt nl pl pt ro sk sl sv. Irish (ga) is absent; it has been an official EU > language since 2007 and its derogation ended 1 January 2022, so "official" is not > arguable here. Catalan (ca) is present and is not an official EU language. > There is no recorded decision to exclude Irish — the count 24 was almost certainly > read off the key count and never checked against the official list. Live > exposure: 12.iesites in the cohort have confirmed widgets, 4 of them in the > readable subset.gais not being added: the instrument is frozen under §6 > until run 2 completes, and adding a language mid-study would confound the delta > with an instrument change — which is exactly the §13a defect §5's freeze exists to > prevent. The correction is to the description, not to the pack. > > The "could not assess (language)" rule is a further UNMET pre-registration > requirement (§21d/5), not a live rule. It was designed, pre-registered here, and > never implemented on the sweep path.detectDisclosure> (src/lib/detectors.js:298-301) computesunclearfrom exactly three inputs — > whether a term matched, whether a persona was detected, andpanelText.trim() > .length >= UNCLEAR_MIN_CHARS(120). Language is not among them.page_langsis > consumed at one place only, the pro-tiereu50.1c.language-coveragerule, which > grades NA when!disclosure.found_at_first_interaction— i.e. it only fires > once a disclosure has already been found, so it cannot catch the case this rule > was written for.grep -rn "could not assess" src/returns zero hits. The > consequence runs in the direction this bullet claims to protect against: a short > greeting in Irish, Russian or Norwegian is published as "No disclosure > detected", not as unassessable. §7's language-exposure counts below are an > offlinehtml_lang-vs-SUPPORTED_LANGSexposure tally, not a grading > outcome, and must not be read as evidence that the rule fired. - Delta: computed as paired within-site transitions on the identical cohort (McNemar-style), with 95% CIs on every published share. Never a difference of two independent percentages.
- Discovery→sweep agreement (spec fixed 2026-07-27, before any subsample data): before the baseline, n ≥ 150 cohort sites stratified across the four signal classes with generic-launcher-only oversampled are run through the production sweep; disagreements are human-adjudicated in a real browser. Gate: widget-presence agreement < 85% overall, or < 80% for the generic-launcher-only class, blocks the baseline until the cohort definition is re-examined. Per-class rates published here.
- Run-2 recomposition (pre-registered): run 2 re-records
final_pldper site and re-appliesredirected_off_frame/duplicate_final_site/ robots from its own run's observations. The headline delta is computed only over sites in-frame and readable in BOTH runs; every composition change (left-frame, newly-unreachable, newly-robots-disallowed, newly-off-frame) is a named attrition class published with the delta, never silently dropped. - Test-retest noise floor (pre-registered): during the baseline window the full cohort is swept twice in the same day and hour band; the per-site scan-to-scan discordance rate is published as the measurement noise floor. A run-1→run-2 delta smaller than the noise floor is reported as not distinguishable from measurement noise.
- Sweep timing (pre-registered): both runs execute inside EU business hours (09:00–17:00 CET), run 2 in the same weekday/hour band as run 1 — availability-gated launchers (tawk
isChatHidden, staffed-hours widgets) otherwise bake the hour choice into the denominator. Per-scan timestamps recorded. - Per-scan evidence archive (pre-registered; the harness must implement): every study scan in both runs stores immutably under its scan ID: the verbatim first-interaction surface text, panel + above-fold screenshots, rendered DOM, observed network hosts (SDK URLs carry vendor/version hints), and grading pack version. This is the delta's audit trail — a vendor shipping a default-disclosure update between runs must be attributable from evidence, not memory.
- Cohort-quality audit (spec fixed; run after the extension completes, before cohort lock): a real-browser screenshot audit of (a) ~60 generic-launcher-only admits — precision; (b) every
sdk_inertsite — the rule's false-negative check; (c) ~50no_widgetsites stratified toward low-prevalence TLDs — recall. Rates published here.
6. Instrument freeze (extends STUDY_PIPELINE.md §8b)#
The §8b freeze covers src/packs/. Extended here to the whole measurement instrument: no deploys touching capture, interaction (consent lists included), robots, fingerprints/recipes, or detectors between the baseline sweep and run 2. The deployed commit hash and Worker version ID are recorded alongside each run. An emergency deploy in the window must be recorded here and the baseline re-graded under the same code before the delta is computed.
Rule pack at baseline: 2026.07.9 (the 24-language lexicon extension; supersedes the earlier 2026.07.8 note — the extension landed BEFORE the baseline, inside the freeze window's opening). Run 2 uses the identical URL list including run-1 failures.
6b. Discovery instrument notes (2026-07-27 run)#
Recorded honestly because they are part of the instrument, and the NDJSON keeps the full attempt history (append-only; assembly takes the last row per domain):
- The run needed two mid-flight instrument fixes, both to throughput/robustness, none to classification: a 45s CDP protocol timeout + 100s per-site hard cap (wedged-page main threads had stalled workers at puppeteer's 180s default), and a browser POOL (a single shared Chrome saturated under 24 workers' context churn until
createBrowserContextitself timed out, instantly failing ~2.5k candidates). - Retry policy (stated): instrument-failure statuses (
browser_error,hard_timeout,timeout) were re-attempted once the instrument was healthy; site-behaviour statuses (dns, robots-disallowed, bot-challenge, rendered) were never re-attempted. Final status = last attempt. - Cohort assembly is
tools/study/assemble-cohort.mjs(deterministic, re-runnable); thenon_eu_operatoradjudication pass runs after it, before the sweep. - Majority of widget detections are generic-launcher-only (vendor unknown). The pre-registered discovery→sweep agreement subsample (§5) is the check on that selector's precision — to be run on the harness before the baseline.
7. Counts — FINAL, cohort locked (2026-07-27; assembly: assemble-cohort.mjs, lock: cohort-lock.mjs)#
Final frame after the 16,000-candidate extension, offline-window salvage, and adjudication. Every candidate accounted for: 2,559 non-rendered + 13,441 rendered (998 of those nav_timeout_partial, rendered-class per §6b) = 16,000.
| Stage | Count | Notes |
|---|---|---|
| Candidates | 16,000 | Tranco XN23N EU-TLD ranks 45–164,793 |
| Rendered | 13,441 | incl. 998 nav_timeout_partial |
| robots-disallowed | 163 | + 3 more at final origin, excluded at lock |
| Unreachable (dns/tls/conn) | 1,139 | |
| Timeout (after retries) | 360 | |
| Bot-challenge | 720 | own exclusion class |
| Nav error (after retries) | 58 | |
| Residual instrument failure | 119 | 0.7% of frame |
| Redirected off-frame | 708 | |
| Duplicate final site | 448 | |
| On-frame unique rendered | 12,285 | prevalence denominator |
| Widget detected | 1,150 | 9.4% (9.2% at the 10k checkpoint — stable) |
| non_eu_operator (adjudicated) | 5 | listed in adjudication-and-review.json |
| Cohort FINAL (run-1 list) | 1,142 | tools/study/cohort-final.csv — frozen |
- Named vendors (331): Zendesk 92 · tawk.to 40 (all 40 exist only via the §4 UA fix) · LiveChat 37 · HubSpot 33 · Intercom 28 · Userlike 21 · Ada 19 · Freshworks 16 · Smartsupp 16 · Chatwoot 10 · Crisp 8 · Zoho 6 · Gorgias 6 · Help Scout 5 · LivePerson 1 · Qualified 1. Unknown-vendor generic-launcher: 811 (71%) — precision measured by the pre-registered agreement subsample before the baseline.
- Per-country cells ≥ 30 (publishable): de 292 · fr 148 · it 96 · nl 72 · pl 70 · cz 52 · se 47 · es 36 · ro 36 · dk 35 · eu 35 · at 30.
- Language exposure under pack 2026.07.9 (24-language lexicon): 1,113 of 1,142 lexicon-supported · 28 undeclared markup · 1 unsupported (a Russian-language site — correctly outside the lexicon's set). The "could not assess (language)" bucket went from ~a third of the cohort to effectively zero; §5's rule still applies to the residue. > CORRECTED 2026-08-02 (§21a). Two things in this bullet. (1) "the EU-24 set" is > the wrong name for the lexicon's 24 keys: they are 23 official EU languages plus > Catalan, with Irish absent (§5's correction). Russian is outside the set either > way, so the adjudication of that one site stands; only the label was wrong. > (2) "The 'could not assess (language)' bucket" is not a grading bucket and never > was. These three counts are an offline tally of recorded
html_langagainst >SUPPORTED_LANGS, computed at cohort assembly — an exposure measure. No sweep > row was ever graded "could not assess (language)", because the rule was never > implemented (§21d/5). The residue — 28 undeclared + 1 unsupported = 29 cohort > sites — does not route to an unassessable outcome; it routes to whatever >eu50.1.chat-disclosureproduces, which for an unmatched greeting is "no > disclosure detected". Read this bullet as "the lexicon now has terms for the > declared language of 1,113 of 1,142 cohort sites", and nothing more.
7-old. First-pass counts (10,000-candidate checkpoint, superseded)#
Every candidate accounted for: 2,415 non-rendered + 7,585 rendered = 10,000.
| Stage | Count | Notes |
|---|---|---|
| Tranco rows | 1,000,000 | list XN23N (30-day window) |
| EU TLD set | 128,124 | |
| Candidates | 10,000 | ranks 45–104,186 |
| Rendered cleanly | 7,585 | after stated one-retry policy (§6b) |
| robots-disallowed | 104 | stated exclusion (hard rule 4) |
| Unreachable (dns/tls/conn) | 636 | dead/parked domains, expected in this rank range |
| Timeout (after 2 attempts) | 1,084 | |
| Navigation error | 70 | |
| Bot-challenge | 453 | own exclusion class, never "no widget" |
| Residual instrument failure | 68 | browser_error 33 + hard_timeout 35, after retry |
| Redirected off-frame | 411 | final PLD left the EU TLD set |
| Duplicate final site | 207 | TLD mirrors collapsed to best rank |
| On-frame unique rendered | 6,967 | denominator for the prevalence stat |
| No widget detected | 6,323 | |
| Widget detected → cohort (pre-adjudication) | 644 | 9.2% of on-frame unique rendered |
Cohort composition (tools/study/cohort-summary.json):
- Named vendors (170): Zendesk 44 · HubSpot 19 · LiveChat 18 · Intercom 16 · tawk.to 16 · Freshworks 11 · Ada 11 · Userlike 9 · Chatwoot 7 · Smartsupp 5 · Crisp 4 · Zoho SalesIQ 4 · Gorgias 2 · Help Scout 2 · LivePerson 1 · Qualified 1. (The 16 tawk.to sites exist in this cohort only because of the UA fix in §4 — the headless UA read all of them as inert SDKs.)
- Unknown vendor, generic launcher only: 474 (74%). The dominant category — custom/regional widgets. Signal mix: 167 strong fingerprints, 110 vendor recipe launchers, 25 API-liveness confirms. The generic selector's precision is exactly what the pre-registered discovery→sweep agreement subsample (§5) measures; treat the 644 as recruitment, not as the final denominator.
- Per-country cohort cells ≥ 30 (publishable under the min-cell rule): .de 154 · .fr 83 · .it 62 · .pl 40 · .nl 35 · .cz 31. All other countries fall below 30 and will be reported as "n too small".
7a. non_eu_operator adjudication (2026-07-27; full table in tools/study/adjudication-and-review.json)#
36 lexical suspects flagged across the 644; every one adjudicated from its own imprint/privacy/legal surface (WHOIS where the site blocked fetches). 5 excluded, all high-confidence: two US-incorporated operators on .it read as "IT" · one US-incorporated operator on .hr read as "HR" · one US operator on a .ee link-shortener pattern · one .eu news site operated from outside the EU. 31 kept with recorded evidence — including EU operators behind apparent hacks (an .it domain run by a Slovenian d.o.o.; a .eu domain run by a MiCAR-licensed Austrian GmbH) and EU mirrors (a Dutch N.V. running matching .de/.it/.se storefronts). Inconclusive evidence kept the site (exclusion requires positive proof, per the pre-registered rule). Cohort after adjudication: 639, superseded by the §3a extension run's assembly.
7b. Frame-as-built review (2026-07-27; findings in adjudication-and-review.json)#
Four hostile lenses over the built frame + triage of the 28 archived serious findings. Consolidated must-fix items and their dispositions:
- Timeout salvage (
nav_timeout_partial) — implemented in discover.mjs; the 1,084 timeouts + 70 nav errors re-attempted under it (§6b retry policy). - Rank extension to 16,000 — pre-registered §3a, running.
- robots.txt at the FINAL origin — to run at cohort lock: every cohort site's final origin re-checked with the production matcher; violations evicted as
robots_disallowed_at_final, count published. - Run-2 recomposition, noise floor, sweep-hour window, evidence archive, agreement gate, cohort-quality audit — pre-registered in §5.
- Public precommitment timestamp before the baseline — needs the founder (push the frozen bundle to a public repo/tag or OSF; hashes will be recorded here).
- Lexicon extension decision — needs the founder (§5); exposure quantification from recorded
html_langto be added at cohort lock.
Next (harness day): /api/_study_run (STUDY_PIPELINE.md §3), deploy + verify the frozen-instrument commits under §6, then the §5 agreement subsample and quality audit before the baseline.
8. Pre-flight adversarial review (2026-07-27)#
Five hostile lenses over the frame before the discovery run; 62 findings (16 blocker / 28 serious / 18 minor), archived in tools/study/preflight-findings.json. Blockers fixed pre-run: consent-language coverage, cross-origin CMP iframes, tawk-class API liveness probes, headless-UA gating, bot-wall classification, redirect-chain + vantage + html_lang recording, candidate count power-sized, estimator + paired delta + adjudication + min-cell rules pre-registered, freeze scope extended. Deferred (recorded above as limitations or open decisions): EU-vantage scanning, and running discovery through Browser Rendering itself. The lexicon language extension — deferred at pre-flight — was subsequently RESOLVED pre-baseline: the founder directed it 2026-07-27 and it shipped as pack 2026.07.9 (§7's language-exposure table shows the effect). The serious/minor findings feed the full §7 review before publishing.
9. Harness (2026-07-27) — built, reviewed, deployed, verified#
POST /api/_study_run is live (STUDY_PIPELINE.md §3; commit 6ed30fc, Worker version fb71ecfd-6a18-415e-a3a4-181e8fab6b93 — the §6 frozen instrument for the baseline). End-to-end smoke verified in production: cohort smoke, one scan of example.com, graded under pack 2026.07.9 in ~24s with the full per-scan evidence archive (10 artifacts) stored under its scan id. Unauthenticated requests to the route are byte-identical to a nonexistent path (verified live). The bearer secret lives in the STUDY_SECRET Worker secret, with a local driver copy at .study-secret (gitignored).
Decisions recorded (each is part of the instrument):
- Cohort tag =
scans.cohort(migration 010, additive). Reserved slugs:smoke(never aggregated),agreement(§5 subsample),run1,run1-noise(same-day noise floor),run2. Every aggregate selects on exactly one slug. - Retention exemption: cohort-tagged scans are exempt from the tier ladder so the run-1 evidence archive survives to the delta and beyond; removed by hand at teardown.
- Robots tri-state at enqueue:
robots_disallowedandrobots_unreachable(RFC 9309 §2.3.1.4 — 5xx/network/timeout is never counted as consent) are separate, driver-visible exclusion classes; 4xx (= no robots.txt) remains an allow (§2.3.1.3). Product paths unchanged. - Capacity deaths are distinguishable: a study scan that dead-letters or is reconciled while
throttledterminates aserror_capacity(study-only status); those are re-run per STUDY_PIPELINE.md §5, never published as site failures. - Sweep pacing:
scan-jobsconsumer pinned atmax_concurrency: 10(~0.4 browser launches/sec vs the 1/sec grant; ~48 min per full-cohort sweep, inside the §5 hour band). The original "let Queues autoscale" plan is corrected in STUDY_PIPELINE.md §2. - Runaway wall: 6,000 total study scans, enforced as a D1 count of cohort rows.
- Isolation: study scans never touch the 24h visitor result cache (read, write, or failure-path purge) and never link
sitesrows.
Adversarial review of the harness diff before deploy: 15-agent workflow, 11 raw findings, 7 confirmed (1 blocker: the autoscale/launch-rate collision; 2 serious: robots-unreachable-as-allow, capacity deaths indistinguishable from site failures), all fixed pre-deploy. 507 tests green (18 harness-specific).
10. Pre-baseline gate results (2026-07-27) — THE AGREEMENT GATE FAILED#
Recorded before any remedial work, so the sequence is auditable: this is the number the pre-registered instrument produced on its first run, not the number after tuning.
10a. Discovery→sweep agreement — FAIL#
150 cohort sites sampled deterministically (tools/study/agreement-sample.py, seed agreement-2026-07-27), disjoint strata by recorded precedence, generic-launcher-only oversampled per §5. Swept through the production harness as cohort agreement (6 excluded at enqueue as robots_unreachable, 7 not terminal at read time).
| Signal class | Confirmed / graded | Agreement | Gate |
|---|---|---|---|
| Overall | 114 / 137 | 83.2% (95% CI 76.1–88.5) | ≥85% → FAIL |
| generic_launcher_only | 79 / 94 | 84.0% (75–90) | ≥80% → pass |
| fingerprint_strong | 29 / 37 | 78.4% (63–89) | — |
| recipe_launcher | 2 / 2 | 100% | — |
| other | 4 / 4 | 100% | — |
The shape is the opposite of the one anticipated. The gate was designed around the suspicion that the generic launcher selector over-recruits; that class passed its own threshold, while the strongest evidence class — sites whose vendor SDK was fingerprinted outright — agreed least (78.4%). A hypothesis worth testing before the baseline: strong fingerprints skew toward large commercial sites, which are also the sites most likely to serve the sweep a consent wall or bot challenge that discovery did not meet. Untested as of this writing; it must not be assumed.
Consequence, per the pre-registration: the baseline is blocked until the cohort definition is re-examined. §5 says so in advance and it is being honoured.
10b. The larger finding: first-interaction surfaces are mostly unreadable#
From the same run, and not something the gate was designed to catch:
| Count | Share of widget-confirmed | |
|---|---|---|
| Widget confirmed by the sweep | 114 | — |
| Panel opened | 21 | 18.4% |
| First-interaction surface readable (headline denominator) | 12 | 10.5% |
At cohort scale that projects to roughly 120 assessable sites out of 1,142, which is fatal in two ways. It collapses the per-vendor and per-country cuts below the pre-registered min-cell of 30, and — the serious one — the sites whose panels open are not a random subset of the cohort. They are the sites with simpler widgets and weaker consent walls. Computing the headline share over them would embed exactly the selection bias §4 of the pipeline doc warns about, and the first competent reader would find it.
Failure causes tallied across all 93 failures (tools/study/agreement-panel-errors.json): 38 panel did not open after the click · 19 launcher not clickable · 12 launcher covered by an overlay · 2 panel selector matched the closed launcher · 2 recipe stale · rest assorted.
10c. Instrument fix applied pre-baseline (and why it was allowed)#
The §6 freeze runs from the baseline to run 2. The baseline has not run, so this is the last legitimate point to fix the instrument; doing it after would confound the delta.
waitForClickableSafe (src/lib/interaction.js): the launcher click resolved document.querySelector's first match, and broad launcher selectors routinely match a zero-size decoy — a skip-link anchor, a collapsed 1366x0 wrapper, a hidden mobile duplicate — sitting ahead of the real button. Puppeteer then threw "Node is either not clickable", which is 19 of the 93 failures above. The launcher wait now polls for the first match with a real box and visible computed styles, and launcherOcclusion receives that handle instead of re-querying the decoy. Verified live on four affected sites: all four now click the real launcher.
This fixed the click, not the read. Re-probed, those sites fail one step later, at "panel did not open after launcher click" — the generic panel selector does not match what opened. Two mechanisms confirmed by hand: panels rendered inside shadow DOM (one large retail site hosts its assistant in a cms-iframe custom element with a shadow root), and launchers that navigate to a dedicated page instead of opening an in-page panel. Neither is a small fix.
10d. Where this leaves the study — decision needed#
The baseline cannot run on 30 July as planned. The options, honestly stated:
- Instrument work first, baseline later. Shadow-DOM-piercing panel detection, navigation-following widgets, and per-vendor recipes for the long tail. Days, not hours, and it moves the baseline past 2 August — which the original §9 reasoning already preferred ("on the day the law took effect" is a measurement, not a prediction).
- Reframe around what IS measurable at scale. Widget prevalence, vendor distribution, and coverage itself — "on N% of EU sites running a chat widget, an automated visitor cannot even reach the first message" is a real, defensible, previously unpublished finding, and it is exactly the honesty metric §6 of the pipeline doc already wanted.
- Publish the baseline with a ~10% denominator and the bias stated. Not recommended: the caveat would have to be larger than the headline, and it would not survive review.
Recommendation: 2, with 1 running behind it. The coverage finding is publishable now, does not depend on the panel-open rate improving, and gets stronger — not weaker — the more the instrument is later fixed.
11. Cohort-quality audit (2026-07-27) — first pass, and why it is a LOWER BOUND#
160 sites captured in a real browser across the three §5 strata (60 generic-launcher-only admits · 50 sdk_inert · 50 no_widget, sampled deterministically by tools/study/audit-sample.py logic embedded in audit-sample.csv), then judged visually — "is a chat launcher visibly present?" — by 14 independent agents, blind-instructed to ignore the stratum label in the filename.
Reliability check (unplanned, and the most reassuring number here): a slicing bug double-judged 10 screenshots by two independent judges. Agreement was 10/10. The visual method itself is reliable; what follows is limited by the capture, not the judging.
| Stratum | n | launcher visible | no launcher | unclear | consent wall blocking |
|---|---|---|---|---|---|
| (a) generic-launcher-only admits | 60 | 23 | 23 | 14 | 19 |
(b) sdk_inert | 50* | 9 | 38 | 13 | 13 |
(c) no_widget | 40* | 3 | 21 | 16 | 10 |
\* run-1 slicing bug (below) judged 10 of (b) twice and missed 10 of (c); counts are of judgements, not sites, for those two rows.
These numbers must NOT be read as cohort precision. The capture instrument was weaker than the instrument under test: audit-capture.mjs originally hand-rolled its consent click (first matching CMP selector, top frame only) instead of calling production's dismissConsentBanner, which pierces shadow DOM, matches accept text in 24 languages, and reaches cross-origin CMP frames. On stratum (a) — sites where discovery had found a launcher — this capture reproduced a launcher match on only 17 of 60. An audit whose instrument is weaker than the one being audited measures the audit, not the cohort. The "unclear" column is the direct evidence: 43 of 160 judgements were unjudgeable, and consent walls were blocking on 42.
Fixed (commit 9887851) and re-captured with the production dismisser; the re-judged figures supersede this table and are recorded below when complete.
Two instrument defects found by running the audit, both fixed:
- Agents writing into the directory they slice from. Judges were told to derive their file list with
ls … | sed -n 'FROM,TOp'. Some wrote zoom-crops back into that same directory, which reordered the listing mid-run and shifted every later slice — 10 files judged twice, 10 never judged. Slices must come from a frozen list, never a livels. - Carriage returns in filenames. Python's
csv.writeremits\r\n; asplit("\n")in the capture script left the\ron the last CSV field, which rode into every screenshot filename.lsrenders it invisibly whilestat/opencannot resolve the path — files that appear to exist and are unreadable. Worth remembering for any future tooling that moves names between the Python and Node halves of this repo.
12. Cohort-quality audit — FINAL (2026-07-27, production-consent capture)#
Supersedes §11. Same 160 sites, re-captured with production dismissConsentBanner, judged by 14 independent agents blind to the stratum label. Full judgements with per-site notes in tools/study/audit-judgements.json; capture record in audit-capture-results.json. Coverage verified: 160 judged, 160 unique, no duplicates, none missed.
The capture fix mattered enormously. Stratum (a) launcher DOM-match reproduction went from 17/60 to 49/60 once consent was dismissed the way production does it. §11's numbers were an artefact of the audit script and should not be cited.
| Stratum | n | launcher visible | none | unclear | rate among decidable |
|---|---|---|---|---|---|
| (a) generic-launcher-only admits — precision | 60 | 23 | 26 | 11 | 46.9% (95% CI 33.7–60.6) |
(b) sdk_inert — false-negative check | 50 | 5 | 37 | 8 | 11.9% (5.2–25.0) |
(c) no_widget — recall gap | 50 | 3 | 38 | 9 | 7.3% (2.5–19.4) |
Page state across all 160: 132 judgeable · 19 consent wall still blocking · 8 error/blank · 1 bot challenge. 28 of 160 judgements were "unclear" and are excluded from the rates above rather than assumed either way. Amended 2026-07-28 (§15b): the exclusion is not neutral — unclears are dominated by page states that HIDE launchers (consent walls, errors, bot challenges), so it plausibly biases each rate down, and the estimator publishes all-no/all-yes and rate-multiplier sensitivity bounds alongside the CI.
12a. What this says about the frame — it is contaminated AND incomplete#
Precision. A generic-launcher admit shows a real chat launcher under visual inspection 46.9% of the time. The DOM selector matched an element on 49 of 60, but that element is visually a chat launcher in fewer than half the decidable cases. The judges' notes name the recurring false positives: accessibility widgets, "Kontakt"/"Contact" nav links that carry a speech-bubble icon, back-to-top arrows, newsletter buttons, search fields. This is precisely what the pre-registered audit existed to detect, and it fired.
Generic-launcher-only is 71% of the cohort (807 of 1,142), so roughly 428 cohort sites likely have no chat widget at all.
Recall. 7.3% of no_widget sites show a launcher on inspection. Projected across 11,135 widget-negative sites that is roughly 815 sites wrongly excluded — comparable in size to the entire cohort.
Illustrative correction (Rogan-Gladen-style, assumptions stated, NOT a result): apparent prevalence 9.36%; correcting admits by measured precision and adding the projected misses gives ~12% true prevalence across a range of assumptions for the unaudited classes. The two errors do not cancel — they compound the uncertainty.
Retracted 2026-07-28. This paragraph originally ended "…and the corrected figure is materially different from the published 9.4%." The variance-propagated estimator (§15b) shows the correction (~230–300 frame-sites) sits inside its own 95% margin at any feasible recall-sample size, so "materially different" overclaimed. The honest statement is the §15 headline-1 framing: a bias adjustment whose CI may include zero. Recorded as a retraction, not silently rewritten, per this log's correction style.
12b. Gap in this audit, stated plainly#
fingerprint_strong precision was never measured. §5 specified precision auditing for generic-launcher-only admits, and that spec was followed. But §10a then found fingerprint_strong had the worst sweep agreement (78.4%), which makes its precision the obvious next question and it is unanswered. Any corrected prevalence therefore rests on an assumed value for 30% of the cohort. This should be audited before anything is published.
12c. Consequence#
Combined with §10, three independent measurements now say the frame is not what the locked cohort claims: the agreement gate failed (83.2%), first-interaction surfaces are readable on 10.5% of widget-confirmed sites, and the dominant admission rule is ~47% precise. The recommendation in §10d is unchanged and strengthened — but note that even the prevalence headline now requires the audit-based correction and its error bars, rather than the raw 9.4%. The honest version of this study is increasingly a study about measurement: what an automated visitor can and cannot establish about AI disclosure at EU scale, with the error rates published. That is a real contribution and it is the one the data supports.
13. Post-fix re-measurement (2026-07-27, Worker 4af9624f)#
The launcher fix (§10c) was deployed — fb71ecfd → 4af9624f, which is now the instrument of record — and the identical 150 agreement sites were re-swept as cohort agreement2. This supersedes §10a/§10b as the current instrument reading. Both runs are retained; the earlier numbers stand as the pre-fix record.
Corrected 2026-07-27, later the same evening. The post-fix column was first computed while 6 scans were still in flight; the sweep has since reached terminal state (140 graded, 4
error— the errors are an exclusion class, counted, not hidden). Figures below are the COMPLETE sweep. Nothing material moved (80.3→80.7, 26.4→26.5, 16.4→16.8) and every conclusion stands, but a number computed mid-sweep should never have been written down as final, and the correction is recorded rather than silently overwritten.
| Measure | Pre-fix (fb71ecfd) | Post-fix (4af9624f), complete sweep |
|---|---|---|
| Widget-presence agreement | 114/137 = 83.2% | 113/140 = 80.7% (CI 73.4–86.4) — still FAILS the ≥85% gate |
| ‑ generic_launcher_only | 84.0% | 82.3% (its own ≥80% bar: still passes) |
| ‑ fingerprint_strong | 78.4% | 73.7% |
| Panel opened | 21/114 = 18.4% | 30/113 = 26.5% |
| Readable first-interaction surface | 12/114 = 10.5% | 19/113 = 16.8% |
The fix did what it claimed and not more. Panel opening improved by 8 points and readable surfaces by 6 — real, and in line with the 19-of-93 "not clickable" failures it targeted. It does not come close to rescuing the headline: 16.8% readable still means a denominator near 190 of 1,142 at cohort scale, still drawn from the sites with the simplest widgets. §10d's conclusion is unchanged.
Agreement fell slightly rather than rising. The fix can only add launcher clicks, so the 2.5-point drop is run-to-run variance, not regression — which the next section quantifies.
13a. Test-retest noise floor (partial, from the two agreement runs)#
§5 pre-registers a same-day double sweep to establish the measurement noise floor. The two agreement runs deliver a first read of it on 135 sites graded in both:
Recomputed on the complete sweep: 137 sites graded in both runs.
| Discordant | Rate | |
|---|---|---|
| Widget presence | 5 (4 lost, 1 gained) | 3.6% |
| Readable surface | 11 (2 lost, 9 gained) | 8.0% |
Caveat that must travel with these numbers: run 2 used the post-fix code, so this confounds the instrument change with genuine noise and is therefore an upper bound on the true noise floor. The asymmetry supports that reading — readable-surface changes are 8 gained vs 2 lost, i.e. mostly the fix working, whereas widget presence is 4 lost vs 0 gained, i.e. sites that simply did not present a widget on the second visit.
Consequence for the delta: a run-1 → run-2 difference smaller than roughly 4 points on widget presence, or 8 points on readable surfaces, is not distinguishable from measurement noise. A clean same-code double sweep is still owed before any delta is published.
14. fingerprint_strong precision — the §12b gap, now closed (2026-07-27)#
60 fingerprint_strong cohort admits sampled deterministically (seed audit-fpstrong-2026-07-27, tools/study/audit-sample-fpstrong.csv; vendor mix Zendesk 19, LiveChat 9, Freshworks/tawk.to/HubSpot 5 each, tail beyond), captured with the production consent dismisser and judged blind by 5 agents. Coverage clean: 60 judged, 60 unique, none missed. Judgements in tools/study/audit-judgements-fpstrong.json.
| Class | launcher visible | none | unclear | precision (decidable) |
|---|---|---|---|---|
fingerprint_strong | 38 | 14 | 8 | 73.1% (95% CI 59.7–83.2) |
generic_launcher_only (§12) | 23 | 26 | 11 | 46.9% (95% CI 33.7–60.6) |
The vendor fingerprint is substantially better than the generic launcher — and still not clean. Roughly a quarter of sites whose vendor SDK was fingerprinted show no chat launcher to a visitor. The judges' notes give the mechanism: the SDK ships in the page bundle while the widget is disabled, out of staffed hours, or gated to logged-in users. This is a real finding in its own right — "the vendor SDK is present" and "a visitor can chat" are not the same claim, and only the second one matters for Article 50.
It also explains §10a's inversion. fingerprint_strong had the worst sweep agreement (73.0%) precisely because ~27% of those admits have no reachable launcher for the sweep to confirm — the two measurements agree, and the class is weaker evidence of a live chat surface than its name suggests.
14a. Corrected prevalence — both precisions now measured#
| Apparent prevalence | 9.36% (1,150 / 12,285) |
| True positives among admits | 807 × 46.9% + 343 × 73.1% ≈ 629 — so ~521 admitted sites likely have no widget |
| Missed among widget-negatives | 11,135 × 7.3% ≈ 815 |
| Corrected prevalence | ≈11.8% (1,444 / 12,285) |
Remaining assumption, now small: api_live and recipe_launcher (~15 cohort sites) take the fingerprint_strong value as a proxy; too few to move the result.
Still owed before this is publishable: proper variance. The figure above is point arithmetic on three measured rates, each with a wide CI (the widest, the recall gap, rests on 41 decidable sites). A Rogan-Gladen estimator propagating all three CIs — and a larger recall sample, which is the binding constraint — is required before any prevalence number is published. Do not publish 11.8% as though it were precise.
15. DECISION (2026-07-28, the founder) — reframe adopted; baseline this week#
Recorded before any post-decision data exists, so the reframe is itself pre-registered.
§10d is resolved: option 2 — reframe around what is measurable at scale — with option 1 (instrument work) running behind it. The published piece leads with:
- Audit-corrected widget prevalence with propagated error bars (the §14a estimator, upgraded per §15b below) — never the raw 9.4%. Reported as a bias adjustment whose CI may include zero: the correction is presented as what the audits imply plus its uncertainty, not as a precise new prevalence (see the §15b power note — the recall term dominates the variance at any feasible sample size).
- Vendor distribution among sweep-confirmed sites, with the §14 caveat attached: "vendor SDK present" and "a visitor can chat" are different claims, and the measured gap between them (~27% for fingerprinted sites) is itself reported. (Corrected 2026-07-31: this read "among admits", which the code has never done —
aggregate.pybuilds the vendor table over the sweep-confirmed rows and the rendered header reads "Share of confirmed". The code wins and the doc is matched to it, not the reverse. Consequence, carried into §18d: a sweep confirmation is an observation made on the sweep day, so the vendor distribution is weekend-sensitive under a Saturday baseline — unlike an admit, which is a fixed 27–28 July artefact.) - Coverage — the share of sweep-confirmed widgets where an automated visitor cannot reach the first-interaction surface, split by failure cause (consent wall, launcher not clickable, panel did not open, shadow DOM/navigation, and "no chat widget present — admission false positive", since §12/§14 prove sweep-style confirmation admits non-chat elements at scale and those land in the "panel did not open" bucket). The denominator is stated as detector-confirmed, never asserted as "sites running a chat widget". This is the honesty metric of pipeline §6 promoted to a headline finding.
- The disclosure-state distribution is still reported, but explicitly scoped to the readable subset with its selection bias stated, below the fold — per §10d option 3's logic it cannot lead.
Timing (the founder's call, same date): the baseline runs THIS WEEK, before 2 August. Target Thursday 30 July, inside the pre-registered 09:00–17:00 CET band. Per §5 the noise floor is a same-day full-cohort double sweep, so baseline day is two sweeps of the identical 1,142-site list (cohort slugs run1, run1-noise), ~48 min each at max_concurrency: 10. Run 2 lands Thursday 6 August — same weekday/hour band per §5, and after Article 50 applies, so the delta brackets the deadline. The §6 instrument freeze starts at the first run1 enqueue. Instrument of record: Worker 4af9624f, pack 2026.07.9, unchanged.
Superseded on two points, both by later sections — this paragraph is kept as the record of what was decided on 2026-07-28, not as current fact. (1) The dates moved one day: baseline Fri 31 July, run 2 Fri 7 August (§17 — date-only amendment, weekday pairing preserved) — and then moved again, after Friday was missed exactly as Thursday had been: baseline Sat 1 August, run 2 Sat 8 August (§18). That second move keeps the pairing but changes the weekday itself, and unlike §17 it is not free — §18d writes the price down before the data exists. (2) The instrument of record is Worker
e9a8965e, not4af9624f(§16g). The "~48 min each" figure is also superseded by the measured ~110 min (§16f).
15a. What the delta means under the reframe#
§5's blanket "readable in BOTH runs" restriction cannot apply to a coverage delta — readability IS the coverage outcome, and conditioning the paired set on the dependent variable would make the delta undefined. The paired denominators are therefore per-outcome, fixed here before the baseline:
- Widget presence: sites in-frame and graded in BOTH runs.
- Panel-opened and surface-readability (coverage): sites in-frame, graded in both, and widget-confirmed in at least one run (noise-floor.py's
confirmed_eitheruniverse), with widget-presence transitions published as their own named class. - Disclosure-state: retains §5's readable-in-BOTH restriction, with its selection-bias framing — the only outcome for which that conditioning is legitimate.
Every §5 attrition class is still published with the delta. Each movement is judged against the noise floor from baseline day's same-code double sweep (superseding §13a's cross-code upper bound).
Scope limit, stated plainly: the sweeps revisit only the 1,142 discovery admits, so the design can observe a cohort site losing (or hiding) its widget but can never observe a discovery-negative site gaining one. "Movement in prevalence" is therefore re-scoped to within-cohort widget-presence movement — at most a one-sided cohort-loss bound; no frame-level prevalence delta will be published.
15b. Pre-registered upgrades landing before the baseline (this decision's work list)#
- Recall sample expansion: ~400 additional
no_widgetsites, same deterministic sampler logic, same production-consent capture, same blind visual judging. The recall gap is the binding CI on corrected prevalence, and it is measured on 41 decidable sites; the pool targets ~370. (Amended when the data landed — §15e: the expansion is a self-weighting SRS and is used ALONE as the primary recall input; pooling the TLD-skewed §12(c) 50 with post-stratification became a sensitivity, since 22 post-strata over ~335 judgements is over-stratified. Both are published and they agree within 0.4 points.) - Power note, recorded before the data exists: at the ~370-decidable target the recall term still carries ~85% of the corrected-prevalence variance (SE ≈ 145 frame-sites on a correction of ~230–300), so the correction is not expected to be distinguishable from zero at 95% — parity with the precision terms would need ~2,500 decidable (~3,000 captured) sites, which is deferred, with this as the stated reason. The published framing follows headline 1: a bias adjustment with honest error bars, not a new point. (Outcome, §15e: borne out — the corrected CI 7.95–15.42% contains the apparent 9.36%.)
- Estimator: Rogan-Gladen-style corrected prevalence with variance propagated from all three measured rates (both precisions + pooled recall), by bootstrap on the judgement counts with Jeffreys-smoothed rate draws (zero-event cells must carry variance). The "unclear" exclusion is bounded by published sensitivity scenarios (all-no / all-yes / unclears-at-2×-decidable-rate) because unclears are dominated by page states that hide launchers — the exclusion plausibly biases every rate down, and the estimator says so. Point arithmetic in §12a/§14a is superseded once this lands.
- Coverage decomposition (pre-registered follow-up): the "panel did not open" failure bucket is contaminated by admission false positives (§12/§14); the already-captured failure-bucket screenshots are to be blind-judged with the §12/§14 method and coverage republished on an audit-corrected basis with error bars, parallel to the §14a prevalence correction. Until then the coverage table carries the admission-false-positive cause row and the detector-confirmed denominator wording of headline 3.
- Aggregation + publish scaffold rewritten to the cuts in 1–4 above. The publish gate (
data/study-results.jsonmust exist AND carry"publish": true) is unchanged. - Precommitment re-anchored (OpenTimestamps) over this section and the reframed aggregation before the baseline runs; the 2026-07-27 manifest and
.otsremain in the repo as the prior record. DONE 2026-07-28:tools/study/precommit.pybuilds a hash-only manifest over 17 files (analysis plan, estimator, aggregation, runbook, frame, judgements, rule pack) and stamps it to three calendars. It took several attempts as pre-baseline hardening continued — the final anchor isprecommit-manifest-6(as of 2026-07-28; the live anchor isprecommit-manifest-11— see §16b and §18); see §16b for the full sequence and the lesson. Verify the live anchor withots verify tools/study/precommit-manifest-11.json.otsonce it is stamped (§16b — it is not stamped as this is written); anchors 1–10 remain in the repo and verify in place.
15c. Disposition of the §5 agreement gate — discharged by re-examination, NOT passed#
The §5 gate fired at 83.2% (§10a) and again post-fix at 80.7% (§13); it has never read ≥85%, and it is not being waved through. Its own release condition is that the baseline stays blocked "until the cohort definition is re-examined." That re-examination is §§11–14: both admission rules' precision measured (46.9% / 73.1%), the recall gap measured (7.3%), and the §10a inversion explained (§14 — ~27% of fingerprinted admits have no reachable launcher, so the two instruments agree; the class is weaker than its name, not the sweep blinder). Its outcome is this reframe: discovery admits are no longer treated as ground truth anywhere — every reframed headline either carries the measured precision/recall correction (prevalence) or states its denominator as detector-confirmed rather than true widget presence (coverage). Per-class agreement rates are published as §5 requires. This is a discharge under the gate's own wording, recorded here so no reader has to discover the gap themselves; agreement remains below the pre-registered bar and the published piece says so.
15d. What existed in D1 before this decision — stated, not certified away#
Disclosure-state grades for the agreement cohorts existed before the 2026-07-28 reframe decision: the gate sweeps were graded under pack 2026.07.9, so §15's "recorded before any post-decision data exists" is true only of post-decision data. For the record, the readable-subset distributions as of the decision: agreement (pre-fix): 12 assessed — 0 detected, 5 not detected, 7 wording-inconclusive (all WARN). agreement2 (post-fix): 19 assessed — 1 detected (PASS), 10 not detected, 8 wording-inconclusive (18 WARN, 1 PASS). This log cannot certify those grades went unqueried before the decision — §10b's readable counts were computed from the same records. The demotion rationale is the §10b/§13 readability collapse, and the disclosed distribution corroborates rather than contradicts it: near-zero detected disclosure is the opposite of an inconvenient primary result to bury.
15e. Recall expansion — RESULT (2026-07-28), and the power note was right#
400 additional plain-negative sites (recall-expansion-sample.mjs, seed recall-expansion-2026-07-28, self-weighting SRS of the 10,935-site remaining universe), captured with the production consent dismisser — 400/400 captured, 0 errors — and judged blind. Coverage verified by merge-recall-judgements.py: 400 judged, 400 unique, none missed, no duplicates. Judgements in recall-expansion-judgements.json.
| Recall gap (SRS, unweighted) | 6.83% — 22 of 322 decidable (95% CI 4.55–10.13) |
| Undecidable | 78 of 400 (19.5%) — consent walls 63, non-rendered 33 |
| §12's estimate it replaces | 7.3% on 41 decidable (95% CI 2.5–19.4) |
Estimator restructured when the expansion landed, and why. The expansion is a self-weighting SRS, so its unweighted rate is design-unbiased and is now the primary recall input. The earlier plan — pool in the §12(c) 50 and TLD-post-stratify — is demoted to a sensitivity: 22 TLD post-strata over ~335 decidable judgements leaves most cells under 20, where a Jeffreys draw on a 0-of-5 cell dominates its own universe weight (that is added noise, not removed bias, and it visibly skewed the interval). The two agree closely — pooled+post-stratified gives 11.01% against the primary's 11.34% — so the choice does not carry the result, which is why both are published.
| Corrected prevalence | |
|---|---|
| Apparent (raw detector) | 9.36% (1,150 / 12,285) |
| Corrected (primary) | 11.3% (95% CI 7.95–15.42), bootstrap median 11.34% |
| Pooled + post-stratified sensitivity | 11.01% |
| Unclear-reassignment bounds | 9.29% (all-no) – 28.51% (all-yes); 2×-rate scenario 13.2% |
The pre-registered power note (§15b) is borne out and must be honoured. The corrected figure's 95% CI contains the apparent 9.36%, so the correction (+2.0 points) is not distinguishable from zero at 95%. That is the outcome §15b predicted before the data existed, and headline 1's framing — a bias adjustment reported with honest error bars, never a precise new prevalence — is what gets published. §12a's retracted "materially different from the published 9.4%" stays retracted; this is the measurement that settles it.
15f. The judging instrument changed mid-measurement — stated and measured#
Items 0–239 of the expansion were judged by one model version and items 240–399 by another (a capacity limit interrupted the run; slices are disjoint and were frozen in advance, so coverage is unaffected). A judging instrument that changes mid-measurement is exactly the kind of thing this log exists to catch, so it was measured rather than noted.
Cross-model reliability check: a 40-screenshot subsample of the first half — deliberately enriched for the rare outcomes (all 10 yes, 15 unclear, 15 no) — was re-judged blind by the second model. Results in crossmodel-reliability-judgements.json:
- Decidable calls agreed 25 / 25 (100%) — every
yesand everynoreproduced. - All 9 disagreements were
unclear→no, and 6 of those were SSL interstitials or blank captures: a convention difference about whether a page that never rendered counts as "no launcher visible" or "nothing to judge".
Consequence, applied uniformly: a page that never rendered cannot evidence launcher absence, so error_or_blank and bot_challenge judgements normalize to unclear in the estimator for every stratum, old and new (NON_RENDERED in prevalence-estimator.py). Impact on the legacy figures is negligible — §12(a) 46.9% and §12(c) 7.3% unchanged, §12(b) 11.9%→12.2%, §14 73.1% unchanged — which is itself evidence the convention was not silently driving those numbers.
The raw between-halves rates (5.1% vs 9.5%, z = −1.53) are not significant and are confounded with alphabetical domain composition, since slices are alphabetical; the same-image comparison above is the stronger evidence and it shows no model effect on decidable calls.
16. Run records — preflight DONE 2026-07-28; sweeps to be filled on the day#
16a. Preflight (executed 2026-07-28, a day early; runbook "Wednesday evening" list)#
This section is a dated snapshot of 2026-07-28, not a live status board: every figure in it is deliberately left as it read that day — including the 507/507 test count, the 341-row D1 total and the "final anchor is precommit-manifest-6" line, all of which have since moved (§17a, §16b, §18f) — so where 16a and a later section disagree, the later section is the current fact and 16a is the record of what was verified before the first sweep was attempted.
Every §6 freeze fact is pinned here BEFORE the first sweep, so the delta can be audited against a stated instrument rather than a remembered one.
| Check | Result |
|---|---|
| Test suite | 507 / 507 pass, 0 fail |
| Live Worker version | e9a8965e-282b-4735-be4b-ae0bbc0fff26 @ 100%, deployed 2026-07-28T19:50:54Z — supersedes 4af9624f; see §16g for the one pre-baseline fix it carries |
| Deployed source == HEAD | yes, re-verified after the §16g deploy |
| Rule pack | 2026.07.9 (latest in src/packs/eu-ai-act-art50/) |
| Cohort rows in D1 before the sweep | agreement 144 · agreement2 144 · smoke 53 (2 preflight + 51 rehearsal, 16f) — 341 of the 6,000 cap |
| Headroom | double sweep adds ≤ 2,284 → ~2,625; run 2 adds ≤ 1,142 → ~3,767. Inside the cap |
| Harness smoke | POST /api/_study_run (cohort smoke, https://example.com) → HTTP 200 in 2.1s, 1 queued, scan 32f5c580-bb03-4194-9a0d-2bd51eea0645 reached graded in ~24 s — matches the ~25 s/scan planning figure (§2) |
| Homepage-only rule | confirmed live: the smoke scan recorded budget_units: 1 |
verified_owner: false / page_cap: 1 | asserted in test/study.test.js (hard rule 1; the nudge must never run on a third-party site) |
| Hour-band filter | aggregate.py:in_band() verified against D1's real timestamp format ('2026-07-28 19:03:41', space-separated, not ISO-T) at the band edges |
| Precommitment | re-anchored repeatedly as hardening continued; final anchor is precommit-manifest-6 — see 16b |
16b. Precommitment: ten anchors, and why there are ten#
Stated plainly because the count looks worse than it is: every anchor below predates any baseline data, the nine already stamped are committed with their .ots proofs, and anyone can diff them against each other and against the tree. None was replaced or rewritten — superseding an anchor here means "a later one exists", never "an earlier one changed". Anchor 10 is listed as the anchor this baseline runs on; per the §18 date change it is cut as a distinct final act after every word of §18 is fixed, so as this paragraph is written it is not yet stamped and nothing here claims otherwise (§18f).
| Anchor | File | Covers |
|---|---|---|
| 2026-07-27 | precommit-manifest.json{,.ots} | pre-reframe frame + method (original record) |
| 2026-07-28 | precommit-manifest-2.json{,.ots} | reframe analysis plan; then drifted by 3 doc files |
| 2026-07-28 | precommit-manifest-3.json{,.ots} | preflight state; then drifted by 1 tooling file |
| 2026-07-28 | precommit-manifest-4.json{,.ots} | post-preflight; then the rehearsal (16f) produced real operational findings |
| 2026-07-28 | precommit-manifest-5.json{,.ots} | post-rehearsal; then the §16g review found the reconciler bug |
| 2026-07-28 | precommit-manifest-6.json{,.ots} | pre-baseline final; then the 30 July date amendment (§17) drifted the three docs |
| 2026-07-30 | precommit-manifest-7.json{,.ots} | the amended tree; then drifted by this very table row — see the note below |
| 2026-07-30 | precommit-manifest-8.json{,.ots} | the amended tree, text-final; then the §17e start-time shift changed the runbook |
| 2026-07-30 | precommit-manifest-9.json{,.ots} | the tree the Friday baseline would have run on; then the §18 move to Saturday 1 August drifted the three docs |
| 2026-07-31 | precommit-manifest-10.json{,.ots} | the Saturday 1 August tree, commit 8ed83c80, instrument_tree_clean: true, Worker recorded in full for the first time since anchor 6 (§18f/2). Cut, stamped and pushed — then superseded within the hour: dry-running Guard C exposed §18f/10 and the runbook had to change again |
| 2026-07-31 (to be cut) | precommit-manifest-11.json{,.ots} | final for this baseline — the same tree plus the §18f/10 fix. Cut after §18 is text-final; not stamped as this table is written |
Why the first six, and the lesson. Anchors 2–5 each drifted, every time because pre-baseline hardening kept producing genuine improvements after stamping: three docs (recording that the anchor happened, the runbook precondition, the pipeline work list), then noise-floor.py's bare FileNotFoundError path, then the rehearsal (§16f) which measured real throughput, then the §16g review — which found a bug that would have destroyed the baseline. Every drift was caught by the runbook's own re-derive-and-diff step, which earned its place four times in one day.
The procedural lesson, recorded for run 2 and any future index: anchor last, as a distinct final act, after rehearsal AND after adversarial review — not as soon as the analysis plan feels done. An anchor taken while work is still improving documents an intention rather than the instrument. At no point was the right move to suppress a later improvement to protect a tidy anchor count; it was to re-anchor and say why. A study whose precommitment is honest about being re-taken six times is more trustworthy, not less, than one that quietly stamped once and kept editing.
Why anchor 7 was superseded within the hour, and why it is still listed. Anchor 7 was stamped over the amended tree, and then the row recording anchor 7 was added to the table above — inside STUDY_METHOD.md, which anchor 7 covers. The anchor was therefore drifted by the act of documenting it. This is self-reference, not a new finding, and it is the sharpest possible restatement of the §16b lesson: the bundle contains the document that describes the bundle, so the text must be final before the hash is taken. Anchor 8 was produced by writing every word of this section first and stamping afterwards, as a distinct final act.
Anchor 7 is left in place, stamped and listed, rather than deleted and re-cut under the same name. Its hash was already submitted to three public calendars; replacing the file while keeping the number would be the one thing this table promises never happens — superseding an anchor means "a later one exists", never "an earlier one changed".
Anchor 8 was in turn superseded the same afternoon by the §17e start-time shift, which changed STUDY_RUNBOOK.md. That is an ordinary drift, not a second self-reference: the schedule genuinely changed after anchor 8 was taken, and the honest response to a changed bundle is a new anchor, never a quiet re-stamp.
Anchor 9 was verified by re-deriving every hash immediately after stamping: zero drift against the tree that would have swept on the Friday, stamped to three calendars. Verify with ots verify tools/study/precommit-manifest-9.json.ots (Bitcoin confirmation lands a few hours after stamping; ots upgrade refreshes the proofs once they are in a block). The anchor-6 → anchor-9 diff is mechanically checkable and is confined to the three schedule docs: the rule pack 2026.07.9.json, cohort-final.csv, aggregate.py, noise-floor.py, study-run.py, prevalence-estimator.py and all remaining bundle files are byte-identical across the amendment. That is the claim §17d makes, and it is the claim a sceptical reader should verify first.
Ten anchors is a lot, and the honest summary is short: six came from pre-baseline hardening that kept improving (§16b), one from documenting an anchor inside the bundle it covers, and three from schedule changes — two made on 30 July and one on 31 July (§18). Every one predates any baseline data, the nine already stamped are committed with their proofs, and the analysis files are byte-identical across all nine; anchor 10 is expected to differ from anchor 9 in the three schedule docs and in nothing else, which is a mechanically checkable claim and the one to verify first. The count is a record of the tree changing, not of the instrument changing — though anchors 7–9 recorded the instrument wrongly, which is its own defect and is written up in §18f/2 rather than quietly repaired.
16f. Dress rehearsal of the Thursday toolchain (2026-07-28)#
study-run.py, noise-floor.py and aggregate.py's funnel path had never executed — the gate sweeps used the older agreement-run.py — so they were unexercised code on a time-critical morning. All three were run end to end against 51 real cohort sites tagged cohort smoke. Full log: tools/study/REHEARSAL-2026-07-28.md (kept outside the precommitment bundle: it is operational, and changes no analysis choice).
The three findings that reach the day:
- Sweeps take ~90 min, not the config's ~48. Saturated throughput measured at ~16 scans/min (~37 s/scan, versus the 25 s single-site calibration) → ~71 min of bulk, plus a straggler tail. Both sweeps still finish inside the pre-registered band; the runbook table now carries the measured figures. ~4% of rows end
error. - Two guards fired for real, in production, on real data: the duplicate-target check (
example.compresent twice insmoke) refused to aggregate, and the hour-band check excluded all 50 completed scans because the rehearsal ran at ~21:00 CEST. Both are adversarial-review fixes from earlier the same day (§15/review findings 10, 18, 19) behaving exactly as intended. noise-floor.pyreproduces §13a on the realagreement/agreement2pair (widget-presence discordance 3.6%, 4 lost + 1 gained of 137). Its readable-surface figure reads 9.6% against §13a's 8.0% — identical numerator (11), denominator moved to §15a's pre-registeredconfirmed_eitherbasis. Intended change, not regression.
Point 2 is also the empirical answer to whether the baseline could simply be swept early: an out-of-band sweep spends ~2,300 real third-party scans and yields zero usable rows.
16g. Second adversarial review (2026-07-28, post-rehearsal) — one study-destroying bug#
The rehearsal's own conclusions were put through a 16-agent adversarial review. It returned 10 confirmed findings, 2 refuted, and the first two matter enough to record in full.
1. My throughput measurement was wrong, and the error was self-inflicted. §16f's first draft claimed a "saturated ~16 scans/min". That was a sampled 52-second window, and with max_batch_size: 1 / max_concurrency: 10 completions arrive in waves as ten slots refill together, so any short window reads roughly double the truth. Measured properly — first finish to last finish, from D1 — the three real sweeps sustain 10.58, 10.17 and 3.37 graded/min; the densest 52 s window in those same runs reads 23.1, 19.6 and 18.5/min. A 1,142-site sweep is therefore ~110 min of bulk, not ~71. Budget ~2 h. This is the tooling-trap pattern the log already knows: when a number surprises, suspect the instrument first — here the instrument was my own sampling window.
2. reconcileStuckScans would have destroyed the baseline, and silently. STUCK_GRACE_HOURS is 1 h, sized in its own comment for "the ~30-min worst-case retry lifecycle" of a single scan. The study inserts all 1,142 rows at T0 and drains them over ~2 h, so from T0+1h the */10 cron would have terminalized every row still queued — ~500 healthy scans — as error, 200 per tick. The failure would not have been caught by any existing guard: error is terminal, and aggregate.py's incomplete-sweep check tests for NON-terminal rows, so the sweep would have looked complete and the headline would have been computed over a denominator missing ~45% of the cohort. No prior run could have exposed it — the 144-site gate sweeps drained in ~13 min and the rehearsal in ~15, all far inside the grace.
Fixed by giving cohort-tagged rows STUDY_GRACE_HOURS = 6 (~3× the measured sweep, still same-day cleanup for genuinely stranded rows), plus re-asserting the SELECT's predicate in the reconciler's UPDATE so a scan that grades in the select→update gap keeps its real result. Three regression tests added, and verified to FAIL against the old 1 h grace. Deployed pre-baseline as Worker e9a8965e-282b-4735-be4b-ae0bbc0fff26 — legitimate under §6 (the freeze starts at the first run1 enqueue) and on the §10c precedent that the last pre-baseline moment is the only honest time to fix the instrument.
Also fixed from the same review:
- The runbook claimed a throttled scan's +4 h retry lands outside the hour band. False for
run1: it runs 08:00–10:15 UTC, so its throttles retry 12:00–14:15 UTC — insideBAND_UTC, and overlappingrun1-noise's window. Onlyrun1-noisethrottles fall outside. Corrected, with the in-band case now planned for rather than assumed away. - The documented
--resubmitrecovery path would have guaranteed an aggregation refusal:/api/_study_runonly INSERTs, so resubmitting leaves the old row and creates a duplicate target, which bothaggregate.pyandnoise-floor.pyhard-refuse. The runbook now retires the superseded row tocohort='run1-superseded'first (re-tagged, not deleted, so the attempt stays auditable) and gives the exact command. noise-floor.pyfiltered band-excluded pairs insideload(), silently shrinkingn_paired. §5 requires every exclusion to be a named published class, so the count is now emitted asexcluded_out_of_band_pairs.- Stale cross-references naming an older manifest as "final" — which would have made Thursday's own drift check raise a false alarm.
16c. run1 — baseline (Sat 1 Aug, date amended twice — §17, §18) — EXECUTED, 4 h 48 m late#
The sweep ran and it is usable. It did not run at the hour it was scheduled for: both scheduled tasks fired on time and then stalled 4 h 47 m awaiting tool-permission approval, so the enqueue set for 09:00 UTC began at 13:48:13 UTC. The full timeline, what that cost and what it destroyed are §19; this section is the run record.
| Scheduled vs actual enqueue start | 09:00 UTC (11:00 CEST) scheduled → 13:48:13 UTC (15:48:13 CEST) actual — 4 h 48 m late (§19a) |
| Enqueue start / end | 13:48:13 – 14:11:52 UTC = 15:48:13 – 16:11:52 CEST |
| Grading window (first / last terminal row) | 13:48:44 – 15:00:33 UTC = 15:48:44 – 17:00:33 CEST. The last graded row finished 14:56:45 UTC — 3 min 15 s inside the 15:00 UTC cliff; the 15:00:33 row is the single error row (§19a) |
| Worker version at sweep | e9a8965e-282b-4735-be4b-ae0bbc0fff26 @ 100%, rule pack 2026.07.9 — unchanged; the §6 freeze binds from this sweep's first enqueue (§19d) |
| Tree at sweep | repo HEAD 674fae4 (the anchor-11 commit); the 17 bundle files are byte-identical to precommit-manifest-11's recorded tree (§19e). cohort-final.csv sha256 c81ae303…f507ed |
| Submitted / enqueued | 1,142 submitted → 1,098 enqueued |
| Robots-excluded / invalid | 44 excluded at enqueue: 40 robots_unreachable + 4 robots_disallowed. Invalid / unresolvable 0 |
| Terminal states (graded / blocked / error / error_capacity) | 1,073 / 2 / 23 / 0. All 1,098 rows terminal, zero non-terminal, zero throttled |
| Out-of-band captures (excluded) | 0 — and genuinely 0, not §18f/5's fail-open. Reconciled below |
| Widget confirmed / panel opened / surface readable | 791 / 251 / 172 |
| Cap usage after this sweep | 343 prior + 1,098 = 1,441 of 6,000 |
Funnel, in aggregate.py's own labels — nothing dropped:
| Stage | Sites |
|---|---|
| Submitted to the sweep | 1,142 |
| Excluded: robots.txt disallows scanning | 4 |
| Excluded: robots.txt unreachable (treated as no) | 40 |
| Excluded: invalid / unresolvable | 0 |
| Scanned to completion | 1,073 |
| Capture blocked (bot protection / redirect) | 2 |
| Scan did not complete (our side; excluded) | 23 |
| Captured outside pre-registered hour band (excluded) | 0 |
| Widget confirmed at sweep time | 791 |
| First-interaction surface readable | 172 |
Band reconciliation — why that 0 is real. §18f/5 records that in_band() fails open: it returns True whenever finished_at is falsy, so a bare 0 on the out-of-band line is as consistent with a filter that never ran as with a sweep that behaved, and §18f/5 says to read it as a red flag until discharged. It is discharged here on the row-level timestamps rather than asserted:
- Every one of the 1,073 graded rows finished inside
BAND_UTC = (7, 15)— 179 in UTC hour 13 and 894 in UTC hour 14, and none in any other hour. 179 + 894 = 1,073, so the accounting closes and there is no unexamined remainder. - Exactly one row of the 1,098 finished outside the band, at 15:00:33 UTC, 33 seconds past the cliff. It is an
errorrow, already excluded one line above as "Scan did not complete", so it cannot reach the out-of-band line at all; its absence there is arithmetic, not filtering. - Zero rows have a NULL
finished_at. The fail-open branch had nothing to fire on. That is the independent confirmation §18f/5 asks for, and it is what distinguishes a genuine 0 from a filter that did not run.
§17e predicted this line would be non-zero, and the reason it is not matters. Under the 05:00 EDT start a throttled scan's +4 h retry lands 13:00–15:15 UTC and straddles the 15:00 UTC cliff, so §17e and §16g both instructed that this line be reconciled against the recorded throttled ids rather than assumed to be ~0. It reads 0 because no row throttled at all — the throttled count for run1 is zero, so the mechanism that would have produced out-of-band rows never engaged. The prediction was sound; its precondition did not occur. That is a different thing from the prediction being wrong, and the distinction is why the reconciliation is written out rather than summarised.
Coverage — §15 headline 3, computed over sweep-confirmed rows:
| Count | Share of confirmed (95% CI) | |
|---|---|---|
| Widget confirmed at sweep time | 791 | denominator |
| Chat panel opened | 251 | 32% (29–35) |
| First-interaction surface readable | 172 | 22% (19–25) |
On 78.3% (95% CI 75.2–81.0) of the 791 sweep-confirmed widget sites, an automated visitor could not read the first-interaction surface. That is the headline coverage figure. Its denominator is detector-confirmed-at-sweep and is never to be described as "sites running a chat widget" — §15 headline 3's wording, which §12 and §14 make non-negotiable.
| Why the surface was unreadable (n = 619) | Count | Share of the 791 confirmed |
|---|---|---|
| Panel did not open after the click | 362 | 46% |
| Panel opened but text unreadable | 71 | 9% |
| Launcher covered by another overlay | 57 | 7% |
| Vendor SDK fingerprinted, no launcher | 41 | 5% |
| Launcher found but not clickable | 32 | 4% |
| Panel selector matched only the closed launcher | 31 | 4% |
| Launcher covered by a consent/cookie layer | 13 | 2% |
| Other | 12 | 2% |
The 362-row bucket carries the contamination §15b pre-registered: admission false positives land there, and blind re-judging of those already-captured screenshots by the §12/§14 method — then republishing coverage on an audit-corrected basis with error bars — is the pre-registered follow-up. Until it lands, the table above is the reported form.
CORRECTED 2026-08-02 (§21c) — this paragraph's last clause is withdrawn. It read: "and the 'vendor SDK fingerprinted, no launcher' row (41) is the only cause that names the false-positive class outright." That reading is backwards, and the code that emits those rows says so.
interaction.js:664-670emits them under the comment "the widget is there, the recipe is stale. A real maintenance-backlog signal, not a 'no widget' case." The branch is reached only whenmatchedis truthy — a named vendor — after both the bespoke recipe selector and the generic salvage failed, and every one of the 41 rows carries the literal substring(recipe stale)in its error string.aggregate.py:126-127then relabels that into "Vendor SDK fingerprinted, no launcher on the page", an assertion about the site, and this log leaned on the relabelled version inside the §5-gate discharge narrative. So the log and the code read the same 41 rows in opposite directions, and the log was wrong. Corrected reading: those 41 rows are the clearest instrument-side bucket in the table, not the clearest false-positive bucket. The table's row label is left as published because it is generated byaggregate.py, which is a precommitment-bundle file and frozen until run 2 (§21f); it is relabelled at the same time the funnel is regenerated.Instrument-side subtotal: 120 rows, 15.2% of the 791 — not 151. The two unambiguously instrument-side buckets are "panel opened but text unreadable" (79, per §21f's corrected count, not the 71 printed above) and these 41 recipe-stale rows: 79 + 41 = 120 = 15.2 points. The 31 "panel selector matched only the closed launcher" rows are mixed, not instrument-side — the selector mis-resolution is ours, but the failure to grow ≥1.5× after the click is a fact about the site — so 151 (19.1%) is an over-claim in the direction of blaming the instrument, and 120 is the number to print. Every remaining bucket is site-side or unattributed. There is still no bucket in this table that names the admission false-positive class outright; the 362-row bucket is where that contamination sits, and only §15b's blind re-judging pass can size it.
Vendor distribution — §15 headline 2, over sweep-confirmed rows. The denominator is confirmed sites, not admits: that is what aggregate.py has always computed and what its rendered header ("Share of confirmed") has always said, and §15/§18d/1 record the doc being corrected to the code rather than the reverse.
| Vendor | Sites | Share of confirmed |
|---|---|---|
| unknown / custom | 567 | 72% |
| zendesk | 63 | 8% |
| livechat | 31 | 4% |
| 13 further vendors, each n < 30 | 130 | 16% |
Only zendesk and livechat clear the pre-registered min-cell of 30; everything else is pooled into the tail row rather than published as a thin per-vendor share.
Disclosure state — §15 headline 4, readable subset only (n = 172), selection bias stated and published below the fold per §10d:
| Count | Share of readable (95% CI) | |
|---|---|---|
| Disclosure detected | 18 | 10% (7–16) |
| No disclosure detected | 64 | 37% (30–45) |
| Could not verify — wording inconclusive | 90 | 52% (45–60) |
Raw grade distribution over the 791 confirmed: UNVERIFIED 619 · WARN 154 · PASS 18.
Country cut (confirmed widgets; share whose surface was readable):
| TLD | Confirmed | Surface readable (95% CI) |
|---|---|---|
| .de | 202 | 12% (8–17) |
| .fr | 99 | 20% (13–29) |
| .it | 63 | 22% (14–34) |
| .pl | 55 | 36% (25–50) |
| .nl | 46 | 9% (3–20) |
| .se | 33 | 27% (15–44) |
| .cz | 31 | 29% (16–47) |
| all others (each n < 30) | 262 | 27% (22–33) |
§15 headline 1 is not in this section and cannot be. Corrected widget prevalence — ~11%, 95% CI 8.0–15.4 over the 12,285-site frame against the raw detector's 9.4% — is computed from the 27–28 July frame, audit and recall-expansion data and reads no sweep data at all (§15e). No sweep can move it, and §18g/2's two-dated copy rule applies: the frame and its bias correction were measured 27–28 July 2026, the surfaces on the sweep dates.
Reading these levels — and the stall's effect on them has a direction, which this section used to get backwards. They are Saturday levels and §18d wrote that cost down before the data existed. The stall makes §18d/1 worse, in a named direction, by an amount this study cannot measure.
Grading ran 15:48–17:00 CEST — the last ~72 minutes of the pre-registered band, late on a Saturday afternoon, which is about as close to the least-staffed hour of the week as the band reaches. §14 supplies the mechanism rather than a suspicion: roughly a quarter of fingerprint_strong admits ship the vendor SDK with the widget disabled, out of staffed hours, or login-gated, and "out of staffed hours" is exactly what a late Saturday afternoon maximises. A less-staffed hour means fewer widgets live at sweep time and fewer launchers that open a panel when clicked — and those are precisely the events the coverage headline counts. So headline 3's 78.3% is plausibly INFLATED by the stall: the panel opened on 251 of the 791 confirmed sites, and some unknown share of the 540 that did not is an artefact of the hour rather than a property of the site. The same direction propagates to the readable subset (172) and, through the confirmed denominator, to headline 2's vendor shares — a vendor whose customers staff chat into the early evening keeps share that a vendor with a 17:00 rota loses, for reasons that have nothing to do with its install base.
The magnitude is unmeasured, and there is no way left to measure it. One same-day repeat capture of the same sites does exist and it cannot do this job: §11 and §12 are two passes over the same 160 audit sites on 27 July, but dismissConsentBanner was rewritten to the production version between them (stratum (a) launcher reproduction moved 17/60 → 49/60), so the pair confounds an instrument change with the hour and cannot size either. An earlier draft of this paragraph asserted the flat universal — "nothing in this repo has ever captured the same sites at two different hours of the same day" — which was false and is withdrawn here; §18d/4 records the parallel correction for weekday-versus-weekend. Beyond that pair there is no within-cohort hour contrast to size the shift against, and one cannot be manufactured after the fact: any sweep from here is post-application and on a different day. No analysis rule moves and no number in this section changes. What changes is that every level above now carries an hour-of-day bias stacked on top of §18d/1's day-of-week bias, direction known, size unknown, and anyone quoting these figures is entitled to both facts. The previous version of this paragraph closed "No analysis rule moves; the disclosure gets sharper." That is withdrawn. It resolved a bias in the measurement as an improvement in the write-up, which is the wrong sign on the only line in this log that addresses the confound at all.
CORRECTED 2026-08-02 (§21b) — the clause "§18d/4 records the identical hole for weekday-versus-weekend" is withdrawn. The paragraph above is left as written, per §21's rule that executed-run records are not rewritten. The weekday-versus-weekend half of that clause is false, and it became false on the very day this section describes.
sweep-agreement2.json(cohortagreement2, enqueued Tue 2026-07-28 03:41:57Z) andsweep-run1.json(Sat 2026-08-01) share 143 targets, 139 graded in both — the same sites, a weekday capture and a weekend capture, both shipping in the snapshot. The hour-of-day half needs qualifying too: the two captures are ~10 hours apart on the clock (05:41 CEST vs 15:48–17:00 CEST), so the pair is early-vs-late as much as weekday-vs-weekend, which is one reason it cannot be read as a clean day-of-week contrast.This section's conclusion is unchanged and no bound is published. The pair cannot size the late-Saturday-afternoon shift, for four independent reasons set out in §21b: both endpoints are unstaffed windows,
agreement2is out of the pre-registered band, the sample is the deliberately skewed agreement subsample, and the three measures move in three different directions. What is withdrawn is only the reason — "no such data exists" is a stronger and more checkable claim than the situation warranted, and it is false.
16d. run1-noise — same-day retest (Sat 1 Aug, date amended twice — §17, §18) — DID NOT RUN; the §5 noise floor was not obtained, and is now permanently unobtainable#
There are no run1-noise rows in D1 — not a partial sweep, not an aborted mid-drain sweep, none. This section carries no table because there is nothing to tabulate, and that is a recorded result rather than an unfinished section or an oversight in the log: the sweep was never enqueued. What follows is why, and what its absence costs.
What happened. study-run1noise-enqueue-sat1aug fired on time at 11:30:15 UTC and then stalled awaiting tool-permission approval exactly as the run1 task had (§19a). When it resumed at 13:51 UTC it aborted at Guard A: the guard requires the current UTC hour to be ≥ 7 and < 13, and 13:51 UTC is past that cutoff. Its own report additionally recorded that its first date -u reading — 11:30 UTC, taken before the stall — no longer described the world, i.e. it re-read the clock rather than trusting a value it already held. The guard did exactly what it was written to do (§19c): a sweep enqueued at 13:51 UTC could not drain before the 15:00 UTC cliff, and aggregate.py would have discarded most of what it produced.
What the absence costs:
- There is no measured same-code test-retest noise floor, and there never will be one for this baseline. §5 pre-registers it as a same-day full-cohort double sweep; §16c is the only sweep that exists.
noise-floor.pyhas no second sweep to pair and did not run. The stronger form of this is stated below rather than left as an implication: the object §5 registered is not delayed, it is gone. - This section's previous claim is retracted. It read: "This supersedes §13a's cross-code upper bound and is the noise floor the run-1 → run-2 delta is judged against." Nothing supersedes §13a. Its bound — ~4 points on widget presence, ~8 points on readable surfaces, measured on 137 sites, confounded with an instrument change, gathered on Monday 27 July — is now the study's only noise estimate, which is precisely the deficiency §5's floor was written to remove. §13a's own closing line still stands unanswered: "A clean same-code double sweep is still owed before any delta is published." It is owed and, in the form §5 registered it, it can no longer be paid. §19f states what may and may not be claimed on that footing — as a standing rule, not as an interim one.
- §18f/8 is moot for this run. Its finding — that
noise-floor.pystamps a "same-day" label the code never checks — has nothing to attach to, because there is no output to label. - The cap was not spent. Zero rows, so usage stands at 1,441 of 6,000 (§16c) rather than §18f/7's projected 3,769. The arithmetic is recorded because §18f/7 warned that a contingency would not fit under the wall, and on these numbers one would: 1,441 + 1,142 for run 2 = 2,583, leaving 3,417 against a full double sweep's 2,284. That is a statement about headroom, not a decision.
A same-day, pre-application noise floor for this baseline is now PERMANENTLY UNOBTAINABLE. Not deferred, not pending, not "to be recovered" — gone. §5 registered a same-day full-cohort double sweep, and the day was Saturday 1 August 2026, the last day before Article 50 applied. That day is over and there is no procedure that returns it. Anything collected from here is a different object, in at least three ways at once: it is measured after the Regulation applies, so it is not a pre-law floor; it is measured on a different day, so it is not the same-day floor §5 registered and its variance is not the variance the baseline actually carries; and if it is taken on the run-2 day it is measured in a different week, on a cohort that has had a week to move, which folds the between-Saturday component §18d/3 warned about into the very number meant to bound it. It could still be worth having. It cannot be the thing that was lost, and no later section of this log may describe it as if it were.
WHETHER AND HOW a substitute is obtained, and what that substitute would and would not be, is OPEN and is deliberately not answered here. The decision was deferred until run1 was written up. This log does not pre-empt it, and no sentence in §16c, §16e or §19 should be read as assuming that a substitute will be collected at all, that any particular one is appropriate, or that obtaining one would restore what 1 August destroyed. §19f states what may and may not be claimed either way.
16e. run2 — EXECUTED Monday 10 August 2026, on time, in band (rescheduled twice — §18, §25; the weekday deviation is §25a's, stated as a deviation)#
FILLED 2026-08-10. The paragraph below ("Date unchanged and still correct: Saturday 8 August") was written before the 8 August miss and is kept per this log's no-rewrite rule; §25 records the reschedule and the weekday deviation. Full run record:
tools/study/RUN2-2026-08-10.md.
Enqueue 11:01–11:02 UTC Mon 10 Aug — task fired 11:00:07 against fireAt11:00:00, the first on-time fire in five dates; no permission stall (§19b's stored approvals held)Submitted → enqueued 1,142 → 1,088 Robots exclusions (stated, not dropped) 54 = 50 robots_unreachable+ 4robots_disallowedTerminal states at drain 1,064 graded / 22 error (2.0%) / 1 blocked; 1 throttlednon-terminal (below)Band all 1,064 graded rows inside BAND_UTC = (7, 15)— 947 in UTC hour 11, 117 in hour 12; last graded row 12:11:42 UTC, ~2 h 48 m clear of the 15:00 cliff (run1's margin was 3 m 15 s)Duplicate targets 0 (checked within a minute of the enqueue returning, per the runbook) Drain ~73 minutes (11:02 → 12:15:18 UTC), ~14.9 graded/min Instrument Worker 0c23d901-9152-45a9-b545-34d3b1e39be0@ 100% (the 6 Aug build — §24; run1'se9a8965ewas superseded 2 Aug, §22), pack2026.07.9. Guard D: all 14 non-doc bundle files byte-identical to manifest-11; drift confined to the two schedule docs, expectedEgress country (§4; §21g/a) Probed from inside the sweep path at 11:01:08 UTC: US, colo MIA, IP 104.28.153.73, both observations agreeing. The baseline has no such record and never will (§21d/3). Prior readings: US/MIA8 Aug, US/IAD3 Aug — country stable, colo variesCap 1,441 prior + 1,088 = 2,529 of 6,000 Weekday and hour, as executed. Monday, 11:00 UTC start — a §5 weekday deviation, stated as one in §25a, not compliance. The hour band itself is the pre-registered one. Item 1's "SETTLED — 09:00 UTC Saturday" answer below (§20c) was overtaken by the 8 August miss; §25 records the reschedule and its reason, and the executed slot is the one recorded here.
Attrition against the identical 1,142-target cohort file (recorded as fact; the run1 rows were retired to
run1-supersededon 2 Aug and are not a publishable comparator — §25a's two confounds travel with any use of this): left-frame 0, newly-off-frame 0 (same locked file, byte-identical, Guard D); newly-robots-disallowed 0 — the same 4 targets in both runs; newly-robots-unreachable 16, recovered 6, stably unreachable 34 — net exclusions 44 → 54, enqueued 1,098 → 1,088. Target-level lists are derivable fromtools/study/run{1,2}-enqueue.json.The §15a per-outcome paired deltas were not computed and will not be published. There is no baseline: all 1,098 run1 rows were retired on 2 August (§20e, §25a). §19f binds any future citation of the retired rows, and with §20 marked not-obtained (§26) its comparator rule revives unchanged. Run 2 is published as what it now is — a single post-application levels snapshot, Monday 10 August, staffed-hours window.
The one throttled row — recorded, NOT re-enqueued (runbook rule: never re-enqueue while the queue message is still retrying). Scan
[scan id withheld](one.becohort site), created 11:04:56,finished_atNULL at drain. Its +4 h retry lands ~16:05 UTC, outside the band, so the funnel's "captured outside pre-registered hour band (excluded)" line must read 1 for run2 once it terminalizes — the opposite expectation from run1, where 0 was genuine. A 0 on that line for run2 meansin_band()failed open on the NULL timestamp (§18f/5). Final state recorded below when terminal; the NULL-finished_atgate is re-run immediately before aggregation either way.TERMINAL 2026-08-10 16:06:41 UTC — graded, out of band, exactly as predicted. The +4 h retry completed on schedule and the row graded in UTC hour 16, past the 15:00 cliff. Both pre-aggregation gates were re-run against live D1 immediately after: graded rows with NULL
finished_at0, out-of-band graded rows 1 (this row and only this row), non-terminal rows 0. The cohort's final census is therefore 1,088 = 1,065 graded (1,064 in band + 1 out-of-band, excluded) / 22 error / 1 blocked, and the funnel's out-of-band line reads 1, as it must.
Date unchanged and still correct: Saturday 8 August 2026 — same weekday/hour band as run1 per §5, after Article 50 applies, inside the §6 freeze that went live at run1's first enqueue (§19d).
Same table as 16c, plus the §15a per-outcome paired deltas and every named attrition class (left-frame, newly-unreachable, newly-robots-disallowed, newly-off-frame).
Two things this section must settle on the day, both created by §19a's stall, neither decided here:
- Which hour run 2 starts. §5 requires run 2 in the same weekday/hour band as run 1, and its stated reason is that availability-gated launchers otherwise bake the hour choice into the denominator. Run 1 was scheduled for 09:00 UTC and actually graded in UTC hours 13–14, so the registered slot and the executed slot are ~5 hours apart and run 2 can match one or the other but not both. Starting at 09:00 UTC honours the registered plan and confounds the delta with an hour-of-day effect; starting at ~13:45 UTC matches run 1's realised hours and leaves almost no margin at all before the 15:00 UTC cliff (§19a — run 1's last graded row landed at 14:56:45 UTC, 3 min 15 s clear of it). Note also that matching run 1's realised hours means matching a measurement window §16c records as biasing the levels in a known direction. Whichever is chosen must be recorded here with its reason, before run 2 runs. SETTLED 2026-08-02, before run 2 — see §20c: the registered slot (09:00 UTC), because matching run 1's realised hours means a ~3-minute margin against the cliff and deliberately reproducing a window §16c records as biased. The delta is therefore confounded with hour of day, in the direction of spuriously showing improvement (§20c).
- What the delta is judged against, given §16d. The constraint is §19f, and it binds whether or not a floor is recovered first.
- ADDED 2026-08-02 — the sweep egress country must be probed and recorded before run 2 fires. §4's vantage paragraph pre-registers that the Cloudflare Browser Rendering egress country be "probed and recorded before each run". It was not done for the baseline, or at all — the third unmet pre-registration requirement (§21d/3). It is recoverable for run 2 and only for run 2: probe from inside the sweep path, record the result in this section beside the Worker version and the drift check, and note in the same line that the baseline has no such record and never will. Full checklist, including the test-fire item §19b requires and the two items §21f defers, is §21g.
17. AMENDMENT (2026-07-30, the founder): baseline slips one day to Friday 31 July#
Pre-registration amendment. Date only. No analysis rule, no instrument, no cohort, and no threshold changed. Recorded here before the baseline exists, which is the only kind of schedule amendment that is worth anything.
17a. What happened#
The 30 July baseline did not run. No sweep was enqueued at any point during the pre-registered 09:00–17:00 CET band. This is stated as observed rather than explained: the verifiable facts are that D1 contained no run1 or run1-noise rows, and the last commit on the tree was Wednesday's db53c7b ("sweep GO"). The gap was discovered at 16:02 CEST (14:02 UTC) — 58 minutes of band remaining against a measured ~110-minute sweep, needed twice.
Every §6 freeze fact still held at that moment, re-verified rather than assumed:
| Check (re-run 2026-07-30 ~16:05 CEST) | Result |
|---|---|
| Test suite | 510 / 510 pass, 0 fail |
| Live Worker version | e9a8965e — unchanged, no deploys since 2026-07-28 |
| Rule pack | 2026.07.9 |
| Precommit re-derive vs anchor 6 | ZERO DRIFT, 17 files, commit db53c7bd |
| D1 cohort rows | agreement 144 · agreement2 144 · smoke 54 = 342, nothing stray |
So nothing was broken. The instrument was ready and the clock was not.
17b. Why a salvage sweep was refused#
Starting run1 at ~16:10 CEST was considered and rejected for two independent reasons, either of which is sufficient:
- The band filter would have discarded most of it.
aggregate.py:in_band()excludes any scan whosefinished_atfalls outside 07:00–15:00 UTC. At the measured 10.2–10.6 graded/min, ~48 usable minutes admits roughly 500 of 1,142 rows. - The surviving half would not have been a random half.
cohort-final.csvis ordered by ascending Tranco rank, andstudy-run.pyenqueues it in 50-row slices in file order. The rows completing first are therefore the most-trafficked sites in the cohort — biased on precisely the variables under study (staffed chat hours, vendor mix, widget reachability). That is the same class of selection bias the §15 reframe was adopted to escape, and re-introducing it to save a calendar date would have been indefensible.
A third consideration made it worse rather than better: the §6 freeze starts at the first run1 enqueue, so a salvage sweep would have frozen the instrument against a baseline that could not be published.
17c. Why Friday, and what was rejected#
Friday 31 July preserves every property the design depends on:
- Still pre-Article 50. The Regulation applies Sunday 2 August 2026, so a Friday baseline remains a genuine before-measurement and the pre/post frame survives intact.
- Weekday pairing preserved. Both runs move together — baseline Fri 31 Jul, run 2 Fri 7 Aug — so the paired design still compares like with like. A one-sided slip would have confounded the delta with a weekday effect.
- Hour band unchanged: 09:00–17:00 CET, the same
BAND_UTCfilter, no re-registration of the analysis rule. - Full-cohort double sweep intact:
run1+run1-noise, identical 1,142-site list.
Rejected alternatives, with the reason each fails:
| Option | Why not |
|---|---|
| Partial sweep today | §17b — ~44% coverage, Tranco-ordered, biased denominator |
| Sweep this evening, outside the band | Amends the analysis rule rather than the date, and measures out of staffed hours — §12/§14 already established that ~27% of fingerprint_strong sites ship the SDK with the widget disabled or out of staffed hours, so an evening sweep would understate reachability by construction |
| Keep Thursdays: baseline 6 Aug, run 2 13 Aug | Baseline lands after the law applies; the delta becomes post-vs-post and the headline in STUDY_PIPELINE.md §0 stops being true |
| Weekend baseline (Sat 1 Aug) | Weekend chat staffing differs sharply from weekdays; run 2 would inherit the same distortion |
17d. Anchor 7, and why the count went to seven#
The date is named inside STUDY_METHOD.md, STUDY_PIPELINE.md and STUDY_RUNBOOK.md — all three are in the precommitment bundle — so amending it necessarily drifts the manifest. The final anchor is precommit-manifest-9.json{,.ots}, covering the same 17 files with the same instrument; the diff against anchor 6 is confined to the three docs, and src/packs/eu-ai-act-art50/2026.07.9.json plus every tooling file is byte-identical. (Anchors 7 and 8 were superseded the same day — see §16b.)
Unlike anchors 2–6 (§16b), this re-anchor is not caused by hardening still in flight. It is caused by a schedule change, and the §16b lesson — anchor last, say why you re-anchored — is what makes it cheap to do honestly. The alternative, sweeping Friday against a bundle that says Thursday, would have left the precommitment not matching what ran, which §16b's own standard calls worse than none.
17e. Start times moved inside the band — operational, not an amendment#
The runbook's sweep slots moved from 10:00/14:00 CEST to 11:00/13:30 CEST (05:00/07:30 EDT). This is deliberately not filed as a pre-registration amendment, and the distinction is worth stating precisely because it would be easy to overclaim in either direction:
- What is pre-registered is the band, 09:00–17:00 CET (§5), enforced in code as
BAND_UTC = (7, 15). Both sweeps still start and finish inside it.BAND_UTCis untouched, no analysis rule moves, and the anchor-6 → anchor-9 diff still shows every analysis file byte-identical. - What moved is an operating choice the runbook made within that band. The old slots put
run1at 04:00 EDT — and a 04:00 local start is the proximate reason the Thursday sweep never happened (§17a). Fixing the cause is not the same as changing the design.
One real consequence, recorded because it is a genuine change to expected data. A throttled scan retries at +4 h. Under the 04:00 start, run1 ran 08:00–10:15 UTC and its retries landed 12:00–14:15 UTC — wholly inside the band. Under the 05:00 start, run1 runs 09:00–11:15 UTC and its retries land 13:00–15:15 UTC, which straddles the 15:00 UTC cliff. So a late-throttling run1 row can now be excluded as out-of-band where previously it would not have been. It remains a named, counted exclusion class in aggregate.py's funnel, which is what §5 requires; the runbook now says to reconcile that line against the recorded throttled ids rather than assume it is ~0. run1-noise is unaffected in kind — its retries fell outside the band before and still do.
18. AMENDMENT (2026-07-31, the founder): baseline moves to Saturday 1 August#
Pre-registration amendment. Date only — but not, this time, a free one. No analysis rule, no instrument, no cohort, no threshold and no band changes. What does change is the weekday of both runs, and that has costs the design has never carried before. They are written down in §18d, in full, before any baseline data exists, because a distortion disclosed after the numbers land is worth nothing and reads as an excuse.
- Baseline (
run1+run1-noise): Saturday 1 August 2026. - Run 2 (
run2): Saturday 8 August 2026. Same-weekday pairing preserved. - Sweep slots unchanged: 11:00 / 13:30 CEST = 05:00 / 07:30 EDT = 09:00 / 11:30 UTC. Deliberately not re-opened — §17e's throttle-retry straddle analysis is computed for a 05:00 EDT start, and moving the slot would invalidate it for no gain.
- Band unchanged: 09:00–17:00 CET,
BAND_UTC = (7, 15), abort line 16:30 CEST. - Instrument unchanged: Worker
e9a8965e-282b-4735-be4b-ae0bbc0fff26, pack2026.07.9, the same frozen 1,142-site list. - New final anchor:
precommit-manifest-11.json{,.ots}, superseding anchors 1–10 — anchor 10 was cut, stamped and pushed on the evening of 31 July and superseded within the hour when dry-running Guard C exposed §18f/10. Per the §16b lesson it is cut as a distinct final act once every word of this section is fixed, so it is not stamped as this is written and nothing below claims it is. Resolving the "final anchor" wording (added 2026-07-31, because three sections now name three different manifests): every earlier sentence in this log that names the final anchor — §15b (precommit-manifest-6), §16a (precommit-manifest-6) and §17d (precommit-manifest-9) — is a dated snapshot of the anchor that was live when that sentence was written, and each is left standing as the record of that moment rather than retro-edited. The live anchor for this baseline isprecommit-manifest-11. Where any earlier section disagrees with this bullet, this bullet is the current fact and the earlier section is history — the same rule §16a states about itself.
Re-verification before this amendment was written (2026-07-31 evening). Recorded in §17a's format, and it is the table STUDY_RUNBOOK.md's precondition banner cites:
| Check (re-run 2026-07-31 evening) | Result |
|---|---|
| Test suite | 510 / 510 pass, 0 fail |
| Live Worker version | e9a8965e-282b-4735-be4b-ae0bbc0fff26 @ 100% — unchanged, no deploys since 2026-07-28T19:50:54Z |
| Rule pack | 2026.07.9 |
| Precommit re-derive vs anchor 9 | ZERO DRIFT across all 17 files — measured against anchor 9 and BEFORE this amendment's edits |
| D1 cohort rows | agreement 144 · agreement2 144 · smoke 55 = 343; no run1 or run1-noise rows |
.study-secret | present locally |
Read the drift row precisely, because it stops being true two paragraphs later. Zero drift held against anchor 9 at the moment of checking. This section, STUDY_PIPELINE.md and STUDY_RUNBOOK.md are all bundle files, so writing this amendment drifts anchor 9 by construction — that is not a failure, it is exactly why precommit-manifest-11 exists and why §16b insists the text be final before the hash is taken. Anyone re-deriving after these edits and comparing against anchor 9 will see drift in the three schedule docs and must compare against manifest-11 instead.
18a. The second miss, and the root cause the first fix never touched#
The 31 July baseline did not run either. Stated as observed, in the same terms as §17a: no run1 or run1-noise rows existed in D1 at any point during the pre-registered 09:00–17:00 CEST band, the band closed at 17:00 CEST with nothing enqueued, and the gap was discovered at ~18:35 CEST — ninety-odd minutes after the band had shut. There was no salvage decision to make this time. §17b's two reasons for refusing a partial sweep apply a fortiori to a sweep that cannot begin inside the band at all, and §16f point 2 already measured what an out-of-band sweep buys: ~2,300 real third-party scans and zero usable rows.
The root cause, diagnosed properly this time. §17e attributed Thursday's miss to the 04:00 EDT start and moved the slots to 05:00 / 07:30 EDT. That was a genuine improvement and it was not the cause. The cause is that no automation ever existed, verified on 31 July rather than assumed:
| Where a trigger would live | What is actually there |
|---|---|
Host crontab -l | no crontab for <user> |
~/Library/LaunchAgents | five agents, all third-party (Epic, Google ×3, Valve) — nothing for the study |
at queue (atq) | empty |
wrangler.jsonc crons | 17 3 * * *, */10 * * * *, 7 * * * * — retention cleanup, the stuck-scan reconciler, the monitor sweep. All three are product jobs; none knows the study exists |
The 30 July "fix" was an edit to an alarm time inside a markdown table, and nothing else. A schedule that lives only in a document depends on a human reading that document at the right hour on the right day. That dependency failed on 30 July, was "fixed" by changing the hour written in the document, and failed again on 31 July. Two consecutive pre-registered baseline dates were lost to one unmitigated cause, and the second loss is the evidence that the first diagnosis was incomplete. This is the least flattering fact in this log and it is recorded at full strength, because the alternative — a third miss explained by a third proximate cause — is the one outcome the study cannot absorb.
Remedy adopted with this amendment: a real scheduled trigger for each of the two enqueues on the sweep date, firing independently of anyone reading the runbook. The runbook's precondition list stays exactly as it is — a trigger that fires into a broken tree is not an improvement — but from now on the trigger, not the operator's memory, is what starts the sweep, and the operator's job is to verify preconditions and watch the drain.
What the trigger actually is, specified rather than asserted. A remedy that is not specified is not auditable, which is the whole point of this section, so it is written out here in full. Created 2026-07-31 evening: two one-shot Claude Code scheduled tasks, one per sweep, stored as SKILL.md files under ~/.claude/scheduled-tasks/. They can be listed — do that rather than assume they exist (the runbook's eve-of-sweep pass requires it):
| Task | Fires | Enqueues |
|---|---|---|
study-run1-enqueue-sat1aug | 2026-08-01T05:00:00-04:00 — 05:00 EDT = 11:00 CEST = 09:00 UTC | study-run.py --cohort run1 |
study-run1noise-enqueue-sat1aug | 2026-08-01T07:30:00-04:00 — 07:30 EDT = 13:30 CEST = 11:30 UTC | study-run.py --cohort run1-noise |
Both are one-shot and auto-disable after firing, so neither can re-fire and duplicate a cohort. Neither task enqueues blind. Each runs hard preflight guards first, and aborts without enqueueing if any of them fails, reporting which one and why:
- Guard A — the band. Reads
date -uand refuses to enqueue unless the current UTC hour is ≥ 7 and < 13. (13:00 UTC, not the band's 15:00 UTC edge, because a sweep takes ~2 hours of drain and sweep 2 must also fit before the 16:30 CEST abort line.) This guard exists because of the mechanism caveat below: a late fire must abort, never silently enqueue out of band, since out-of-band rows are discarded byaggregate.pyand would burn ~1,142 of the 6,000-row cap for nothing on the last pre-Article-50 day. - Guard B — no rows already exist. Queries D1 for
run1/run1-noiserows and aborts if any are present, because a second enqueue creates duplicate targets and both analyzers hard- refuse on duplicates (§18f/9). - Guard C — instrument unchanged.
wrangler deployments status(notdeployments list | head, which reads the oldest end — §18f/10); the deployed Worker must still bee9a8965e-282b-4735-be4b-ae0bbc0fff26at 100%, or the §6 freeze is broken. - Guard D — zero precommitment drift against
precommit-manifest-11. - Guard E —
npm testgreen (run1task only).
Then, and only then, it enqueues — and immediately runs the mandatory duplicate-target query from §18f/9 (SELECT target, COUNT(*) c … GROUP BY target HAVING c>1, which must return zero rows) and reports the result before reporting success. The run1-noise task carries two differences: it additionally aborts if no run1 rows exist at all (a noise floor without a baseline is meaningless), and it deliberately does not require run1 to be terminal — per the runbook, sweep 2 starts on time regardless, because delaying it is what actually risks the 15:00 UTC cliff.
The weakness of this mechanism, stated plainly rather than dressed up as a daemon. Claude Code scheduled tasks fire only while the app is open. If the app is closed at the scheduled minute, the task runs at next launch — which could be hours late. This is not an OS-level hardened trigger and must not be described as one. The operator requirement is therefore explicit: the app open and the Mac awake at 11:00 and 13:30 CEST. Guard A is what makes a late fire safe rather than catastrophic — a task that wakes up at 16:00 CEST aborts and says so, instead of quietly filling D1 with rows the analysis will throw away. The residual risk is a missed sweep, which is recoverable; the risk the guard removes is a silently wasted sweep, which is not.
Two further things about this remedy that must not be overstated, because overstating a control is exactly what §18a is about. First, the guards are instructions in a SKILL.md carried out by a model, not code with an exit status. They are far better than an alarm and they are not a hard interlock; the only enforcement that is truly mechanical is aggregate.py's own band filter, downstream, after the fact. Second, the trigger has never fired. It was created on the evening of 31 July and Saturday is its first execution, with no dry run against a real slot. So the honest residual probability of a third consecutive miss is not zero — it is merely much lower than a human alarm that has now failed twice. A test-fire before the sweep (which aborts harmlessly at Guard A outside the band, while still exercising and storing the tool approvals the task needs) is the cheapest way to reduce it, and is on the eve-of-sweep checklist for that reason.
Why not a Cloudflare cron, which would genuinely be more robust: adding one requires a deploy, a deploy mints a new Worker version, and that breaks the §6 instrument freeze against e9a8965e-282b-4735-be4b-ae0bbc0fff26 — grading the baseline under un-precommitted code to fix a scheduling problem. Correctly ruled out; revisit after run 2 (Sat 8 Aug), when the freeze lifts.
18b. Why Saturday 1 August, stated as a purchase#
Article 50 applies Sunday 2 August 2026. Saturday 1 August is therefore the last pre-application day that will ever exist, and a pre-law measurement of these surfaces is a one-time, permanently non-repeatable observation: after Sunday there is no way to go back and measure what EU chat surfaces looked like before the obligation attached. Every other property of the design — the frame, the cohort, the instrument, the band, the paired analysis — is reproducible on demand. This one is not.
That is precisely what is being bought, and §18d is what it costs. Both halves of that trade belong in the same section, so that no reader has to reconstruct it from a date change.
Two secondary properties survive intact and are worth stating: the pre/post frame still brackets the deadline (baseline 1 August is before, run 2 on 8 August is after), and the paired design still compares like with like, because both runs move together to the same weekday. A one-sided move — Saturday baseline against a weekday run 2 — was never on the table; it would confound the delta with exactly the weekday effect §17c was right to worry about.
18c. §17c's weekend rejection is overturned here — and it was not wrong#
§17c rejected a weekend baseline in one line: "Weekend chat staffing differs sharply from weekdays; run 2 would inherit the same distortion." That rejection is overturned by this amendment, and it must not be characterised as misreasoned, because it was not.
Read precisely, §17c is internally consistent. Its weekend row is a statement about the levels — what share of sites show a reachable widget, a readable surface, a disclosure — and it presupposes a paired move, which is exactly what is happening now ("run 2 would inherit the same distortion" only makes sense if run 2 also moves). Its weekday-pairing bullet, in the list immediately above that table, is a statement about the delta — that a one-sided slip would confound the comparison. The two passages address different quantities and they do not contradict each other: the pairing protects the delta, and the weekend row concedes the levels. Anyone re-reading §17 should find it coherent, because it is.
What changed is the alternative it was weighed against. On 30 July the choice was "weekend, distorted levels" versus "Friday 31 July — clean weekday, still pre-law". Against that, rejecting the weekend was obviously correct and cost nothing. Friday is gone (§18a). The choice on 31 July is "Saturday, distorted levels, pre-law" versus "a clean weekday, post-law" (§18e). That is a different question, and it has a different answer.
§17c is overtaken, not refuted. It answered the question in front of it correctly; the question changed underneath it. That distinction is the whole difference between a study that amends itself honestly and one that rewrites its own reasoning to match the schedule it ended up with.
18d. What a Saturday baseline costs — five items, recorded before the data exists#
None of these is a reason not to proceed; §18b is why they are being accepted. All five are recorded here so that they constrain the published claims rather than being discovered later as convenient explanations.
1. The absolute levels are distorted, and the published outputs are levels. Weekend chat staffing differs from weekday staffing. Three of the four reframed headlines (§15) are levels measured on the sweep day: coverage by failure cause, vendor distribution, and the disclosure-state distribution on the readable subset. All three are weekend-sensitive, and §14 supplies the mechanism rather than a suspicion — roughly a quarter of fingerprint_strong admits ship the vendor SDK with the widget disabled, out of staffed hours, or login-gated, and "out of staffed hours" is exactly the category a Saturday inflates. Note the reconciliation made in the same pass: headline 2's denominator is sweep-confirmed sites, not admits — that is what aggregate.py has always computed and what the rendered header ("Share of confirmed") has always said, and §15's text has been corrected to match the code. The consequence is that the vendor distribution is weekend-sensitive too: an admit is a fixed 27–28 July artefact, while a confirmation is an observation made on the sweep day, so a vendor whose customers run staffed weekday-only chat loses share on a Saturday for reasons that have nothing to do with its install base.
2. The widget-presence delta is BIASED, not merely under-powered — and the bias points at the study's own headline risk. §15a fixes that outcome's paired universe as sites in-frame and graded in BOTH runs. A site whose widget is dark at weekends is graded in both runs — as absent on 1 August and absent on 8 August — so it stays in n and can never be discordant. noise-floor.py publishes (lost + gained) / n over exactly that universe, so every such site inflates the denominator without ever being able to touch the numerator: the published widget-presence shares are attenuated toward zero.
The tension a skeptic will press here, stated before they get to. The same universe rule fixes both the noise floor and the run-1 → run-2 delta, so the attenuation hits the two together — and to the extent it is a common multiplicative factor, the delta-versus-floor comparison is partly invariant to it, which is an argument that a Saturday costs less than the paragraph above implies. Two reasons that defence is not sufficient. (i) It holds only if the inert weekend-dark mass is the same mass in both runs; between-Saturday rota variance is precisely what cost 3 says the same-day floor cannot see, so the invariance is assumed, not measured. (ii) Invariance of the ratio does not rescue the published absolute shares, which are what the page reports. So: the headline numbers are attenuated for certain, and the delta-vs-floor judgement is attenuated by an unknown and probably smaller amount. Both belong in the published limitations; neither should be used to wave the other away. The direction is the problem. It pushes the delta toward "nothing changed when Article 50 began to apply" — which is the single conclusion this study is least able to defend, because it is also what a null result from an under-powered design looks like. For the coverage outcomes (§15a's confirmed_either universe) and for disclosure-state (readable-in-both), weekend-dark sites drop out of n instead of sitting inertly inside it, so there a Saturday costs power and not direction. This distinction — bias for one outcome, power for the others — is recorded nowhere else in this log and is the sharpest of the five costs.
3. The noise floor is weekend-scoped, and is therefore systematically too tight. §5's pre-registered floor is a same-day double sweep: run1 and run1-noise, hours apart, on Saturday 1 August. It measures within-Saturday, within-band variance. The delta it is used to judge spans Saturday to Saturday, a week apart, and therefore carries a component the floor cannot see by construction: between-Saturday variance — a different weekend rota, a holiday weekend, a vendor's weekend defaults changed, an August staffing dip. A Sat→Sat movement can clear a floor built from within-day variance while still sitting inside ordinary week-to-week variability. Two honest qualifications: this structure is not unique to weekends (a Thu→Thu delta judged against a same-Thursday floor has the identical gap), and the Saturday case is worse only to the extent that weekend staffing is the more variable of the two — which is itself unmeasured (cost 4). It is stated here because it is a stronger objection than the one §17c actually raised, and because it appeared in no document until now.
4. The magnitude of the weekend effect is UNMEASURED, and the ~27% figure cannot size it. The temptation is to reach for §14's ~27% as a weekend penalty. It is not one. That number is the complement of fingerprint_strong precision — 73.1%, 95% CI 59.7–83.2 — measured on 60 admits captured on a Monday, and it is attributed to a union of three undecomposed causes (widget disabled / out of staffed hours / login-gated) that the audit never separated. It contains no weekday-versus-weekend contrast and cannot be turned into one after the fact. Nothing in this repo measures a weekday or weekend effect at all: no discovery pass, no audit and no sweep has ever captured the same sites on both. So both sides of the argument are unquantified priors — §17c's "differs sharply" and the counter-argument that a paired design washes it out. Saying so is the honest position; putting a number on it, in either direction's favour, would not be.
CORRECTED 2026-08-02 (§21b). The middle sentence — "Nothing in this repo measures a weekday or weekend effect at all: no discovery pass, no audit and no sweep has ever captured the same sites on both" — is false as of 1 August 2026. It was true when written on 31 July, because the only weekend sweep in the repo did not exist yet; it was never amended once run 1 landed.
sweep-agreement2.json(Tue 28 Jul) andsweep-run1.json(Sat 1 Aug) share 143 targets, 139 graded in both — a weekday capture and a weekend capture of the same sites, both shipping in the snapshot, falsifiable in one command. The conclusion of this item survives unchanged: the pair still cannot bound the weekend effect (§21b gives four reasons — both endpoints are unstaffed,agreement2is out of band, the sample is the deliberately skewed agreement subsample, and the three measures move in three different directions), and the ~27% figure still cannot be repurposed as a weekend penalty. What is withdrawn is the claim that no such data exists, which is a stronger and more checkable statement than the one this item needed, and the wrong one to leave standing on a page whose subject is not overclaiming.
5. §5's phrase "EU business hours" is genuinely strained by a Saturday. What is pre-registered is a clock band — 09:00–17:00 CET, enforced as BAND_UTC = (7, 15) — and both sweeps run inside it exactly as registered; no code and no analysis rule moves. But the phrase carries a weekday connotation that a Saturday does not satisfy, and a competent reader is entitled to say so. The band is honoured to the letter and the label is imprecise. That is disclosed here and must be disclosed in the published methodology; it is not argued away by redefining the phrase.
And the stronger form of the objection, which the paragraph above understates. §5 does not merely name the band, it gives its reason: the runs go inside EU business hours because "availability-gated launchers (tawk isChatHidden, staffed-hours widgets) otherwise bake the hour choice into the denominator." A Saturday does not just strain the label — it is the condition under which that stated hazard is largest, since staffed-hours gating is exactly what a weekend maximises. So the honest statement is: the letter of §5 is honoured and its purpose is partly defeated. What preserves the design is that §5's other clause — run 2 in the same weekday/hour band as run 1 — is satisfied literally by Sat↔Sat, so the hour choice is baked into both denominators equally rather than into one. That is a real defence for the delta; it is not a defence for the levels, which is cost 1. Anyone quoting this study's prevalence figures is entitled to know they were measured on a Saturday.
18e. Rejected alternatives#
Real options, each with a real case for it. Stated fairly, in §17c's format.
| Option | Why not |
|---|---|
| Salvage sweep on the evening of 31 July | Outside the band by construction — not an amendment to the date but to the analysis rule, and §16f point 2 measured the yield: ~2,300 real third-party scans for zero usable rows. Refused on the same grounds as §17b |
| Mon 3 Aug + Mon 10 Aug | The analytically safer option, and the one that would be chosen if the study's question were different. Clean weekday levels, a paired weekday delta, and two full days to build and test the automation §18a says never existed — it removes costs 1, 2 and 3 outright. Rejected because the baseline would then be measured after the Regulation applies: the pre/post frame collapses to post-vs-post, and the pre-law surface observation — available on exactly one remaining day — is forfeited permanently. (§17c also rejected a post-law baseline on the ground that §0's headline "stops being true". That reason is not carried forward here: per §18g/1 the "on the day Article 50 took effect" wording was never satisfiable by any candidate date, 31 July included, so it favours no option and must not be used to tilt this one.) Analytical cleanliness is recoverable; the pre-law day is not |
| Fri 7 Aug + Fri 14 Aug | Keeps a weekday, keeps the pairing, and gives a full week to build the trigger — but both runs land after the law applies, so it carries the same fatal post-vs-post defect as Mon 3 + Mon 10 while measuring the baseline five days further from the transition and pushing publication two weeks past the deadline. Strictly worse than the Monday option on the axis that actually decides it |
Saturday baseline plus a weekday calibration sweep (e.g. run1-weekday, Mon 3 Aug) | Intellectually the best answer to costs 1–4: it converts an unquantified prior into a published number — the weekend effect measured on the identical cohort with the identical instrument — instead of a caveat. Rejected on capacity, not on merit. 343 rows already used + 2,284 (Saturday double) + 1,142 (calibration) + 1,142 (run 2) = 4,911 of the 6,000 cap, leaving 1,089 — less than one full sweep, so any contingency re-run would breach the wall mid-study. It also calibrates the weekend effect at a post-law moment against a pre-law baseline, which weakens the correction it exists to supply. Recorded as the option to revisit if the cap is ever raised |
18f. Defects found on 31 July while re-verifying the anchors#
Nine defects, all verified against the actual files and commands that evening, none of them caused by the date change — the date change is only what caused someone to look. The tenth, the missing automation, is §18a.
1. The OTS proofs had never been upgraded, so every "anchored" claim was a calendar commitment rather than a confirmed blockchain timestamp. Checked on 31 July, precommit-manifest-9.json.ots contained three PendingAttestations and zero BitcoinBlockHeaderAttestations — and so did anchors 2, 6 and 8; anchor 2 had been stamped three days earlier and was confirmable long before. ots upgrade had never been run on any anchor. OpenTimestamps is deliberately two-stage: stamping submits the hash to calendars, and only the upgrade — once a calendar's aggregate lands in a Bitcoin block — turns the receipt into a proof that can be verified without trusting the calendar operator. Until that evening, every "OTS-anchored" claim in these docs was true only in the weaker, calendar-trusting sense, which is not what the word anchored is doing in a precommitment.
Remedy, stated as three checkable facts rather than a claim of diligence. ots upgrade was run across all nine stamped .ots files then on disk — anchors 1–9, i.e. precommit-manifest.json.ots plus -2 through -9; manifests 10 and 11 did not exist yet and so has no proof of its own and inherits none of this — and the refreshed proofs committed:
- The proofs now carry Bitcoin.
precommit-manifest-9.json.ots— the anchor §17d calls final — now carriesBitcoinBlockHeaderAttestation(960266)andBitcoinBlockHeaderAttestation(960302)alongside the still-pending calendar paths, where that morning it carried three pending attestations and zero Bitcoin. Anchors 2–9 each carry block-header attestations; any calendar path still pending upgrades when its aggregate confirms. - The proof was confirmed to actually cover its file — a check that had never been run on any anchor in this study.
precommit-manifest-9.jsonhashes tof8911ccea6008c659e499a27a6bfd6dbcfdaa42fe0bae2b9a2b9df1e34897cef, and that same digest is the message committed insideprecommit-manifest-9.json.ots; it matches in both directions. Until this was done, "the manifest is anchored" rested on the filename of the.otsnext to it, which is not evidence of anything. - The pre-upgrade state is preserved as evidence, not merely described. Every pending-only proof was kept alongside its upgraded replacement as
precommit-manifest-*.json.ots.bak. A reader does not have to take this paragraph's word for the defect: the.bakfiles are the pre-upgrade proofs and can be inspected directly.
The lesson is procedural and unflattering: ots verify was written into §15b and §16b as the check and was never actually run for four days.
2. worker_version regressed at anchors 7, 8 and 9 — RECORDED, not silently corrected. Those three manifests record worker_version: "4af9624f (instrument of record, §13)" — the superseded Worker. Anchor 6 records the full e9a8965e-282b-4735-be4b-ae0bbc0fff26 correctly. Cause: precommit.py takes --worker-version as an optional flag over a hardcoded default; the flag stopped being passed at anchor 7 and the stale default was substituted with no warning. Why this is worse than a cosmetic slip: the 17-file bundle contains no Worker source at all — its only src/ path is the rule-pack JSON — and the runbook's drift check compares the file hashes only. worker_version is therefore the single field in the manifest that binds the grading instrument, and it was wrong in the three most recent anchors, including the one §17d calls final. Nothing about the deployed instrument moved: the live Worker has been e9a8965e since 2026-07-28T19:50:54Z (§16a) and was re-verified on 30 July (§17a). What was wrong was the manifest's record of it. Per this log's correction style the stamped manifests stay exactly as they are; precommit-manifest-11 carries the correct full UUID (it is cut after this section is text-final — §16b), and the anchor-9 → anchor-10 difference in that field is itself the audit trail.
3. precommit.py hardcoded two more fields, so anchors 2–9 misstate their own lineage. supersedes is hardcoded to tools/study/precommit-manifest.json in every manifest from 2 through 9 — so anchor 9 claims to supersede anchor 1 rather than anchor 8 — and note is a fixed §15-reframe sentence with no CLI flag, carried unchanged into anchors that exist for entirely different reasons (a date amendment, a start-time shift). Neither field is load-bearing for verification, which is done on hashes; both are the manifest's own account of why it exists, and both were wrong. Fixed in precommit.py on 31 July. precommit.py is not a bundle file, so the fix costs zero drift — which is also exactly why the defect survived nine anchors without any drift check noticing it.
The fix changed the CLI contract, and that is why the documented command grew three flags. The defect in §18f/2 was possible only because --worker-version was optional over a hardcoded default, so dropping it substituted a stale value silently. The remedy is to remove every silent default: --worker-version, --supersedes and --note are now all required arguments, alongside --out. The consequence is deliberate and worth stating so the next reader is not confused by a command that used to work: the old bare invocation python3 tools/study/precommit.py --out /tmp/check.json now exits with a usage error. A throwaway drift check is now written:
python3 tools/study/precommit.py --out /tmp/check.json \
--worker-version 'e9a8965e-282b-4735-be4b-ae0bbc0fff26' \
--supersedes tools/study/precommit-manifest-9.json \
--note 'drift check only, not an anchor'
STUDY_RUNBOOK.md precondition 3 is updated to this form in the same 31 July pass; without that edit the runbook's own drift check — the check that caught real drift four times on 28 July — would fail on a usage error at 11:00 on a Saturday. Failing loudly on a missing flag is the point: the three fields are the manifest's account of what it binds and why, and a wrong default is worse than a refusal.
4. Anchor 8 was stamped over a dirty tree. precommit-manifest-8.json records instrument_tree_clean: false. It was superseded the same afternoon and was never the live anchor, so nothing published rests on it — but §16b's account of the anchor sequence should have said so and did not. Recorded here.
Turned into a guard for manifest-11, since a historical note prevents nothing: precommit-manifest-11.json must record instrument_tree_clean: true — meaning everything, including the refreshed .ots proofs and their .ots.bak predecessors (§18f/1), is committed before the anchor is stamped, not after. An anchor taken over a dirty tree commits to a state no one can reconstruct from the repository, which is the one property a precommitment exists to provide.
5. in_band() can fail open, so a 0 on the out-of-band funnel line is a RED FLAG, not a pass. Both implementations parse the hour as int(str(ts)[11:13]) out of D1's space-separated timestamp; aggregate.py returns True when finished_at is falsy (documented, so non-graded rows keep their own exclusion class) and noise-floor.py's not ts or … does the same. If that column is ever NULL, absent, or reshaped by a different pull, the band filter silently keeps everything and the funnel's "Captured outside pre-registered hour band" line reads 0 while excluding nothing. This is not hypothetical for a pull shape: both sweep pulls already in the repo (sweep-agreement*.json) contain no finished_at key at all. §17e establishes what the correct reading looks like — under the 05:00 EDT start, run1's throttle retries land 13:00–15:15 UTC and straddle the 15:00 UTC cliff — so for run1 that line must not be ~0. Reconcile it against the recorded throttled scan ids, as §16g already instructs, and read a bare 0 as a filter that did not run rather than a sweep that behaved.
6. The anchors are stamped to THREE calendars, not four. Every .ots from anchor 2 onward carries paths to bob.btc.calendar.opentimestamps.org, btc.calendar.catallaxy.com and finney.calendar.eternitywall.com — three (anchor 1, the 27 July manifest, carries two). "Four calendars" appeared three times in this log (§15b, and twice in §16b) and is corrected in place, since it is a plain factual error about files anyone can read, not a historical judgement that deserves preserving.
7. Cap headroom is tighter than §16a's line suggests. The runaway wall is 6,000 cohort rows (§9). Actual arithmetic: 343 already in D1 (agreement 144 · agreement2 144 · smoke 55) + 2,284 for the Saturday double sweep + 1,142 for run 2 = 3,769, leaving 2,231. That is less than one further full double sweep (2,284). So if the baseline has to be redone in full — a botched slice, a duplicated cohort, a reconciler surprise — the contingency does not fit under the wall without an explicit decision to raise it. Stated now so that decision is not improvised at 11:00 on a Saturday.
8. noise-floor.py never checks that its two sweeps ran on the same calendar day, yet labels its output "same-day". in_band() discards the date and reads only the hour, and the tool unconditionally stamps "note": "same-day same-code double sweep; supersedes §13a's cross-code upper bound" into its output. Same-day is a runbook convention, not a property the code enforces or the output can evidence. Nothing about the plan changes — both sweeps are on 1 August — but the label asserts more than the tool can support, and §16d's noise floor should be read knowing that.
9. study-run.py's enqueue retry is not idempotent, and the failure surfaces days late. post() retries up to three times on transport errors, timeouts and 5xx/429. /api/_study_run mints a fresh scan UUID per URL and blind-INSERTs. So if a 50-URL slice's response is lost after the Worker has already committed its 50 inserts, the retry re-posts the same 50 URLs and creates 50 duplicate rows. The resume path cannot see it — done = len(record["results"]) is index-based over the slices the driver believes it completed. aggregate.py and noise-floor.py both hard-refuse on duplicate targets (§16f point 2 watched that guard fire on real data), so the consequence is not a quietly corrupted headline; it is an aggregation refusal discovered after the band has closed and after the §6 freeze has started, with a full re-sweep as the only remedy and defect 7's headroom to pay for it. Mitigation added to the runbook: immediately after each enqueue, run SELECT target, COUNT(*) c FROM scans WHERE cohort='<cohort>' GROUP BY target HAVING c>1 — it must return zero rows. A non-empty result caught while the band is still open can be retired with §16g's run1-superseded procedure; the same result found at aggregation cannot.
10. The freeze-verification command in the runbook read the WRONG END of the deployment list — found by dry-running the guard rather than by reading it. Precondition 2 said to verify the instrument with npx wrangler deployments list | head, and both scheduled tasks inherited the same idiom (| head -20). wrangler deployments list prints in ascending date order: head shows the oldest deployments — as of 31 July, a 27 July version — and the live deployment is last. Two ways this bites, and the second is worse than the first:
- An operator verifying the §6 freeze sees a version that is not
e9a8965e, concludes the instrument changed, and aborts a sweep that should have run. On a date with no slack left, that is a lost baseline. 4af9624fsits ABOVEe9a8965ein that list. So the wrong command shows the superseded instrument — the exact version defect 2 records three anchors as having wrongly named — in a position that reads as current. A check designed to catch instrument drift would have actively confirmed the wrong instrument.
Fixed in the runbook and in both scheduled tasks: use npx wrangler deployments status, which prints the current deployment and nothing else (deployments list | tail also works).
The anchor consequence, recorded rather than tidied away. This was found after precommit-manifest-10 had already been cut, stamped and pushed. Fixing it meant editing STUDY_RUNBOOK.md, which is a bundle file — so anchor 10 no longer matched the text that would run, and precommit-manifest-11 supersedes it within the hour of its being stamped. Anchor 10 stays in the repo with its proof, as anchors 7 and 8 do, because the sequence is the audit trail: it shows a defect being found by rehearsal and paid for immediately, rather than a document that was right the first time. Per §16b, a stamped manifest is never re-cut under its own number — its hash is already on the calendars.
The general lesson, which is the same one §18a draws about the trigger: this was found by executing the guard, not by reading it. It had survived every eve-of-sweep pass on 29, 30 and 31 July because each of those verified the instrument — correctly, by other means — and never ran the precise command the document prescribes. A verification step is not verified until someone runs it verbatim and looks at what comes back.
18g. Two copy requirements for the published piece, settled here rather than at publication#
1. The "on the day" headline was never satisfiable by ANY candidate date. STUDY_PIPELINE.md §0's drafted headline — "On the day Article 50 took effect, N% of EU-facing sites running a chat widget showed no AI disclosure" — describes a day on which nothing was measured, and that has been true of every schedule this study has ever held: 30 July + 6 August, 31 July + 7 August, and now 1 August + 8 August. Article 50 applies on Sunday 2 August and no sweep runs on 2 August. This is preflight-findings.json entry 43 (first-reader lens, serious), whose fix is to phrase the claim as measured and put the scan dates in the headline graphic and the first paragraph. Adopted: the published headline names the scan dates — before and after Article 50 began to apply (scans 1 and 8 August 2026) — and never claims a measurement on the day itself.
2. Headline 1 and headlines 2–4 are measured on different dates, and the copy must carry both. Headline 1 — corrected prevalence 11.3% (95% CI 7.95–15.42), §15e — is computed by prevalence-estimator.py from the discovery NDJSON, the frozen cohort file, and the audit and recall-expansion judgements. It reads no sweep data at all, and every one of those inputs was produced on 27–28 July 2026. No sweep can move it. Stamping a single scan date across the whole piece would therefore misdate the study's lead number by four to five days and attach it to a Saturday on which it was not measured. Required wording, explicit and two-dated: the frame and its bias correction were measured 27–28 July 2026; the surfaces were measured on the sweep dates. This is a copy rule, not an analysis change, and it is recorded here so it survives the gap between this log and the published page.
19. The baseline as executed (2026-08-01)#
§17 and §18 record two pre-registered baseline dates that did not happen. This section records the one that did. It is written the same day, after the sweep and before any decision about §16d, and it is not a clean story: the baseline exists, the noise floor does not, and the reason is a failure mode §18a named in writing the evening before.
19a. The timeline, as fact#
Both scheduled tasks fired on time, to the second. Everything that went wrong went wrong after that.
| Time (UTC) | Event |
|---|---|
| 09:00:15 | study-run1-enqueue-sat1aug fires — on schedule |
| 09:00:20 | Guard A reads date -u → 09:00 UTC, in band. PASS |
| 09:00:27 | The task issues two npx wrangler calls (Guards B and C) — and blocks, awaiting tool-permission approval |
| 11:30:15 | study-run1noise-enqueue-sat1aug fires — also on schedule — and blocks the same way |
| 13:47:40 | The run1 task's tool calls return. 4 h 47 m of wall clock, with nobody at the keyboard: the scheduled hour was 05:00 EDT |
| 13:48:13 | run1 enqueue begins — 4 h 48 m late |
| 13:51 | run1-noise resumes, re-reads the clock, and aborts at Guard A (13:51 UTC is past the 13:00 UTC cutoff) |
| 14:11:52 | run1 enqueue complete: 1,098 rows |
| 14:56:45 | The last graded row finishes — 3 min 15 s before the 15:00 UTC band cliff. This is the margin the sweep actually had |
| 15:00:33 | The last row of the sweep reaches terminal state — an error row, 33 seconds past the 15:00 UTC band cliff |
What the stall cost run 1's validity as an in-band pre-registered sweep: nothing. All 1,073 graded rows landed inside BAND_UTC = (7, 15), the out-of-band funnel line is a genuine 0 rather than §18f/5's fail-open (§16c), and the single row that crossed the cliff was an error row already excluded upstream. The baseline is usable exactly as pre-registered, on the frozen instrument, on the frozen cohort.
What it cost run 1's levels is not nothing, and the unscoped version of the sentence above is withdrawn. It read "What the stall cost run 1: essentially nothing", which the rest of this section then contradicts. The stall moved the measurement window out of the 11:00 CEST hour it was scheduled for and into 15:48–17:00 CEST on a Saturday — an unmeasured shift toward the least-staffed hours of the week. Fewer widgets live and fewer panels opening push headline 3's "could not read the first-interaction surface" share up, and the vendor and disclosure-state levels with it: the direction is known, the magnitude is not, and §16c records that this study has no instrument left that could size it. So the scoped statement is the whole statement — nothing was lost to the sweep's validity, and something unquantified was done to its numbers.
How close that was — stated rather than glossed, and the figure this paragraph used to give was wrong by about 6×. The sweep finished with minutes to spare, not tens of minutes and not hours. The last graded row finished at 14:56:45 UTC: 3 min 15 s before the 15:00 UTC cliff. Bulk grading ran deep into UTC hour 14 — 894 of the 1,073 graded rows landed in that hour alone — so a further three and a quarter minutes of stall, not twenty, would have begun discarding graded rows. By §17b's ordering argument that tail would not have been a random tail, because cohort-final.csv is ordered by ascending Tranco rank and study-run.py enqueues it in file order, so the rows finishing last are systematically the least-trafficked sites, biased on precisely the variables under study. run1 did not survive on design margin. It survived on about three minutes. An earlier draft of this paragraph said twenty; the error was roughly sixfold and it ran in the direction of complacency, which is the direction this log is least entitled to err in.
What the stall destroyed outright: the noise floor (§16d) — and destroyed is the right verb, not delayed: a same-day pre-application floor for this baseline cannot be obtained on any later date. That is the largest single item of the damage. It is not, per the paragraph above, the only one: the displaced measurement window is a second, smaller, unquantified cost that lands on the levels rather than on the design.
One uncomfortable detail that belongs beside that timeline. run1's Guard A was evaluated once, at 09:00:20, and was not re-evaluated when the blocked tool calls returned; the enqueue at 13:48:13 proceeded on a band reading that was 4 h 48 m old. Had it re-read the clock the way the run1-noise task did, Guard A's own rule — current UTC hour ≥ 7 and < 13 — would have aborted run1 as well, and there would be no baseline at all. So the honest summary of the day is that the study got a usable baseline out of a guard reading a stale clock, and lost its noise floor to a guard reading a fresh one. Both sentences are true and neither is dropped for being awkward. The consequence for run 2 is concrete: the band guard must re-read the clock immediately before the enqueue call, not only at task start. A check taken before a blocking call is not a check on the state at the call.
19b. The lesson: an automated trigger inherits the permission model of the thing it automates#
Stated as bluntly as §18a states its own, because this is the second time in three days that a control was believed to be stronger than it was.
The remedy adopted on 31 July was correct, and it still nearly failed. §18a's diagnosis — that no automation had ever existed, and that a schedule living in a markdown table depends on a human reading that table at the right hour on the right day — was right, and the trigger it produced is why a baseline exists at all. It converted a total miss into a partial one. That is a real improvement and it is not the point.
The point is that a scheduled task which runs privileged commands is only as automatic as its tool approvals. The trigger fired itself, evaluated its own guard, and then stopped dead at the first npx wrangler call, because that call needed an approval that had never been granted. An approval prompt with nobody at the keyboard is indistinguishable from a hang: no timeout, no alarm, no exit status, no notification. The automation was complete right up to the first thing it actually needed to do.
§18a predicted this, named the mechanism, and prescribed the fix — and the fix was not performed. Its closing paragraph is unambiguous on all three counts: the guards are "instructions in a SKILL.md carried out by a model, not code with an exit status"; "the trigger has never fired", with no dry run against a real slot; and a test-fire "aborts harmlessly at Guard A outside the band, while still exercising and storing the tool approvals the task needs" — a parenthesis that names the exact mechanism that then failed. It was on the eve-of-sweep checklist. It was not done. This is the same failure shape as §18f/10, which found a broken verification command by running it after three consecutive eve-of-sweep passes had read it: a control is not verified until it has been executed end to end. §18f/10 drew that lesson on 31 July, and §19 is the price of not applying it to the trigger the same evening.
The rule for run 2 — a rule, not a suggestion:
- Test-fire both scheduled tasks days ahead, from a cold app start, deliberately outside the band, so Guard A aborts them harmlessly while every privileged call is exercised and its approval stored.
- Test-fire to exhaustion, not to first success — and the earlier statement of what remains untried was wrong, in this task's favour. It read that "Guards B, D and E, the enqueue itself, and the mandatory §18f/9 duplicate-target query are all still untried in a real slot," which contradicts the timeline directly above it. Once the approval landed at 13:47:40,
run1's Guards B and C returned and the enqueue ran to completion — 1,098 rows, all terminal (§16c). Those steps are tried. What is untried is narrower, and it is the part that actually failed: no privileged call in either task has ever completed from a cold start without a human present to grant an approval at the moment it was demanded. Everything after 13:47:40 ran attended, four hours and forty-eight minutes late; that is a test of the task, not of the automation. And onrun1-noise, nothing past Guard A has ever executed at all — its Guards B through E, its enqueue and its duplicate query remain untried in any slot, real or dry. A test-fire that stops at the first approval prompt has verified one line of the task; a test-fire that a human unblocks has verified the commands and left the unattended path exactly as unproven as before. - Treat an un-test-fired trigger as NOT a mitigation. §18a rated the trigger's residual risk as "much lower than a human alarm that has now failed twice." Un-test-fired, it was not lower in the way that mattered — it moved the point of failure from the alarm clock to the permission dialog and left the outcome (a sweep that does not start on time) unchanged. A control whose first execution is the real event has been designed, not tested.
- §18a's operator requirement is upgraded. "The app open and the Mac awake at 11:00 and 13:30 CEST" is insufficient if an approval can block indefinitely. Either the approvals are pre-stored (1–2 above), or a human must be reachable within minutes of the fire time — and the first is the one to rely on, because the second is the dependency §18a was written to remove.
19c. Credit where it is due — narrowed to what Guard A actually did#
Recorded for the same reason the failures are: it is evidence about the design, not self-congratulation.
Guard A refused a sweep that would have been thrown away. At 13:51 UTC the run1-noise task re-read the clock, found itself 51 minutes past its own cutoff, aborted without enqueueing, and reported which guard had failed and why. It also flagged that its pre-stall date -u reading no longer matched reality — it did not trust a value it already held across a four-hour gap.
The comparison this section used to draw was rigged, and it is withdrawn. It read: "A cron entry, a launchd agent, or an alarm with no band check would have enqueued 1,142 scans at 13:51 UTC." That is false for the two named alternatives, and false in the direction that flatters this design. A cron entry or a launchd agent would have fired at 09:00 UTC and would not have stalled at all — an OS-level trigger does not sit behind a tool-permission dialog, so the 4 h 47 m block is a property of this mechanism, not a hazard that cron shares and this design survived. Against cron, the day's scorecard is worse, not better: cron would have swept in band at the registered hour, and run1-noise would exist. Crediting a guard for surviving a failure the comparator does not have is not evidence about the design; §18a's "revisit a Cloudflare cron" is parked for the freeze reason in §19d, not because this trigger beat it.
The like-for-like comparison — this task without its band check — is the one that holds, and the claim it supports is conditional. Given a resumption at 13:51 UTC, Guard A prevented a wasted sweep. The same task with Guard A removed would have enqueued 1,142 scans at 13:51 UTC, drained a fraction before the 15:00 UTC cliff, and reported success. §16f point 2 measured that yield on real data: ~2,300 real third-party scans for zero usable rows. It would additionally have spent a fifth of the 6,000-row cap on the last pre-Article-50 day, and — worse — left run1-noise rows in D1 that Guard B must then refuse and that both aggregate.py and noise-floor.py hard-refuse to pair on duplicates (§18f/9), turning an empty slot into a contaminated cohort. That is a real benefit and it is the whole of the benefit. The condition attached to it was manufactured by the mechanism the guard is part of.
§18a's justification for Guard A does not survive the day intact either. It reads: "The residual risk is a missed sweep, which is recoverable; the risk the guard removes is a silently wasted sweep, which is not." The first clause is the one that failed. It was written on 31 July, when the phrase "recoverable" still had a calendar behind it; on 1 August it did not. What was missed was the same-day pre-application noise floor, and §16d records that object as permanently unobtainable — there is no later date on which it can be collected, because the property that defined it was the date. So only the second clause stands: a silently wasted sweep would indeed have been worse, because it would have cost the cohort as well as the slot. Both outcomes were unrecoverable on the last pre-law day; one of them was also destructive. Guard A chose the cheaper unrecoverable outcome. That is the correct choice, and it is a smaller thing than "a missed sweep is recoverable" made it sound.
The narrow claim, stated precisely so it is not read as more than it is: the guards remain model-executed instructions rather than code with an exit status, one of them evaluated a stale reading (§19a), and the only truly mechanical enforcement is still aggregate.py's downstream band filter. Within those limits, Guard A behaved as written on the one day it mattered, and its first real execution was a correct abort on a real slot. That is a claim about one guard, not about the mechanism containing it: the mechanism started 4 h 48 m late and lost the noise floor, and §19b is where that is scored.
19d. The §6 instrument freeze is live#
Active from 13:48:13 UTC on 1 August 2026 — the first run1 enqueue — and binding through the completion of run 2.
| Frozen | Value |
|---|---|
| Worker | e9a8965e-282b-4735-be4b-ae0bbc0fff26 @ 100% |
| Rule pack | 2026.07.9 |
| Tree at sweep | commit 674fae4 |
| Cohort | the frozen 1,142-site cohort-final.csv, sha256 c81ae303…f507ed |
No deploy touching capture, interaction (consent lists included), robots, fingerprints/recipes or detectors until run 2 is complete. This binds any sweep between now and run 2, whatever is decided about §16d. It is also why §18a's "revisit a Cloudflare cron" stays parked until after 8 August: a deploy to fix a scheduling problem would mint a new Worker version and break the freeze mid-study, which is the trade §18a already ruled on.
19e. There will be no anchor 12#
Writing this section drifts precommit-manifest-11, by construction, and that is correct. STUDY_METHOD.md is a bundle file, so filling §16c, rewriting §16d and appending §19 changes its hash. §18 said this would happen the moment the text moved, and §16b has said since 28 July that the bundle contains the document that describes the bundle.
The response is not a twelfth anchor. Anchors stop at 11. The reason is what a precommitment is for: it proves the analysis plan was fixed before the data existed. precommit-manifest-11 was stamped on 31 July — before the sweep. That stamp, and only that stamp, is what predates the data. Its ots upgrade to a confirmed Bitcoin timestamp — three BitcoinBlockHeaderAttestations, anchor 10 upgraded in the same pass — ran at ~16:05 UTC on 1 August, after the sweep had finished. An earlier draft of this paragraph let the "before the sweep ran" qualifier slide onto the upgrade as well as the stamp; it is reattached here to the stamp alone.
Why the correction matters even though the proof is unaffected. What an OpenTimestamps proof establishes is that the document existed no later than the moment its digest was submitted to the calendars — 31 July. The upgrade changes what a verifier must trust, replacing the calendar operator's word with a Bitcoin block header; it cannot move the committed time in either direction. So the precommitment is exactly as strong as it would be had the upgrade run a week later. But writing "upgraded to Bitcoin before the sweep" states a fact about when evidence was gathered in the grammar of a fact about when the commitment was made, and a reader checking the dates would find the second one false. That is precisely the class of overclaim §18f/1 caught, when every stamped anchor turned out to be carrying pending-only proofs while these documents called them anchored — anchors 1–9 held nothing but calendar receipts until the evening of 31 July. Having found that defect by reading dates carefully, this section does not get to be sloppy with its own.
An anchor cut now would timestamp a document that has read the data; it would prove nothing about pre-registration, and it would force every future reader to work out which anchors are pre-data and which are not. That is a worse record, not a better one. §16b's argument for re-anchoring — anchor last, say why — applies only while every anchor still predates the data, and from 13:48:13 UTC on 1 August that condition no longer holds.
What a sceptic should check instead, all of it mechanical:
ots verify tools/study/precommit-manifest-11.json.ots— the proof now carries Bitcoin, and §18f/1's stronger check (that the proof commits to the manifest's own digest, in both directions) applies to it. Read the attestation for what it certifies: the submission is dated 31 July, pre-sweep; the upgrade that turned the receipt into a block-header attestation was run at ~16:05 UTC on 1 August, post-sweep.- The 17 bundle files as of commit
674fae4— the tree that swept — against manifest-11's recorded hashes: byte-identical. Not against today's tree, which contains this section. - One seam, recorded rather than tidied away in the style of §18f/2: manifest-11 records
instrument_commit: 28938c525b55506b58b91cf3cbba27b1f09804c7while the tree that swept was674fae4. The commits differ; the bundle contents at674fae4are byte-identical to anchor 11's recorded tree, which is the claim that carries the precommitment and the one to verify first. - Verify against
674fae4, not against today's tree. This is stated above but is the one instruction that makes the check pass, so it is repeated as an instruction rather than left as context: at HEAD exactly two of the 17 files drift —STUDY_METHOD.mdandSTUDY_PIPELINE.md, both by construction (§19e, §21). A reader told to expect one expected mismatch and finding two has been given a wrong prediction, not a caveat.
ADDED 2026-08-02 (§21e) — this section understates its own evidence, and the correction runs in the study's favour. The paragraphs above are careful to keep "stamped 31 July, pre-sweep" apart from "
ots upgraderan ~16:05 UTC on 1 August, post-sweep", and that separation is right. But they leave a reader with the impression that no Bitcoin confirmation predates the sweep, and the published draft compressed them into exactly that claim: "Nothing here claims a confirmed blockchain timestamp predating the sweep." All three do. Parsingprecommit-manifest-11.json.ots, its threeBitcoinBlockHeaderAttestations sit at heights 960456, 960481 and 960506, mined 2026-07-31T19:53:57Z, 2026-07-31T23:51:35Z and 2026-08-01T03:50:58Z. The sweep enqueued at 2026-08-01T13:48:13Z. The earliest attestation therefore precedes the sweep by 17 h 54 m, and all three precede it.Why both statements are true at once.
ots upgradeis a retrieval, not a commitment: at ~16:05 UTC on 1 August it fetched attestations that had already been committed to blocks mined the previous evening and overnight. The command ran after the sweep; the blocks it retrieved were mined before it. §19e's original wording is not false — it describes when the command ran — but it describes the weaker of the two available facts, and this is the one paragraph on the page written for a reader with a terminal. Publish the three heights and their times so the check needs no repository:ots info tools/study/precommit-manifest-11.json.ots, then compare each block's timestamp against the 13:48:13Z enqueue. This does not change what the precommitment proves — the 31 July submission is still what carries it, and an upgrade cannot move a committed time in either direction — but understating one's own evidence in a section about not overstating it is the same class of carelessness §18f/1 was written to catch.
19f. What the study can claim without a noise floor, and what it cannot#
Stated now, before any recovery decision exists, so that the decision cannot be shaped to fit a claim already published.
The premise, stated first because everything below depends on it: the floor §5 registered is permanently unobtainable, not pending. §16d gives the argument. A same-day, pre-application double sweep of this cohort could only have been collected on 1 August 2026, and it was not. Any later measurement is a post-application floor, from a different day, and possibly a different week — a different object, useful or not on its own merits, and never a retrospective satisfaction of §5. This section is therefore not an interim position awaiting recovery. It is the standing rule, and it is written to hold even if nothing further is ever collected.
Unaffected by the missing floor — all four §15 headlines. Every one of them is a level, and a level needs no noise floor: a floor is a property of a comparison, not of a measurement. That is a statement about §16d only. Three of the four are separately affected by the stall's other consequence — the measurement window moved to 15:48–17:00 CEST on a Saturday (§16c) — and the status column carries both.
| §15 headline | Source | Status |
|---|---|---|
| 1. Corrected widget prevalence — ~11% (95% CI 8.0–15.4) over 12,285 frame sites, against the raw detector's 9.4% | 27–28 July frame, audit and recall-expansion data (§15e). Reads no sweep data at all | Unaffected. §16d cannot touch it, and no sweep can move it (§18g/2's two-dated copy rule). The stall cannot touch it either: it reads no row that the stall's hours produced |
| 2. Vendor distribution over sweep-confirmed sites | run1 (§16c) | Unaffected by §16d. Carries §18d/1's Saturday sensitivity and the late-afternoon measurement window (§16c) — confirmation is an observation made at 15:48–17:00 CEST, so vendors whose customers run early-closing rotas lose share for reasons unrelated to install base. Direction known, magnitude unmeasured |
| 3. Coverage — 78.3% (CI 75.2–81.0) of 791 confirmed widgets unreadable, split by cause | run1 (§16c) | Unaffected by §16d. Carries §18d/1, §15b's pending audit correction of the "panel did not open" bucket, and the late-afternoon measurement window (§16c), which plausibly INFLATES the 78.3% — fewer live widgets and fewer panel opens in the least-staffed hours. Direction known, magnitude unmeasured, and unmeasurable from here |
| 4. Disclosure-state distribution on the readable subset (n = 172) | run1 (§16c) | Unaffected by §16d. Keeps its selection-bias framing and stays below the fold — and the selection is more severe than §10d anticipated, because the readable subset was drawn at 15:48–17:00 CEST on a Saturday (§16c). Direction known, magnitude unmeasured |
CORRECTED 2026-08-02 (§21b). Headline 3's status cell above ends "Direction known, magnitude unmeasured, and unmeasurable from here". "Unmeasurable from here" is too strong and is withdrawn as written. A weekday/weekend capture of the same sites does exist —
sweep-agreement2.json(Tue 28 Jul) againstsweep-run1.json(Sat 1 Aug), 143 shared targets, 139 graded in both — and it was already in the repo when this table was written. The cell's conclusion stands: that pair cannot size the late-Saturday-afternoon shift, for the four reasons in §21b, and no number from it is published as a bound or as a correction. The accurate statement is "magnitude unmeasured, and the only same-site weekday/weekend pair in the repo cannot bound it (§21b)" — which is a claim about what the available data can support, not the stronger and false claim that no such data was ever captured.
Weakened — the run1 → run2 delta, and only that. With no same-code floor, any movement must be judged against §13a's cross-code upper bound: ~4 points on widget presence, ~8 points on readable surfaces. Every reason §13a called itself an upper bound still applies, and two more now attach:
- It is cross-code: its second run used post-fix code, so it confounds an instrument change with genuine run-to-run noise. §5's floor existed to remove exactly that confound and now does not.
- It rests on 137 sites, not 1,142.
- It was measured on Monday 27 July, so it cannot see between-Saturday variance. That is §18d/3's objection made worse rather than answered: §18d/3 complained that a same-Saturday floor is systematically too tight for a Sat→Sat delta, and the floor actually available is not even a Saturday floor.
- A bound this loose is unlikely to distinguish a real Article 50 effect from noise at any plausible effect size — and §18d/2 has already established that the widget-presence delta is attenuated toward zero by weekend-dark sites sitting inertly in the denominator. The two failure modes point the same direction: toward a null that is indistinguishable from an under-powered design, which §18d/2 identifies as the single conclusion this study is least able to defend.
Why this is a real but bounded loss. §15 lists four headlines and the delta is not among them; §15a already scoped it to within-cohort widget-presence movement, an at-most one-sided cohort-loss bound, with no frame-level prevalence delta published at all. What has been lost is a secondary result's precision, not a headline. That is the honest size of it — genuine, and not fatal.
The rule that follows, binding on the published piece. Until and unless a same-code floor exists, any run1 → run2 movement is published with §13a's bound named as the comparator and its cross-code, 137-site and Monday limitations stated in the same sentence — never as "the noise floor" unqualified — and §5's pre-registered wording is honoured: a delta smaller than the floor is reported as not distinguishable from measurement noise. If a substitute is later obtained, this paragraph is superseded by whatever section records it, and the comparator changes in one place rather than silently across the piece — but that section inherits §16d's constraint and must state, in the same breath as the number, that what it reports is a post-application floor from a different day standing in for a pre-application same-day floor that no longer exists. A substitute may replace the comparator. Nothing can replace the object, and the published piece does not get to imply otherwise by quietly calling a later measurement "the noise floor."
20. The substitute noise floor (decided 2026-08-02, before run 2) — NOT OBTAINED (2026-08-10). §19f's rule revives unchanged — see §26#
MARKED NOT-OBTAINED 2026-08-10, exactly as §20d's final bullet prescribes and exactly as §16d was.
run2-noiseaborted un-enqueued at 12:46 UTC on 10 August — the frozen driverstudy-run.pyrejects therun2-noisecohort label, which was coined in the scheduling docs after the driver was frozen and could never have been added to it (tools/study/PREFLIGHT-2026-08-10-run2noise-ABORT.md). Zero rows were written. The founder decided the same day: drop the floor (§26). The study holds no noise floor of any kind — not the pre-registered §5 object (§16d, permanently unobtainable) and not this substitute. §13a's cross-code bound is again the study's only noise estimate, cited only under §19f's rules. Nothing in §§20a–20e below is rewritten; it is the decision record of a measurement that was never made.
the founder's call: obtain a substitute. Run-2 day becomes a same-day full-cohort double sweep — run2 at 11:00 CEST and run2-noise at 13:30 CEST on Saturday 8 August 2026, the same procedure §15 registered for baseline day. Both scheduled tasks exist (study-run2-enqueue-sat8aug, study-run2noise-enqueue-sat8aug) with the §19b guards.
This section is the one §19f said would supersede its comparator rule, and it inherits §16d's constraint in full. It is written before run 2 so that what the substitute can support is fixed before its numbers exist.
20a. What the substitute is#
A same-day, same-weekday, same-band, same-code test-retest measurement of the instrument's run-to-run variation, graded under the frozen Worker e9a8965e-282b-4735-be4b-ae0bbc0fff26 and pack 2026.07.9 that graded run1 — the §6 freeze has been live since 13:48:13 UTC on 1 August precisely so this identity holds. That is a real measurement and it is worth having: it bounds how much of a run1 → run2 movement is the instrument disagreeing with itself.
20b. What it is not — four limits, none of which a later number retires#
- It is not the pre-registered object. §5 registered a floor measured on baseline day, before Article 50 applied. This one is measured six days after the Regulation applies. Per §16d that object is permanently unobtainable; nothing here changes that, and the published piece never calls this "the noise floor" without the qualifier.
- It characterises 8 August, not 1 August. Applying it to run1 is an assumption — that the instrument's self-disagreement is stable across the week — not a measurement. The assumption is plausible under a frozen instrument and is stated rather than relied on silently.
- It still cannot see between-Saturday rota variance (§18d/3). A same-day floor never could; this inherits that limit rather than fixing it, so a Sat→Sat delta is still judged against a floor that is systematically too tight.
- The hour-of-day mismatch, which is the sharpest of the four — see §20c.
20c. Which hour run 2 starts, and its reason (this answers §16e/1, as §16e requires)#
Decision: the registered slot — run2 at 09:00 UTC, run2-noise at 11:30 UTC. Run 1 was scheduled for 09:00 UTC and, after the §19a stall, actually graded in UTC hours 13–14. Run 2 can match the registered plan or run 1's realised hours, not both. Reasons for the registered slot:
- Matching run 1's realised hours means enqueueing at ~13:45 UTC. Run 1's last graded row landed at 14:56:45 UTC — 3 min 15 s before the cliff (§19a). That is not a margin, it is a coin flip, and losing run 2 loses the delta outright and the substitute floor with it.
- It would also deliberately reproduce a measurement window §16c records as biasing the levels in a known direction — late Saturday afternoon, close to the least-staffed hours of the week. Run 2's levels are the post-application snapshot in their own right; degrading them on purpose to buy comparability trades a primary result for a secondary one.
The cost, stated in the direction it actually runs. The run1 → run2 delta is now confounded with hour of day, and the confound is not sign-neutral: run 1 was measured in a poorly-staffed window and run 2 will be measured in a better-staffed one, so more widgets should be live and more panels should open on 8 August even if nothing whatsoever changed. The delta is therefore biased toward showing improvement after Article 50 began to apply — which is exactly the finding this study has the least standing to announce, and exactly the direction a reader would suspect us of wanting. It is recorded here, before the data exists, for that reason.
Note this points opposite to §18d/2's attenuation-toward-zero, and the two do not cancel: they are different mechanisms of unknown relative size, and claiming they offset would be a third error.
20d. What follows for the published piece — superseding §19f's comparator rule#
- The substitute, once measured, becomes the comparator in place of §13a's cross-code bound, and §13a's bound is retired from that role rather than quoted alongside it.
- Every citation of it carries, in the same sentence, that it is a post-application floor from a different day standing in for a pre-application same-day floor that no longer exists.
- Any run1 → run2 movement is additionally reported with the hour-of-day confound and its direction (§20c). A movement toward better coverage that is smaller than the confound could plausibly produce is reported as not distinguishable from the timing difference, not as an Article 50 effect.
- If
run2-noisealso fails to run, §19f's rule revives unchanged and this section is marked not-obtained, exactly as §16d was.
20e. Bookkeeping#
Cap after run-2 day: 343 prior + 1,098 (run1) + ~1,098 (run2) + ~1,098 (run2-noise) ≈ 3,637 of 6,000, leaving ~2,363 — still less than a further full double sweep (~2,196 plus overhead) would comfortably allow twice, so this is the last double sweep the cap supports with contingency intact.
There is no anchor 12, and this section does not create one. It is written after the baseline data exists, so it is not a pre-registration of anything — it is a decision record, and it says so. Anchor 11 remains the precommitment, pre-data and Bitcoin-confirmed; §19e explains why nothing is re-anchored. What binds this section is that it was written before run 2 ran, which the git history shows and which is the only claim it makes about its own timing.
21. Corrections from the Phase 5 adversarial review (2026-08-02)#
The Phase 5 adversarial review (tools/study/ADVERSARIAL-REVIEW-2026-08-02.md) returned a DO NOT PUBLISH gate decision with 19 blocking defects. Most of them are on the published draft rather than in this log, but seven reach this log, and they are recorded here in full rather than quietly patched at the point of occurrence. Each one is also marked inline where it occurs, so a reader arriving at the original sentence is told immediately that it was corrected and where to read why.
Three standing rules for this section.
- Nothing above is rewritten. §17, §18 and the executed-run records (§16c, §16d, §19) stay as they were written, including the parts later falsified. What was believed at the time is the record; corrections are additive and dated. Two treatments, applied deliberately and named here so the difference is not mistaken for sloppiness:
- Amendment and run records — §16c, §18d/4, §19e, §19f — keep their original sentences untouched, with a dated correction block appended immediately below.
- Design and counts sections — §4, §5, §7 — have the false number corrected in place, because those sentences are read as statements of what the instrument is, and leaving "24 official EU languages" standing would keep propagating it. In every case the original wording is quoted verbatim inside the correction block that follows, so nothing is lost.
- Either way the pre-correction bytes remain recoverable: this file is a bundle file, and its state at commit
674fae4is the treeprecommit-manifest-11hashes.
- Every correction moves toward more disclosure, not less. Two of them (§21b, §21e) strengthen the study's position and were made anyway, because a log that only corrects in its own favour is not a log.
- This section drifts
precommit-manifest-11by construction, and there is still no anchor 12.STUDY_METHOD.mdis a bundle file; §19e already settled that writing after the data means drift, and that a post-data anchor would prove nothing about pre-registration. Verify the bundle at commit674fae4, never at today's tree, whereSTUDY_METHOD.mdandSTUDY_PIPELINE.mdboth differ on purpose.
21a. The lexicon covers 23 of the 24 official EU languages, plus Catalan — and the consent list covers 23#
Review item #4. Corrected inline at §4 (consent accept phrases), §5 (the pre-registration bullet) and §7 (the exposure tally). The facts, so the check needs no repository:
| Claim as written | Truth |
|---|---|
| "extended pre-baseline from 6 to 24 EU languages" (§5) | 24 keys in TERMS_BY_LANG: 23 official EU languages + Catalan, with Irish (ga) absent |
| "covers the 24 official EU languages in rule pack 2026.07.9" (published draft) | Same defect, on the page |
| "covers all 24 EU official languages in its accept-phrase list" (§4, consent) | CONSENT_ACCEPT_PHRASES carries 23 languages; Irish absent here too |
| "correctly outside the EU-24 set" (§7) | The set is not "EU-24"; the adjudication of that one Russian-language site is unaffected |
Keys, in full, so the count is checkable by eye: bg ca cs da de el en es et fi fr hr hu it lt lv mt nl pl pt ro sk sl sv. Irish has been an official EU language since 2007 and its derogation ended 1 January 2022, so "official" is not arguable. There is no recorded decision to exclude it: tools/study/lexicon-proposals.json — the archive of the 2026-07-27 extension, including its kill list and reasoning — carries 18 language keys, ca among them, and the strings ga, Irish and Gaeilge appear in it zero times. The number 24 was read off the key count and never checked against the official list, twice, in two different lists.
Live exposure: 12 .ie cohort sites have a confirmed widget at sweep time and 4 of them are in the readable subset of 172. So this is not a hypothetical gap; it sits under roughly 2% of the disclosure-state denominator.
Irish is not being added. The §6 instrument freeze runs until run 2 completes, and adding a language mid-study would confound the run1 → run2 delta with an instrument change — the exact defect §13a already carries and §5's freeze exists to prevent. The correct fix is the description, and it is made here. Adding ga is a post-run-2 item.
21b. The design did capture the same sites on a weekday and on a weekend#
Review item #7. The claim was made three times in this log and once on the published draft, in these words or close to them: "Nothing in this repo measures a weekday or weekend effect at all: no discovery pass, no audit and no sweep has ever captured the same sites on both." It is false, and it is falsifiable in one command against two files that ship in the snapshot. Corrected inline at §16c, §18d/4 and §19f.
What exists. sweep-agreement2.json (cohort agreement2, enqueued Tue 2026-07-28 03:41:57Z per agreement2-enqueue.json) and sweep-run1.json (Sat 2026-08-01, graded 13:48–15:00 UTC) share 143 targets, of which 139 are graded in both.
Why the sentence was written and why it survived. It was true when first written on 31 July: the only weekend sweep in the repo did not exist yet. Run 1 created the falsifying pair on 1 August, and none of the three copies was re-read against the new data. The same class of miss as §18f/1 — a sentence that was accurate on the day, then quietly aged into an overclaim.
What the pair actually shows. Computed 2026-08-02 over the 139 sites graded in both, reproducible from the two shipped files:
| Measure | Tue 28 Jul (05:41 CEST) | Sat 1 Aug (15:48–17:00 CEST) | Movement |
|---|---|---|---|
| Widget present (paired, n = 139) | 112 (80.6%) | 103 (74.1%) | −6.5 pts; 11 lost, 2 gained; McNemar 95% CI −11.4 to −1.5 |
| Panel opened, over each run's own confirmed rows | 30/112 = 26.8% | 33/103 = 32.0% | +5.3 pts |
| Readable first-interaction surface, over each run's own confirmed rows | 19/112 = 17.0% | 20/103 = 19.4% | weekday − weekend = −2.4 pts, paired bootstrap 95% CI −8.7 to +3.7 (endpoints move ~0.1 pt with the resampling seed) |
This is NOT published as a bound on the staffing gap, for four independent reasons.
- Both endpoints are unstaffed windows.
agreement2enqueued at 05:41 CEST on a Tuesday — before the EU working day — and run 1 graded at 15:48–17:00 CEST on a Saturday. Sizing the gap to the registered 11:00 CEST slot requires at least one staffed endpoint, and neither of these is one. A contrast between two unstaffed hours cannot bound the distance from either to a staffed one. agreement2is out of band. Its enqueue at 03:41:57Z is more than three hours beforeBAND_UTC = (7, 15)opens — the band §5 pre-registers and every published sweep figure is filtered against. The Phase 5 review reports it surviving into analysis only via the missing-finished_atfail-open (the same fail-open §16c's band reconciliation had to check by hand to prove its zero was genuine). Promoting an out-of-band capture into a published bound, in a design whose band is pre-registered and whose §16c run record makes a point of reconciling every row against it, would trade one overclaim for a worse one. If it were ever published it would need an explicit, named band carve-out.- The sample is the agreement subsample, not the cohort. §5 specifies it as stratified across the four signal classes with
generic_launcher_onlydeliberately oversampled. That skew does not bias a within-site paired delta, but it does mean these are not cohort levels and the delta is estimated on 139 purposively chosen sites, not on 1,142. - The three measures move in three different directions. Presence falls 6.5 points, panel-open rate rises 5.3, readable rate is flat within noise. A staffing effect has a coherent signature — fewer live widgets and fewer panels opening and fewer readable surfaces. This is not that signature. The most likely reading is that the contrast is dominated by four days of ordinary site churn plus the §13a instrument noise the study has already measured, not by day of week.
One thing the pair does establish, and it is uncomfortable rather than convenient. The presence movement (−6.5 pts, 13 of 139 discordant = 9.4%) is larger than §13a's own test-retest floor of 3.6% discordance on widget presence. Whatever is driving it — churn, hour, day, or the residual instrument variance §13a called an upper bound — the study's only noise estimate does not cover it. That widens, not narrows, the uncertainty around any run1 → run2 movement, and it is recorded here for that reason. The deployed Worker also differs between the two captures (4af9624f → e9a8965e), though the §16g change was the stuck-scan reconciler's grace window and touches neither capture, interaction, nor detection.
What §18d/4's conclusion loses and keeps. It keeps everything that mattered: the ~27% fingerprint_strong complement is still a precision figure measured on a Monday and still cannot be repurposed as a weekend penalty, and both sides of §17c's argument are still unquantified priors. What it loses is the reason — "no such data exists" was a stronger, more checkable, and false claim, where "the data that exists cannot bound it" is true.
21c. The 41 "recipe stale" rows: this log read them backwards#
Review item #15. Corrected inline at §16c. Recorded here because a log contradicting the code it documents, inside the §5-gate discharge narrative, is worse than the row label itself.
interaction.js:664-670emits the 41 rows under the comment "the widget is there, the recipe is stale. A real maintenance-backlog signal, not a 'no widget' case."- The branch is reachable only when
matchedis truthy — a named vendor — and only after both the bespoke recipe selector and the generic salvage failed. All 41 rows carry the literal substring(recipe stale). All 41 have a named vendor. aggregate.py:126-127relabels this to "Vendor SDK fingerprinted, no launcher on the page" — an assertion about the site.- §16c then cited that relabelled row as "the only cause that names the false-positive class outright". The code says the opposite of what the log says, about the same 41 rows.
Corrected reading: the 41 rows are the clearest instrument-side bucket in the coverage table, not the clearest false-positive bucket. There is still no bucket in that table that names the admission false-positive class outright; that contamination sits in the 362-row "panel did not open" bucket and only §15b's pre-registered blind re-judging pass can size it.
Instrument-side subtotal: 120 rows = 15.2% of the 791 — not 151 (19.1%). The unambiguously instrument-side buckets are "panel opened but text unreadable" (79 — the corrected count, §21f) and these 41: 79 + 41 = 120. The 31 "panel selector matched only the closed launcher" rows are mixed — the selector mis-resolution is ours, but the failure to grow ≥1.5× after the click (interaction.js:732-733) is a fact about the site — so quoting 151 over-attributes to the instrument. Every failure bucket should be tagged instrument-side, site-side or mixed when the table is regenerated (§21f).
The published row label is not hand-patched: aggregate.py is a precommitment-bundle file and is frozen until run 2 (§21f). The relabelling ships in the same regeneration pass as the funnel fix.
21d. Unmet pre-registration requirements: the count is five, not two#
The published draft says "this is the second of two unmet pre-registration requirements on this page." A section titled "Five things this page cannot tell you" that publishes an undercount of its own misses is the first thing a competent reviewer leads with. The full list, as of 2026-08-02:
| # | Pre-registered | State | Where |
|---|---|---|---|
| 1 | Discovery→sweep agreement ≥ 85% overall (≥ 80% generic-launcher-only) | NOT MET — 80.7%, never read ≥ 85% on any run; discharged by re-examination under the gate's own release condition, not passed | §5, §10a, §13, §15c |
| 2 | Same-day full-cohort double sweep as the test-retest noise floor | NOT OBTAINED — run1-noise never enqueued; a pre-application floor is now permanently unobtainable | §5, §16d, §19f |
| 3 | Sweep egress country "probed and recorded before each run" | NEVER PERFORMED, for the baseline or at all — no record in this log, either PREFLIGHT file, or data/study-results.json | §4 (corrected), §21g |
| 4 | The vendor-split caveat §3 directs be stated with the vendor split | NOT STATED on the page | §3, below |
| 5 | "Could not assess (language)" outcome for undeclared/unsupported page language | DESIGNED, PRE-REGISTERED, NEVER IMPLEMENTED on the sweep path | §5 (corrected), §7 (corrected) |
On #3 (review item #12c). §4 pre-registers the probe in the same paragraph that offers the vantage disclosure. It was never run, so the study cannot say from which country Cloudflare Browser Rendering egressed on 1 August, and no run-1 field permits reconstructing it. It is NOT recoverable for run 2 either — probing it needs a deploy, which would abort both 8 August sweeps at Guard C (§21g/1). It is unmet for both runs; the post-freeze list carries it. Two adjacent claims in the same paragraph are also wrong and are corrected there: discovery ran from three networks, a majority of fetches from something other than the described residential line; and the sweep persists no redirect chain and no consent outcome, so there is no per-site consent record to audit for any of the 619 unreadable rows.
On #4 (review item #19). §3 does not merely note a consequence, it says where to state it: "Known consequence to state with the vendor split: a popularity-ranked frame over-represents enterprise sites relative to the long tail, so SMB-leaning vendors are under-represented versus their raw install base." The vendor section on the draft carries one caveat — the SDK-present-versus-reachable one — which is a different claim; "long tail", "over-repres" and "enterprise" all return zero matches on the rendered page. This lands on the most frame-sensitive number the study publishes: 72% unknown/custom is precisely what a rank-truncated frame inflates, because Tidio, Crisp, tawk.to and Smartsupp concentrate below the frame's rank floor of 164,793. The sentence is being added to the page verbatim.
Discrepancy between this log and the page — RESOLVED 2026-08-02, same day. The first revision of the published payload enumerated four ("one of four pre-registered requirements this page did not meet"), listing items 1–4 above, while this log said five. The difference was item 5, which the review raised as its own finding (#11) without folding it into its tally of misses. The page's count has been moved to five and item 5 added to its enumeration, so log and page now agree. Recorded rather than silently corrected because the sequence is the point: the undercount was itself the defect the review named — a section titled "Five things this page cannot tell you" publishing a short count of its own misses — and it survived one full fix iteration before being caught by the verification pass. That is the same failure shape as §18f/10: an error that survives review because each pass checks the thing it was asked about rather than the claim in front of it.
On #5 (review item #11). The review raised this without counting it in its own tally; it is counted here because §5 pre-registers it in the same voice as the other rules. Verified: detectDisclosure computes unclear from term match, persona and panelText length only (detectors.js:298-301, UNCLEAR_MIN_CHARS = 120); page_langs is read only by the pro-tier eu50.1c.language-coverage rule, which grades NA when !disclosure.found_at_first_interaction and therefore fires only after a disclosure was found; grep -rn "could not assess" src/ returns zero hits. The residual error runs in the direction §5 claims to protect against — a short greeting in an uncovered language publishes as "No disclosure detected". Like Irish, this is not being implemented inside a frozen instrument; it is described accurately and fixed after run 2.
21e. All three Bitcoin attestations predate the sweep — the log understated its own evidence#
Review item #9. Recorded inline at §19e. The correction runs in the study's favour and is made for that reason as much as any other.
| Attestation | Block height | Block time (UTC) | Relative to the 2026-08-01T13:48:13Z enqueue |
|---|---|---|---|
BitcoinBlockHeaderAttestation | 960456 | 2026-07-31T19:53:57Z | 17 h 54 m before |
BitcoinBlockHeaderAttestation | 960481 | 2026-07-31T23:51:35Z | 13 h 57 m before |
BitcoinBlockHeaderAttestation | 960506 | 2026-08-01T03:50:58Z | 9 h 57 m before |
The ots upgrade command ran at ~16:05 UTC on 1 August, after the sweep — that part of §18f/1 and §19e is accurate and stays. But an upgrade retrieves attestations already committed to blocks that were mined earlier; it does not create them and cannot move a committed time. So the correct statement is: the submission predates the sweep, the block confirmations predate the sweep, and only the retrieval command postdates it. The draft's line "Nothing here claims a confirmed blockchain timestamp predating the sweep" is therefore false in the direction of self-deprecation, and it is the wrong error to leave in the one paragraph written for a reader with a terminal. Publish the three heights and times so the check needs no repository.
21f. Two defects NOT fixed here, why, and exactly what they are#
Review items #2 and #6 are the only two of the nineteen that require editing code. Both are deferred until after run 2, and neither is deferred because it is in doubt. Both are stated here in full so that a reader who never sees the fix still has the corrected numbers. The working file that carries them to the far side of the freeze is tools/study/DEFERRED-FIXES-post-run2.md; this section is the durable record and the two must not be allowed to drift apart.
The reason for the deferral, in two parts.
aggregate.pyandprevalence-estimator.pyare precommitment-bundle files. Guard D on both 8 August scheduled sweeps (study-run2-enqueue-sat8aug,study-run2noise-enqueue-sat8aug) aborts on any drift againstprecommit-manifest-11. Editing either file before 8 August aborts both sweeps — losing the run1 → run2 delta and the substitute noise floor §20 was commissioned to obtain, which is the second time the same floor would have been lost.- Worse than the operational cost: editing the analysis code now would mean the run1 → run2 delta is computed by code chosen with knowledge of run 1. The whole point of anchoring the estimator before the data was to make that impossible. A correct fix applied at the wrong moment converts a fixable arithmetic defect into an unfixable pre-registration defect.
Both fixes land after run 2, and both pre-fix and post-fix numbers are published, side by side, with the defect described — not silently replaced.
#2 — the coverage funnel does not add up. 251 panels opened − 172 readable = 79, but the table at §16c prints 71, with 8 rows misfiled into "Other" (12). The 8 are userlike rows with panel_opened = 1 whose error string ("bespoke launcher selector stale … salvaged via generic fallback") matches no branch of aggregate.py:119-135, so they fall through to the residual bucket. Corrected: "panel opened but text unreadable" = 79 and "Other" = 4; the unreadable buckets still sum correctly to 619 either way, which is why no existing check caught it. Regeneration also moves two adjacent cells — an independent re-run gives 58 "another overlay" and 12 "consent layer" against the published 57 / 13. The fix is to match on interaction_succeeded rather than on error-string vocabulary and regenerate the table wholesale; hand-patching the one cell would leave the other two wrong. The failure-bucket instrument-side / site-side / mixed tagging (§21c) ships in the same pass. A printed funnel that fails subtraction needs no domain knowledge to spot, and this one is on the page.
#6 — the published "95% CI" is really about 99%. draw_rate (prevalence-estimator.py:122-131) resamples the cell trinomially and then draws Beta(k+½, n−k+½) from the resampled counts. Both layers are complete models of the same sampling variance, so they add. Confirmed by direct simulation rather than by reading: 200k draws give an SD ratio of 1.417 ≈ √2; Jeffreys-only gives 1.006, plain resampling 0.999, and the analytic delta-method SD (1.353 pp) matches the single-layer versions, not the published one. The defect is visible on the page without any code: the four component intervals are Wilson and correct, and the headline interval is √2 wider than they imply.
| Published | Correct | |
|---|---|---|
| Corrected prevalence (point estimate) | 11.3% | 11.3% — unaffected |
| Interval | 8.0–15.4 (labelled 95%, actually ≈99%) | 8.99–14.31 |
State the direction, because it runs against the author's interest. The corrected interval is narrower. The study's framing — that the apparent 9.36% sits inside the corrected interval, so the correction is a bias adjustment consistent with zero — therefore gets weaker, not stronger. 9.36% is still inside 8.99–14.31, so the framing survives; it survives by 0.37 points rather than by 1.36. That is the whole reason this defect gets its own paragraph instead of a line in an errata list: an author who found a defect that narrowed his own margin and deferred it for operational reasons owes the reader the number he would have published, in advance, in writing, before the fix.
In the interim, the published wording must disclose both defects rather than wait for them. The page prints 71 with a note that the arithmetic is wrong and the corrected value is 79, and prints the interval with a note that it is ≈99% rather than 95% and that the correctly-sized interval is 8.99–14.31. Disclosing a defect you have chosen not to fix yet is not the same as shipping it quietly.
BOTH FIXES APPLIED 2026-08-10, after run 2, per §26/3. #6 landed exactly as specified: the resample layer is gone, the Jeffreys draw reads the original decidable counts, and the re-run reproduces the prediction — point estimate 11.34% unchanged, interval 9.00–14.30 against the simulation's 8.99–14.31 (seed-level difference only). The pre-fix output is preserved at
tools/study/prevalence-estimate-prefix-21f6.json, and both intervals publish side by side. #2 landed as specified (bucket oninteraction_succeeded, wholesale regeneration, §21c attribution tags, instrument-side subtotal printed by the code); regenerating the retired run1 pull reproduces 79 panel-opened-unreadable, "Other" = 4, and the 120-row / 15.2% instrument-side subtotal exactly. One predicted cell did not reproduce: the shipped classifier yields 57 "another overlay" / 13 "consent layer" — the pre-fix published values — not the 58 / 12 this section predicted. The prediction came from the 2 August draft fix that was written and reverted the same day (Guard D) and cannot now be inspected. Both disagreeing candidates were re-read directly from the frozen run1 pull: one launcher was covered by a#privacy-overlaydiv, the other by a TrustCommandertc_privacyconsent-button container — both are consent tooling on inspection, so the shipped 13 is content-correct and the reverted draft's 12 appears to have been the misclassification. Recorded rather than tuned away: the code was not adjusted to match a number whose derivation no longer exists.
21g. Run-2 checklist — before Saturday 8 August#
Actions, not intentions. Each is either performed and recorded in §16e before run 2 fires, or recorded as not performed.
- Do NOT probe the sweep egress country before run 2 — it is NOT recoverable. An earlier draft of this item said it was, and following that instruction would have destroyed run 2. Observing Browser Rendering's egress requires a code change and therefore a deploy, and Guard C in both 8 August tasks is a strict version-equality check against
e9a8965e-282b-4735-be4b-ae0bbc0fff26— so any deploy (includingwrangler secret put, which mints a new Worker version) aborts both sweeps and takes the delta and the §20 substitute floor with them.request.cf.countryis inbound-only, and no stored column carries page text for a widget-less page, so there is no read-only route either. Record egress as unmet for BOTH runs and move the probe to the post-freeze list. Seetools/study/DEFERRED-FIXES-post-run2.md, which is canon for this and must not drift from this section again. - Test-fire both scheduled tasks, days ahead, out of band, from a cold start, to exhaustion (§19b). The Phase 5 review reports both
study-run2-enqueue-sat8augandstudy-run2noise-enqueue-sat8augas enabled with anextRunAtand nolastRunAtat all — i.e. neither has ever fired, not even as a test — so §19b's prescribed remedy is still outstanding, exactly as it was outstanding on 1 August when it cost the noise floor. Independently verified here, on the filesystem: bothSKILL.mdfiles still issuenpx wrangleras their first command (ad1 executeat line 26 and line 22 respectively),~/.claude/settings.jsoncontains four booleans and nopermissionsblock, and the project's.claude/holds onlylaunch.jsonandskills/. So the exact approval that stalled 1 August for 4 h 47 m is still unstored. An un-test-fired trigger is not a mitigation, and the page must not describe this remedy in the past tense. - Do not touch any bundle file before both sweeps complete.
aggregate.py,prevalence-estimator.py,noise-floor.py,study-run.py,recall-expansion-sample.mjs, the rule pack,cohort-final.csvandcandidates.csvare all Guard-D-checked. §21f's two fixes wait for this. - After run 2 completes: apply §21f's two fixes; regenerate the coverage funnel wholesale with instrument-side / site-side / mixed tagging (§21c); republish the prevalence CI with both pre-fix and post-fix intervals; add Irish to the lexicon and implement the "could not assess (language)" branch (§21a, §21d/5); run §15b's blind re-judging pass over the failure-bucket screenshots already in the evidence archive.
21h. Also recorded, not fixed here#
- §20's header forward-dates itself. §20 is titled "decided 2026-08-02" and §16e/1's SETTLED line carries the same date, but the commit that created both (
765d538) is authored and committed 2026-08-01T18:26:57Z with a clean tree. §20e's self-defence — that "before run 2 ran" is "the only claim it makes about its own timing" — is refuted by its own header. The decision itself is genuinely pre-run-2 and nothing about the substitute floor changes; what is wrong is a date. It is left in place and recorded, per this section's rule 1, and it becomes a must-fix the moment the repository is published, since from then on a reader can rungit logon it. - The rule pack's
report_footermisstates its own version (2026.07.8inside a2026.07.9pack) and twograding_notesentries still describe a reply-latency probe that no longer works as described. The pack is frozen; both are recorded as known defects and fixed after run 2. discovery-run1.ndjsonis not in version control — it is gitignored, so it is not merely unanchored but absent from any published tree. It feeds the frame denominator. Any claim that the discovery funnel is independently checkable depends on shipping this file.
22. §6 RECORDED DEPLOY (2026-08-02) — the freeze window ended before run 2#
§6 says an in-window deploy "must be recorded here". This is that record. It is written the evening it happened, before run 2, so it cannot be reconstructed favourably afterwards.
What happened. Stripe identity verification cleared on 2 August and the founder chose to complete the live-payments cutover the same evening rather than wait for run 2 — Article 50 applies 2 August, and being sellable on the day was judged worth more than an unbroken freeze. The trade was made with the cost stated in advance, not discovered afterwards.
The instrument of record therefore changes mid-study:
| Worker version | Rule pack | |
|---|---|---|
| run 1 (Sat 1 Aug) | e9a8965e-282b-4735-be4b-ae0bbc0fff26 | 2026.07.9 |
| run 2 (Sat 8 Aug) | 0df77510-2cea-4f84-93c4-100481848e6a | 2026.07.9 (unchanged) |
Six Worker versions exist between the two sweeps, and only one of them carries a code change. The full trail, so the version history is auditable rather than merely asserted: e8eff4d3 and 9f677260 (installing the live Stripe key and webhook secret), 53cd756d (the commerce deploy — commit 827fd32 plus the Managed Payments flag), then 0c1e8c3b, a52d55ef and 5f90d007 (rolling the live restricted key, three writes because the first two pasted the wrong clipboard entry), 2217b4e6 (the accounts change: registration, email verification, and account-bound purchases, plus migration 011), and finally c71fd465, a15b8e1f and 0df77510 (nav, copy and pricing corrections — no server logic). Every wrangler secret put mints a Worker version even though it changes no code; that is why the count is high and why it is worth stating plainly rather than leaving a reader to wonder what five undocumented deploys did.
The run-2 scheduled tasks' Guard C was re-pointed at 5f90d007 deliberately — left alone they would have aborted both sweeps and the study would have ended with run 1.
What changed, enumerated rather than characterised. Commit 827fd32 plus a MANAGED_PAYMENTS_ENABLED flip. Files: src/lib/stripe.js (Checkout Session parameters, TRIAL_PERIOD_DAYS), src/auth.js (post-checkout copy for an abandoned checkout), templates/{pricing,report,dashboard}.html (buy handlers and copy), test/billing.test.js, docs/FOUNDER_CHECKLIST.md, and in wrangler.jsonc: six Stripe price ids, DISABLE_CHECKOUT true→false, DISABLE_MONITORING true→false, MANAGED_PAYMENTS_ENABLED false→true.
What did NOT change — the reason run 2 is still worth sweeping. §6's enumerated instrument is capture, interaction (consent lists included), robots, fingerprints/recipes, and detectors. Every one of those is byte-identical across the two versions, as is the rule pack. The full 17-file precommitment bundle is untouched, which Guard D re-proves independently on sweep day rather than taking this paragraph's word for it. run 1 and run 2 are graded by the same grader.
What must be published regardless, and must not be softened. Run 2 did not execute under the same Worker build as run 1. Cite both version ids wherever the delta appears, describe the freeze as recorded and broken at a known point, never as unbroken, and do not let "the measurement code was identical" do the work of "the deployed artefact was identical". It was not.
One live-behaviour change with a possible sweep-day interaction. DISABLE_MONITORING="false" activates the hourly monitor sweep, which enqueues into the SAME scan-jobs queue this study uses. Verified 2026-08-02: 0 monitors and 0 paid users exist, so it is currently inert. If a subscriber starts monitoring before 8 August, a small amount of queue capacity is shared and the sweep-day drain should be watched rather than assumed.
23. The sweep egress country, measured at last (2026-08-03)#
§4 pre-registers that the Browser Rendering egress country be "probed and recorded before each run". §21d/3 records that it never was — the third unmet pre-registration. This section closes it for run 2 and states plainly what it cannot close.
The measurement. /api/_egress_probe (src/egress-probe.js, secret-gated, additive to the instrument and outside the precommitment bundle) launches the same Browser Rendering browser the sweep launches and asks two independent third parties what address they saw:
| Probed | Source | Country | Colo |
|---|---|---|---|
| 2026-08-03T16:45:56Z | cloudflare.com/cdn-cgi/trace | US | IAD |
| 2026-08-03T16:45:56Z | ifconfig.co/json | US | — |
Both agree: US. request.cf.country was never the answer to this question — that is the inbound visitor's country. The only honest measurement is what the browser looks like from outside, which is what this is.
What it means for the study, and it is not nothing. Discovery already ran from three US networks (§21, Comcast / T-Mobile CGNAT / Verizon Business). The sweep now measurably egresses from the US too. So both halves of the instrument observe EU-facing sites from a United States vantage — which is exactly the condition §4's vantage paragraph warns about: geo-dependent redirects, consent flows and widget configurations may differ from what a visitor inside the EU is served. That is no longer a hypothetical caveat about an unknown vantage; it is a measured property of this study, and the page should say so in those terms.
What it does not fix. Run 1's egress country is not recoverable. No run-1 row carries a field that would let anyone reconstruct it, and the sweep is done. A run-2-only reading cannot discharge the vantage limitation the pre-registration exists to make auditable — it stops the count of unmet requirements growing and lets the page name a country instead of a provider. Publish it with that scope attached.
Re-probe on sweep day. This reading is from 3 August, not run-2 day, and Cloudflare may egress from a different colo on 8 August. The pre-registration says before each run, so both run-2 tasks now run the probe in their preflight and record the result beside the Worker version. A reading that differs from US is not a failure — it is the finding, and it gets published as the pair.
24. §6 RECORDED DEPLOYS (2026-08-06) — the launch-week fix wave, and a bookkeeping failure stated plainly#
§22 recorded the freeze ending with a version trail "auditable rather than merely asserted." That standard was then not kept: after §22's trail closed at 0df77510, Guard C in the run-2 tasks was at some point re-pointed to 7a355e55 without a record here, and eleven further Worker versions shipped on 4–6 August without the pin moving or any entry being written. By the time this section was written (18:45 UTC, 6 Aug), 7a355e55 and everything before it had aged out of wrangler deployments list, so the id-to-commit mapping for the early fix wave is partially unrecoverable. That is a bookkeeping failure of exactly the kind §22 exists to prevent; it is recorded here as such, not smoothed over.
The trail as far as list retention allowed, captured 18:45 UTC 6 Aug:
| Created (UTC, 6 Aug) | Worker version | Attribution |
|---|---|---|
| 03:42 | b8a27134 | unrecorded at deploy time; not reconstructed |
| 16:11 | 2a42af5b | fix-wave deploys of that day's main commits (31ac0c0 … 2c606c8: dashboard verification values, entitled-report labeling, member pricing, paid-tier audit gaps, retention tier logic, crawl-deadline partial results, privacy/sitemap stamps, post-deploy review fixes). Per-version mapping was not recorded; fabricated precision would be worse than this stated gap. |
| 16:28 | be9fabdf | 〃 |
| 16:40 | 369af8c9 | 〃 |
| 16:55 | 06347cbe | 〃 |
| 17:33 | bf64ab64 | 〃 |
| 17:50 | 52a3cbac | 〃 |
| 17:54 | eee7aef4 | 〃 |
| 18:07 | 47dd94d0 | 〃 |
| 18:27 | 82c687f1 | git 2c606c8 + fd2ae2c (stale-PDF regeneration becomes render-then-swap); attested by its deploying session |
| 18:41 | 0c23d901 | git 8ad026b — merge of 5fb4a61 (report coverage caveats: crawl_error / sitemap_truncated / candidates_capped; 5s DoH abort deadline in ssrf.js), cc1064c (anonymous result cache never repoints at a deep scan id — founder decision), fd2ae2c. Live at 100% as of this record; Guard C in both run-2 tasks now pins it. |
The claim that matters, checkable rather than asserted. At 8ad026b, git log -1 on every file §6 enumerates — src/lib/capture.js, src/lib/interaction.js, src/lib/robots.js, src/lib/vendor-recipes.js, src/lib/detectors.js, src/detect.js, src/lib/lexicon.js, src/lib/eu-icons.js, src/grade.js, src/packs/ — shows a last modification of 2026-07-27 or earlier, i.e. before run 1. Whatever the per-version attribution gaps above, every version in this table grades with a byte-identical instrument and the identical rule pack 2026.07.9. Guard D independently re-proves the 14-file analysis bundle on sweep day.
One behaviour note a sweep reader deserves. 5fb4a61 added a 5-second abort to the DoH lookups in src/lib/ssrf.js target validation. Verdicts are unchanged except when a DoH query hangs longer than 5s: previously the scan hung on it (surfacing as an instrument-failure status), now it proceeds with the same fail-open verdict DoH errors always produced. The crawl/page-index changes in that commit are inert for study rows (single-page scans never enter the multi-page path).
What run 2's record must cite. run 1 = e9a8965e. run 2 = whatever deployments status reads on sweep day (0c23d901 if nothing else ships — §22's projected 0df77510 row is superseded by this section). Describe the freeze as recorded and broken at known points, never as unbroken. If anything deploys before the sweep: re-point Guard C in BOTH task files and extend this table the same hour — this section is the proof that waiting even a day makes part of the trail unrecoverable.
25. Run 2 missed 8 August and moves to Monday 10 August — a §5 weekday deviation, stated as one#
The 8 August attempt aborted without enqueueing. Full preflight record in tools/study/PREFLIGHT-2026-08-08.md. Guard A aborted on two independent clock reads (13:40:47Z and 13:42:50Z, both past the 13:00Z start cutoff). Guards C, D and E passed: 0c23d901 still 100%, all 14 non-doc bundle files byte-identical, egress US / colo MIA. Nothing was enqueued, so no scans were spent — cohort rows stand at 1,441 of the 6,000 wall.
The failure mode was NEW and the 1 August remedy does not cover it. Both tasks have lastRunAt 13:40:43Z against fireAt 09:00:00Z and 11:30:00Z — the same instant, neither at its own slot. On 1 August lastRunAt was 09:00:15Z, on time, and the loss came from a post-fire permission stall (§19a). Here the triggers never fired at all until the machine was switched on. Test-firing a task to pre-approve its tools cannot make it fire while the app is closed. Recorded because §19a's lesson ("an automated trigger inherits the permission model of the thing it automates") is now only half the story: it also inherits the uptime of the machine. The registered slot has been missed on four consecutive dates — 30 Jul, 31 Jul, 1 Aug (late), 8 Aug.
25a. The weekday change, and why it is a deviation rather than a correction#
Run 2 is rescheduled to Monday 10 August 2026, run2 at 11:00 UTC and run2-noise at 12:45 UTC. This departs from §5's "run 2 in the same weekday/hour band as run 1", and the departure is published as such — it is not compliance re-described.
The clause's referent no longer exists. §5 pins run 2 to run 1's weekday so the run 1 → run 2 delta stays comparable. All 1,098 run1 rows were retired to run1-superseded on 2 August on the founder's reaffirmed instruction (§20e). There is no run 1 and no delta, so there is nothing for the weekday to match. Saturday itself was never chosen on merit: it is where the baseline landed after Thursday 30 July and Friday 31 July were both missed outright (§17, §18), and §18d records its costs as accepted, not preferred.
Monday is the better measurement for what run 2 now is. With the delta gone, run 2's load-bearing job is the post-application levels snapshot. §18d/1 is explicit that levels are the weekend-sensitive output — roughly a quarter of fingerprint_strong admits ship the SDK disabled, login-gated, or out of staffed hours, and "out of staffed hours is exactly the category a Saturday inflates". §21b's paired weekday/weekend contrast points the same way (widget present 80.6% Tue vs 74.1% Sat over 139 sites graded in both, McNemar 95% CI −11.4 to −1.5), though that pair remains not a bound on the staffing gap for the four reasons §21b gives, and must not be cited as one here either.
What this costs, stated plainly. Any future comparison between the retired run1-superseded rows and run 2 is now confounded by weekday as well as by hour of day (§20c), and in the same direction: a staffed Monday morning should show more live widgets and more open panels than a late Saturday afternoon even if nothing changed. Since the retired baseline is not a publishable comparator anyway, this is recorded to prevent the retired rows being quietly promoted back into a delta later. If they ever are, both confounds must travel with them.
25b. Spacing trimmed to 1 h 45 m, and why not 2 h#
The pair runs 11:00 UTC / 12:45 UTC. A 2 h spacing off an 11:00 start puts the second sweep at 13:00 UTC, which trips Guard A's own hour >= 13 cutoff and would abort the floor a third time. The cutoff was left alone rather than raised: it is derived from the 15:00 cliff minus a pessimistic ~110 min sweep, so a 13:30 start finishes at 15:20 and out of band. Spacing is not pre-registered — §5 requires only that both sweeps fall in the same day and hour band — and its operational job is to avoid queue contention at max_concurrency: 10. run2 should be done by ~12:15 UTC, so 1 h 45 m serves that as well as 2 h would.
25c. Guard B's stale run1 expectation, fixed in both task files#
Both run-2 task files told the operator to expect run1 at 1,098 rows and to treat its absence as "something is wrong with the baseline". That text predates the 2 August retirement and was never updated, so on any day Guard A had not already aborted it would have raised a false alarm against a state the founder deliberately created. Both rewritten files now expect run1-superseded ~1,098, state that the absence of run1 is correct, and abort only on existing run2 / run2-noise rows. Logged as an instance of the §18f/1 and §21b class — a sentence that was true when written and aged into an error because nothing re-read it.
26. run2-noise did not run; the floor is dropped (founder decision, 2026-08-10)#
What happened. run2 executed cleanly (§16e). Its pair, run2-noise, fired on time at 12:45:08 UTC, passed all five preflight guards — band (twice, host clock corroborated off-host), interlock, instrument pin, byte-identical bundle, egress — and then died at Step 2 before any network call: study-run.py:29 hardcodes ALLOWED_COHORTS = {"run1", "run1-noise", "run2", "smoke"} and rejects run2-noise at argparse. The label was coined in the scheduling docs when the floor moved from baseline day to run-2 day (§20); the driver is one of the 14 frozen bundle files, so it could never have been taught the name. Zero rows, zero side effects — full record tools/study/PREFLIGHT-2026-08-10-run2noise-ABORT.md. Both in-run escapes were correctly refused: editing the frozen driver would have broken the same-code identity that is the floor's entire point, and writing Monday's sweep under the run1-noise label would have counterfeited the pre-registered object this log records as permanently unobtainable (§16d).
By decision time the same-day pair was already unobtainable. Guard A's 13:00 UTC enqueue cutoff passed with run2-noise un-enqueued, and a sweep started after it could not drain ~73 minutes (§16e) ahead of the 15:00 cliff without losing the Tranco-ordered tail — the §17b/§19a non-random loss. §20a defines the substitute as a same-day pair with run2, so no later action can produce it. The choices that remained were a different object — a later-day second sweep under a new label, which per §16d measures instrument self-disagreement plus one or more days of real site churn (and §21b already measured 4-day churn at −6.5 points presence, larger than §13a's 3.6-point discordance, so churn would dominate what such a sweep returned) — or no floor at all.
The decision: drop the floor. The founder, 10 August 2026, on being shown the abort record, the closed window, and both remaining options: no further noise sweep will be run. Consequences, each already prescribed elsewhere and executed here rather than re-decided:
- §20 is marked not-obtained at its heading, exactly as §20d's final bullet prescribes and exactly as §16d was. §19f's comparator rule revives unchanged: §13a's cross-code bound (~4 points presence, ~8 points readable, 137 sites, Monday 27 July, instrument-confounded) is the study's only noise estimate, never called "the noise floor" unqualified, its three limits named in the same sentence. In practice the rule is currently moot — run1 is retired (§25a), so there is no delta to judge — but it binds any future citation of the retired rows.
- The published piece states plainly that no test-retest noise floor exists — neither the pre-registered pre-application object (§16d) nor the §20 substitute — and why, with pointers to both failure records. The §15 headlines are levels and are unaffected; run2 publishes as a single post-application snapshot (§16e).
- The §6/Guard-D deferral rationale is discharged. §21g/3 forbade touching the bundle "before both sweeps complete"; with run2 complete and no further sweep ever to run under this instrument, the deferred fixes (§21f #2 and #6, SF-3, and the lexicon's missing Irish) are unblocked and applied as dated amendments — pre-fix and post-fix values published side by side, per §21f. Study teardown (harness route,
STUDY_SECRET, scheduled tasks) is likewise authorized; cohort rows stay, as the audit trail. - There is still no anchor 12 (§19e's reasoning is unchanged by any of this), and this section, like §20, is a decision record written after data existed — it pre-registers nothing and says so.
The lesson, one sentence wide. Five guards checked the world — clock, database, deployed Worker, file hashes, egress — and none checked whether the enqueue command accepted its own argument; grep run2-noise tools/study/study-run.py at any point in the eight days the task sat scheduled would have caught it off-band for free. Rehearse the check, don't read it (§18's meta-lesson, now with a second scalp).
27. The §15b blind re-judging, run at last — two-thirds of the largest failure bucket shows no launcher at all (2026-08-10)#
§15b's fourth bullet pre-registered it and §21 called not running it the study's weakest point: the "panel did not open after the click" bucket carries the admission-false-positive contamination, and only a blind re-judging of the already-captured screenshots can size it. It has now been run — against run 2's bucket, since run 1 is retired and run 2 is the published table.
Protocol (the §12/§14 record format and conventions, applied to the sweep's own captures — one stated difference: §12/§14 judged fresh audit-day re-captures, this judged the sweep-time captures themselves, which is the right ground truth for auditing what the sweep saw but is not byte-identical method): all 361 run2 rows bucketed "Panel did not open after the click" — the corrected interaction_succeeded-keyed bucket, §21f #2 — had their sweep-captured above-fold screenshots (R2 screenshot_abovefold, one per graded scan) pulled and judged blind. Judges saw only the image; filenames were scan ids, so no machine metadata — domain, vendor, error string, admission class — was available. Blindness is to metadata and to bucket membership, not to site identity where the page itself displays it: a logo in the screenshot was visible to the judge, as it was in every prior audit. Recording format: launcher_visible yes/no/unclear, page_state page_visible / consent_wall_blocking / error_or_blank / bot_challenge, non-rendered states normalizing to unclear per the estimator — including two rows judged yes on error/placeholder pages, which the uniform rule counts as unclear; they are flagged per-row (normalized_rows) and a sensitivity rate that skips the normalization is published below. Fourteen independent judges, per-row judge index recorded; ~27 files each; 20 files planted as duplicates across two different judges. Dup disagreements resolve to unclear with both raw readings kept on the row. One judgement came back with a one-character filename transcription error, repaired to the unique matching id and recorded in the file's repairs field. Full record: tools/study/run2-panelfail-blind-judgements.json.
Results.
| Bucket | 361 rows (of run2's 620 unreadable, of run2's 794 confirmed) |
| Decidable | 260. The 101 counted-unclear decompose by page_state as 66 page_visible (dominated, per the notes, by full-width cookie bars and centred consent popups covering the corners where launchers sit), 31 consent_wall_blocking, 2 error/blank, 2 bot challenges. Consent walls do not auto-exclude: 34 further consent-wall rows were decided (25 no, 9 yes) and sit inside the 260 |
| No real launcher visible in the capture | 176 / 260 = 67.7% (Wilson 95% 61.8–73.1) |
| Sensitivity, no yes-on-error normalization | 176 / 262 = 67.2% (61.3–72.6) |
| Launcher corroborated in the capture | 84 / 260 = 32.3% (26.9–38.2) |
| Dup reliability | launcher_visible 19 / 20; page_state 16 / 20 (the four splits are consent-wall-vs-page_visible boundary calls, all recorded); the one launcher split is no-vs-unclear, resolved to unclear |
Reading it. Two-thirds of the decidable rows in the study's largest failure bucket show no chat launcher a blind judge could see in the sweep's own above-fold capture — the judges' notes name the §12 false-positive classes over and over: accessibility widgets, carousel arrows, trust badges, comment-count icons, contact links. This corroborates §12's precision audit (46.9% on generic admits) from an independent direction, and it means the coverage headline's denominator contamination is not hypothetical: a large share of "panel did not open" is "there was never a panel to open". The direction is the one the draft page already names — correcting the denominator would move the coverage figure down.
What it does and does not license. The measured rate publishes on the page beside the coverage table, with its CI and the unclear share (the new panelfail_audit_note field). The pre-registered follow-up's second half — republishing coverage on an audit-corrected basis — is an estimator design change and goes through adversarial review before any re-based headline is printed; the raw headline stays raw and labelled raw until then. Two limitations travel with the rate: the above-fold capture is the same ground truth the §12/§14 audits used, so a launcher below the fold or rendered after the screenshot reads as "no" (biasing the rate up by an unmeasured amount); and the 101 unclears are dominated by consent walls, the §15b-noted non-neutral exclusion — both stated wherever the number appears.
Cite as: DisclosureProof, “The State of AI Disclosure 2026 — method log”. Licensed CC BY 4.0, like the aggregate tables on the study page. Publishing this log changes nothing about the data policy: no dataset, no repository snapshot and no per-site record is published, and per-site records never will be. Questions and corrections: hello@disclosureproof.com.