docs(taxonomy): #600 link census reproducibility — independent re-run byte-identical - #606
Conversation
… byte-identical Re-run the full link_* langlinks census (741 Fallacies + 88 Virtues articles, probe `0 fallacies` then `0 virtues`) on master c20d5d2 — an independent confirmation of the §5.1 decision-grade figures measured 2026-06-25. Result: byte-identical. Fallacies 2739/4823 (57%), Virtues 180/322 (56%), combined ~2919, 0/829 errors, and the same link_en categorization (Fallacies 900/433/75, Virtues 185/9/29). Per-language rows match cell-for-cell. Adds a one-paragraph "Reproducibility — re-confirmed 2026-06-27" note under §5.1 stating the figures are stable and not a one-run artefact. The "named number replacing the 57% estimate" was already the measured value on master (dispatch was prepared 1 min before #600 merged §5.1 at 14:40); this PR's contribution is the second-run pin that makes it reproducibly-confirmed. Read-only. 0 write under Cards/. Dispatch ai-01 2026-06-27 tertiaire (#600 census link_* complet). Co-Authored-By: Claude-Code <noreply@anthropic.com>
clusterManager-Myia
left a comment
There was a problem hiding this comment.
[NanoClaw]
docs(taxonomy): #600 link census reproducibility — independent re-run byte-identical relu (+2/-0, 1 fichier markdown, Argumentum base master). Ajout d'un paragraphe de confirmation de reproductibilité : un re-run indépendant (po-2024, master c20d5d2c) du census langlinks #600 retourne des figures byte-identical.
Vérification interne faite : les chiffres du nouveau paragraphe (Fallacies 2 739 / 4 823 = 57 % ; Virtues 180 / 322 = 56 % ; catégorisation Fallacies 900 Wikipédia · 433 non-Wikipédia · 75 vides ; Virtues 185 · 9 · 29) matchent les figures §5.1 pré-existantes du même doc (vérifié head : 2 739 / 4 823 / 900 / 433 / 75 / 180 / 322 / 185 présents dans les sections de mesure d'origine). Cohérence interne solide → la claim « byte-identical » est auto-cohérente. Le probe 192-link-coverage-langlinks-probe.py existe bien (7 555 bytes, read-only, no API key — MediaWiki public).
Note de transparence : je vérifie la cohérence interne + l'existence du script, pas un fresh re-run indépendant de ma main (le census MediaWiki 741+88 articles n'est pas reproductible dans ce tour). La claim de reproductibilité est crédible et bien documentée (auteur po-2024, master SHA, date, conditions de throttle).
Additif pur (0 suppression), release-freeze respecté (fichier docs/taxonomy/, pas Cards/), 0 concern sécurité. LGTM-with-artifact.
[COMMENT only — validation humaine]
…sal (#609) Grounds #141 in the CURRENT taxonomy state (issue text is ~2 years old; DatasetUpdater + translation PRs have moved the state considerably). Census finding (read-only script on ba8e4a6, reproducible like #600/#606): - Non-card nodes: Fallacies 1232, Virtues 110 (Scenarii/Rules out of scope). - TEXT enrichment (desc/example/title x 8 langs) = 100% on both datasets => original #141 scope item 3 (descriptions + examples + translations) is DONE. No text-enrichment gap to script. - The "GPT-4 enrichment script" = already adapted: modern DatasetUpdater (PR #210, OpenAI SDK v2.10.0, gpt-5.5). Not a rewrite. Reframes #141: the genuinely-open residue is the AIF cross-reference graph (crossLink_* 8 relationship cols + AIF_skos* 4 SKOS mappings). Schema is READY (columns exist), content ~0% on non-cards: - Fallacies: crossLink 0.2%, AIF_skos 0.5% (barely started) - Virtues: crossLink_Opposes + skosDirectRef/MappingType 100% via #498 pilots, but 7 OTHER crossLink verbs = 0% - ~16700 empty cross-ref cells where schema is waiting. Proposal: 4-stage method (ground -> gpt-5.5 candidate via /v1/responses effort=low -> drift-free sidecar dry-run -> expert/jsboige ratification gate -> OWL/2sxc export). NOT auto-write: AIF/Walton mappings are specialised (memory: "Anti-Fab Validator: Walton scheme = WARN"), and Cards/ is frozen pre-tag. "Adapt the script" = add a crossLink prompt + task config to DatasetUpdater, GATED post-release on jsboige GO + schema decision (single decimal_path vs structured {target,note}). Overlaps #498 (AIF scale-up pilots produced the few filled cells). Feeds dependencies #141 declares: #130 (OWL/SKOS) + #136 (2sxc). Scope: docs + read-only script only. 0 write Cards/, 0 AssetConverter code change (pre-tag safe). Base ba8e4a6. Relates to #130, #136, #498. Co-authored-by: Your <your.email@example.com> Co-authored-by: Claude-Code <noreply@anthropic.com>
…or, 0 Cards/) (#610) * docs(taxonomy): #141 non-card census + AIF cross-ref enrichment proposal Grounds #141 in the CURRENT taxonomy state (issue text is ~2 years old; DatasetUpdater + translation PRs have moved the state considerably). Census finding (read-only script on ba8e4a6, reproducible like #600/#606): - Non-card nodes: Fallacies 1232, Virtues 110 (Scenarii/Rules out of scope). - TEXT enrichment (desc/example/title x 8 langs) = 100% on both datasets => original #141 scope item 3 (descriptions + examples + translations) is DONE. No text-enrichment gap to script. - The "GPT-4 enrichment script" = already adapted: modern DatasetUpdater (PR #210, OpenAI SDK v2.10.0, gpt-5.5). Not a rewrite. Reframes #141: the genuinely-open residue is the AIF cross-reference graph (crossLink_* 8 relationship cols + AIF_skos* 4 SKOS mappings). Schema is READY (columns exist), content ~0% on non-cards: - Fallacies: crossLink 0.2%, AIF_skos 0.5% (barely started) - Virtues: crossLink_Opposes + skosDirectRef/MappingType 100% via #498 pilots, but 7 OTHER crossLink verbs = 0% - ~16700 empty cross-ref cells where schema is waiting. Proposal: 4-stage method (ground -> gpt-5.5 candidate via /v1/responses effort=low -> drift-free sidecar dry-run -> expert/jsboige ratification gate -> OWL/2sxc export). NOT auto-write: AIF/Walton mappings are specialised (memory: "Anti-Fab Validator: Walton scheme = WARN"), and Cards/ is frozen pre-tag. "Adapt the script" = add a crossLink prompt + task config to DatasetUpdater, GATED post-release on jsboige GO + schema decision (single decimal_path vs structured {target,note}). Overlaps #498 (AIF scale-up pilots produced the few filled cells). Feeds dependencies #141 declares: #130 (OWL/SKOS) + #136 (2sxc). Scope: docs + read-only script only. 0 write Cards/, 0 AssetConverter code change (pre-tag safe). Base ba8e4a6. Relates to #130, #136, #498. Co-Authored-By: Claude-Code <noreply@anthropic.com> * tools(i18n): link_* langlinks resolver (sidecar candidate-URL generator, 0 Cards/) SECONDAIRE of ai-01 deep-queue dispatch. Extends the #600/#606 coverage track from MEASUREMENT to RESOLUTION. The existing probe (docs/taxonomy/192-link-coverage-langlinks-probe.py) measures the resolvable ceiling (~2919 cells, 57%, #600 §5.1) but discards the target title (returns only lang codes). This tool captures the target-language title via the MediaWiki langlinks API and emits the candidate fill URLs as a sidecar report — step 1 of #600 §6 methodology. For every node with an en.wikipedia.org/wiki/<Title> link_en missing link_<lang>, queries langlinks, captures the target title, emits: dataset,key,link_lang,resolved_url to stdout or --out <path>. NEVER writes under Cards/. Verified on a strided sample of 10 fallacies: 0 errors, 37 candidate fills (~57% rate, consistent with the probe ceiling), URLs correctly URL-encoded (Cyrillic %D0, PT accented %C3). Sample also surfaced the homonym risk #600 §6.4 warns about (English "Engagement" for a fallacy node) -> confirms human spot-validation is non-optional for AR/FA/ZH. Safety: - 0 write Cards/ (sidecar only, pre-tag freeze). - Public MediaWiki API, no key, 0.3s throttle, descriptive UA (default urllib UA is 403-forbidden). - Resolves FROM link_en (Wikipedia only); 433 non-Wikipedia curated sources excluded + preserved as-is (#600 §6.2). Next (post-release, gated): follow-up PR consumes the sidecar -> apply candidate URLs cell-by-cell (drift-free QUOTE_MINIMAL+CRLF, method #595), skip non-empty cells, human spot-validate ~5% residue (AR/FA/ZH, ~150). Relates to #600, #606. Co-Authored-By: Claude-Code <noreply@anthropic.com> --------- Co-authored-by: Your <your.email@example.com> Co-authored-by: Claude-Code <noreply@anthropic.com>
… (0 errors, 0 Cards/) (#618) PRIMAIRE of ai-01 deep-queue #2 (dispatch msg-...92kynx). Produces the candidate-URL output behind the #600 §5.1 probe ceiling. Ran tools/link-langlinks-resolve.py (PR #610, merged) over full Fallacies + Virtues taxonomy: 829 articles probed -> 2934 candidate link_<lang> URLs (fallacies 2754, virtues 180), 0 errors. Materializes the ~57% probe ceiling (~2919 cells) into concrete reviewable URLs. Step 1 of #600 §6 methodology. Per-lang: ru 371, pt 430, es 513, ar 603, fa 480, zh 537. Resolves FROM link_en (Wikipedia only); 433 non-Wikipedia curated sources excluded + preserved (#600 §6.2). Quality verified: - URL encoding correct per script (Cyrillic %D0, Arabic %D8, Persian %D8, CJK %E8, PT/ES Latin unencoded). - Homonym scan (#600 §6.4) of 1620 AR/FA/ZH candidates: 0 real leaks. AR 0/603, FA 0/480 Latin-path; ZH 3/537 (0.6%) = legit loanwords/ acronyms (FUD x2, Creepypasta) — not errors. The "Engagement" homonym from PR #610 sample does NOT recur at scale (isolated case). - Asymmetry: virtues ru/pt = 4 each (sparse abstract-concept coverage) vs ar/fa/zh 46-56 — coverage reality, flagged for spot-validation. Sidecars (reviewable snapshot, UTF-8 no-BOM LF): - docs/taxonomy/600-link-resolve-fallacies.csv (2754 rows) - docs/taxonomy/600-link-resolve-virtues.csv (180 rows) schema: dataset,key,link_lang,resolved_url Robustness: try/except + continue (skip on network error, never abort) -> 0 errors/829. Public MediaWiki API, no key, 0.3s throttle, descriptive UA. Next (gated post-release): apply PR consumes sidecar -> cell-by-cell drift-free (QUOTE_MINIMAL+CRLF, method #595), skip non-empty, human spot-validate ~5% AR/FA/ZH residue (~150 cells), re-run probe to confirm gap closed. Scope: docs/taxonomy/ only. 0 write Cards/, 0 AssetConverter code change (pre-tag safe). Base 18b4d02. Relates to #600, #606, #610. Co-authored-by: Your <your.email@example.com> Co-authored-by: Claude-Code <noreply@anthropic.com>
Summary
An independent reproducibility re-run of the
link_*langlinks census (#600, merged32dd809b), confirming the §5.1 decision-grade figures are stable — not a one-run artefact or mislabeled sample.Re-run (2026-06-27, master
c20d5d2c)Full probe, both datasets (
0 fallaciesthen0 virtues), 741 + 88 = 829 unique articles, ~0.3s throttle, 0 API key.link_encategorization also matches: Fallacies 900 Wikipedia · 433 non-Wikipedia · 75 empty; Virtues 185 · 9 · 29. Per-language rows match cell-for-cell (ru/pt/es/ar/fa/zh).What changed
One paragraph added under §5.1 — a "Reproducibility — re-confirmed 2026-06-27" note stating the figures are stable and decision-grade. No table/code changes; the data is the data.
Why this matters
The 2919 figure feeds jsboige's WE priority decision (how bounded the post-release link-fill effort is). A single-run measurement is a claim; a re-run that returns identical numbers is a reproducibly-confirmed claim — the same epistemic bar the characterization test (PR #604) sets for the Virtues FR-frozen behaviour.
Scope
Cards/(read-only docs, release freeze respected)c20d5d2c, 1 file, +2 linesRelates to #600. Does not change its conclusion — confirms it.