Skip to content

docs(taxonomy): #600 link census reproducibility — independent re-run byte-identical - #606

Merged
jsboige merged 1 commit into
masterfrom
docs/600-link-census-reproducibility-confirm
Jun 28, 2026
Merged

docs(taxonomy): #600 link census reproducibility — independent re-run byte-identical#606
jsboige merged 1 commit into
masterfrom
docs/600-link-census-reproducibility-confirm

Conversation

@jsboige

@jsboige jsboige commented Jun 27, 2026

Copy link
Copy Markdown
Contributor

Summary

An independent reproducibility re-run of the link_* langlinks census (#600, merged 32dd809b), confirming the §5.1 decision-grade figures are stable — not a one-run artefact or mislabeled sample.

Context: dispatch ai-01 2026-06-27 (tertiaire) asked to "replace the ~57% estimate with a named number via a full census." Reading the merged #600 doc revealed the named number was already on master — §5.1 (measured census, 2026-06-25) sits one minute after the dispatch was written (which referenced the older §5 theoretical ceiling of 8110). This PR therefore contributes what's genuinely missing: a second independent run that pins the figure as reproducibly-confirmed.

Re-run (2026-06-27, master c20d5d2c)

Full probe, both datasets (0 fallacies then 0 virtues), 741 + 88 = 829 unique articles, ~0.3s throttle, 0 API key.

Dataset §5.1 (2026-06-25) Re-run (2026-06-27) Errors Match
Fallacies 2739 / 4823 (57%) 2739 / 4823 (57%) 0/741 ✅ byte-identical
Virtues 180 / 322 (56%) 180 / 322 (56%) 0/88 ✅ byte-identical
Combined ~2919 ~2919

link_en categorization also matches: Fallacies 900 Wikipedia · 433 non-Wikipedia · 75 empty; Virtues 185 · 9 · 29. Per-language rows match cell-for-cell (ru/pt/es/ar/fa/zh).

What changed

One paragraph added under §5.1 — a "Reproducibility — re-confirmed 2026-06-27" note stating the figures are stable and decision-grade. No table/code changes; the data is the data.

Why this matters

The 2919 figure feeds jsboige's WE priority decision (how bounded the post-release link-fill effort is). A single-run measurement is a claim; a re-run that returns identical numbers is a reproducibly-confirmed claim — the same epistemic bar the characterization test (PR #604) sets for the Virtues FR-frozen behaviour.

Scope

  • 0 write under Cards/ (read-only docs, release freeze respected)
  • ✅ Base c20d5d2c, 1 file, +2 lines
  • Dispatch ai-01 2026-06-27, tertiaire (docs(i18n): #192 link_* coverage research track (own the URL gap) #600 census link_* complet)
  • Pre-existing MD060 markdown-lint table-style warnings in the file are out of scope (cosmetic, on tables I did not touch)

Relates to #600. Does not change its conclusion — confirms it.

… byte-identical

Re-run the full link_* langlinks census (741 Fallacies + 88 Virtues articles,
probe `0 fallacies` then `0 virtues`) on master c20d5d2 — an independent
confirmation of the §5.1 decision-grade figures measured 2026-06-25.

Result: byte-identical. Fallacies 2739/4823 (57%), Virtues 180/322 (56%),
combined ~2919, 0/829 errors, and the same link_en categorization
(Fallacies 900/433/75, Virtues 185/9/29). Per-language rows match cell-for-cell.

Adds a one-paragraph "Reproducibility — re-confirmed 2026-06-27" note under
§5.1 stating the figures are stable and not a one-run artefact. The "named
number replacing the 57% estimate" was already the measured value on master
(dispatch was prepared 1 min before #600 merged §5.1 at 14:40); this PR's
contribution is the second-run pin that makes it reproducibly-confirmed.

Read-only. 0 write under Cards/. Dispatch ai-01 2026-06-27 tertiaire (#600
census link_* complet).

Co-Authored-By: Claude-Code <noreply@anthropic.com>

@clusterManager-Myia clusterManager-Myia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[NanoClaw]

docs(taxonomy): #600 link census reproducibility — independent re-run byte-identical relu (+2/-0, 1 fichier markdown, Argumentum base master). Ajout d'un paragraphe de confirmation de reproductibilité : un re-run indépendant (po-2024, master c20d5d2c) du census langlinks #600 retourne des figures byte-identical.

Vérification interne faite : les chiffres du nouveau paragraphe (Fallacies 2 739 / 4 823 = 57 % ; Virtues 180 / 322 = 56 % ; catégorisation Fallacies 900 Wikipédia · 433 non-Wikipédia · 75 vides ; Virtues 185 · 9 · 29) matchent les figures §5.1 pré-existantes du même doc (vérifié head : 2 739 / 4 823 / 900 / 433 / 75 / 180 / 322 / 185 présents dans les sections de mesure d'origine). Cohérence interne solide → la claim « byte-identical » est auto-cohérente. Le probe 192-link-coverage-langlinks-probe.py existe bien (7 555 bytes, read-only, no API key — MediaWiki public).

Note de transparence : je vérifie la cohérence interne + l'existence du script, pas un fresh re-run indépendant de ma main (le census MediaWiki 741+88 articles n'est pas reproductible dans ce tour). La claim de reproductibilité est crédible et bien documentée (auteur po-2024, master SHA, date, conditions de throttle).

Additif pur (0 suppression), release-freeze respecté (fichier docs/taxonomy/, pas Cards/), 0 concern sécurité. LGTM-with-artifact.

[COMMENT only — validation humaine]

@jsboige
jsboige merged commit 28f70e6 into master Jun 28, 2026
3 checks passed
@jsboige
jsboige deleted the docs/600-link-census-reproducibility-confirm branch June 28, 2026 11:53
jsboige added a commit that referenced this pull request Jun 29, 2026
…sal (#609)

Grounds #141 in the CURRENT taxonomy state (issue text is ~2 years old;
DatasetUpdater + translation PRs have moved the state considerably).

Census finding (read-only script on ba8e4a6, reproducible like #600/#606):
  - Non-card nodes: Fallacies 1232, Virtues 110 (Scenarii/Rules out of scope).
  - TEXT enrichment (desc/example/title x 8 langs) = 100% on both datasets
    => original #141 scope item 3 (descriptions + examples + translations)
    is DONE. No text-enrichment gap to script.
  - The "GPT-4 enrichment script" = already adapted: modern DatasetUpdater
    (PR #210, OpenAI SDK v2.10.0, gpt-5.5). Not a rewrite.

Reframes #141: the genuinely-open residue is the AIF cross-reference graph
(crossLink_* 8 relationship cols + AIF_skos* 4 SKOS mappings). Schema is
READY (columns exist), content ~0% on non-cards:
  - Fallacies: crossLink 0.2%, AIF_skos 0.5% (barely started)
  - Virtues: crossLink_Opposes + skosDirectRef/MappingType 100% via #498
    pilots, but 7 OTHER crossLink verbs = 0%
  - ~16700 empty cross-ref cells where schema is waiting.

Proposal: 4-stage method (ground -> gpt-5.5 candidate via /v1/responses
effort=low -> drift-free sidecar dry-run -> expert/jsboige ratification
gate -> OWL/2sxc export). NOT auto-write: AIF/Walton mappings are
specialised (memory: "Anti-Fab Validator: Walton scheme = WARN"), and
Cards/ is frozen pre-tag. "Adapt the script" = add a crossLink prompt +
task config to DatasetUpdater, GATED post-release on jsboige GO + schema
decision (single decimal_path vs structured {target,note}).

Overlaps #498 (AIF scale-up pilots produced the few filled cells).
Feeds dependencies #141 declares: #130 (OWL/SKOS) + #136 (2sxc).

Scope: docs + read-only script only. 0 write Cards/, 0 AssetConverter code
change (pre-tag safe). Base ba8e4a6.

Relates to #130, #136, #498.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 29, 2026
…or, 0 Cards/) (#610)

* docs(taxonomy): #141 non-card census + AIF cross-ref enrichment proposal

Grounds #141 in the CURRENT taxonomy state (issue text is ~2 years old;
DatasetUpdater + translation PRs have moved the state considerably).

Census finding (read-only script on ba8e4a6, reproducible like #600/#606):
  - Non-card nodes: Fallacies 1232, Virtues 110 (Scenarii/Rules out of scope).
  - TEXT enrichment (desc/example/title x 8 langs) = 100% on both datasets
    => original #141 scope item 3 (descriptions + examples + translations)
    is DONE. No text-enrichment gap to script.
  - The "GPT-4 enrichment script" = already adapted: modern DatasetUpdater
    (PR #210, OpenAI SDK v2.10.0, gpt-5.5). Not a rewrite.

Reframes #141: the genuinely-open residue is the AIF cross-reference graph
(crossLink_* 8 relationship cols + AIF_skos* 4 SKOS mappings). Schema is
READY (columns exist), content ~0% on non-cards:
  - Fallacies: crossLink 0.2%, AIF_skos 0.5% (barely started)
  - Virtues: crossLink_Opposes + skosDirectRef/MappingType 100% via #498
    pilots, but 7 OTHER crossLink verbs = 0%
  - ~16700 empty cross-ref cells where schema is waiting.

Proposal: 4-stage method (ground -> gpt-5.5 candidate via /v1/responses
effort=low -> drift-free sidecar dry-run -> expert/jsboige ratification
gate -> OWL/2sxc export). NOT auto-write: AIF/Walton mappings are
specialised (memory: "Anti-Fab Validator: Walton scheme = WARN"), and
Cards/ is frozen pre-tag. "Adapt the script" = add a crossLink prompt +
task config to DatasetUpdater, GATED post-release on jsboige GO + schema
decision (single decimal_path vs structured {target,note}).

Overlaps #498 (AIF scale-up pilots produced the few filled cells).
Feeds dependencies #141 declares: #130 (OWL/SKOS) + #136 (2sxc).

Scope: docs + read-only script only. 0 write Cards/, 0 AssetConverter code
change (pre-tag safe). Base ba8e4a6.

Relates to #130, #136, #498.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

* tools(i18n): link_* langlinks resolver (sidecar candidate-URL generator, 0 Cards/)

SECONDAIRE of ai-01 deep-queue dispatch. Extends the #600/#606 coverage
track from MEASUREMENT to RESOLUTION.

The existing probe (docs/taxonomy/192-link-coverage-langlinks-probe.py)
measures the resolvable ceiling (~2919 cells, 57%, #600 §5.1) but
discards the target title (returns only lang codes). This tool captures
the target-language title via the MediaWiki langlinks API and emits the
candidate fill URLs as a sidecar report — step 1 of #600 §6 methodology.

For every node with an en.wikipedia.org/wiki/<Title> link_en missing
link_<lang>, queries langlinks, captures the target title, emits:
  dataset,key,link_lang,resolved_url
to stdout or --out <path>. NEVER writes under Cards/.

Verified on a strided sample of 10 fallacies: 0 errors, 37 candidate
fills (~57% rate, consistent with the probe ceiling), URLs correctly
URL-encoded (Cyrillic %D0, PT accented %C3). Sample also surfaced the
homonym risk #600 §6.4 warns about (English "Engagement" for a fallacy
node) -> confirms human spot-validation is non-optional for AR/FA/ZH.

Safety:
  - 0 write Cards/ (sidecar only, pre-tag freeze).
  - Public MediaWiki API, no key, 0.3s throttle, descriptive UA
    (default urllib UA is 403-forbidden).
  - Resolves FROM link_en (Wikipedia only); 433 non-Wikipedia curated
    sources excluded + preserved as-is (#600 §6.2).

Next (post-release, gated): follow-up PR consumes the sidecar -> apply
candidate URLs cell-by-cell (drift-free QUOTE_MINIMAL+CRLF, method #595),
skip non-empty cells, human spot-validate ~5% residue (AR/FA/ZH, ~150).

Relates to #600, #606.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

---------

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 1, 2026
… (0 errors, 0 Cards/) (#618)

PRIMAIRE of ai-01 deep-queue #2 (dispatch msg-...92kynx). Produces the
candidate-URL output behind the #600 §5.1 probe ceiling.

Ran tools/link-langlinks-resolve.py (PR #610, merged) over full Fallacies +
Virtues taxonomy: 829 articles probed -> 2934 candidate link_<lang> URLs
(fallacies 2754, virtues 180), 0 errors. Materializes the ~57% probe ceiling
(~2919 cells) into concrete reviewable URLs. Step 1 of #600 §6 methodology.

Per-lang: ru 371, pt 430, es 513, ar 603, fa 480, zh 537. Resolves FROM
link_en (Wikipedia only); 433 non-Wikipedia curated sources excluded +
preserved (#600 §6.2).

Quality verified:
  - URL encoding correct per script (Cyrillic %D0, Arabic %D8, Persian %D8,
    CJK %E8, PT/ES Latin unencoded).
  - Homonym scan (#600 §6.4) of 1620 AR/FA/ZH candidates: 0 real leaks.
    AR 0/603, FA 0/480 Latin-path; ZH 3/537 (0.6%) = legit loanwords/
    acronyms (FUD x2, Creepypasta) — not errors. The "Engagement" homonym
    from PR #610 sample does NOT recur at scale (isolated case).
  - Asymmetry: virtues ru/pt = 4 each (sparse abstract-concept coverage)
    vs ar/fa/zh 46-56 — coverage reality, flagged for spot-validation.

Sidecars (reviewable snapshot, UTF-8 no-BOM LF):
  - docs/taxonomy/600-link-resolve-fallacies.csv (2754 rows)
  - docs/taxonomy/600-link-resolve-virtues.csv (180 rows)
  schema: dataset,key,link_lang,resolved_url

Robustness: try/except + continue (skip on network error, never abort) ->
0 errors/829. Public MediaWiki API, no key, 0.3s throttle, descriptive UA.

Next (gated post-release): apply PR consumes sidecar -> cell-by-cell drift-free
(QUOTE_MINIMAL+CRLF, method #595), skip non-empty, human spot-validate ~5%
AR/FA/ZH residue (~150 cells), re-run probe to confirm gap closed.

Scope: docs/taxonomy/ only. 0 write Cards/, 0 AssetConverter code change
(pre-tag safe). Base 18b4d02.

Relates to #600, #606, #610.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants