Skip to content

fix(i18n): #192 passe 2 — harmonize 24 Virtues term-structure cells (RTL/CJK consistency) - #595

Merged
jsboige merged 1 commit into
masterfrom
fix/192-virtues-terminology-consistency-rtl-cjk
Jun 24, 2026
Merged

fix(i18n): #192 passe 2 — harmonize 24 Virtues term-structure cells (RTL/CJK consistency)#595
jsboige merged 1 commit into
masterfrom
fix/192-virtues-terminology-consistency-rtl-cjk

Conversation

@jsboige

@jsboige jsboige commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

#192 passe 2 — harmonize 24 Virtues term-structure cells (RTL/CJK consistency)

Worker: po-2024 · Base: e789ae08 (master) · Scope: Virtues CSV, 1 file, 24 cells.

Context

#192 (multi-pass i18n) idle dispatched by ai-01 (jjvdah). Before any LLM batch I ran a
coverage diagnostic (scratchpad, FR-relative): Virtues text content is already 100%
covered
across 7 langs — the "pass 1" is done. So I pivoted to the mechanically
verifiable
slice of pass 2: terminological consistency on structure labels.

A taxonomy family/subfamily/subsubfamily MUST have ONE translation per language (nodes are
grouped by shared label). I detected 12 FR terms with >1 distinct translation; 6 had a
clear >=80% majority
(OBVIOUS = data-entry drift from the bulk RTL/CJK passes #364) —
harmonized. 6 were <80% / near-ties (ARBITRARY) — left intact, need native/glossary
(ASK pending).

Harmonized (24 cells, 6 groups, all >=80% majority)

field.lang FR term harmonized to (majority) outliers
family.fa Échange enrichissant تبادل غنی‌ساز 5
family.zh Échange enrichissant 充实性交流 5
subfamily.ar Effort d'objectivité جهد موضوعي 2
subfamily.fa Effort d'objectivité کوشش برای عینیت 2
subfamily.zh Effort d'objectivité 客观性努力 2
subsubfamily.ar Raisonnement concluant استدلال حاسم 8

Drift-free proof (cell-level)

  • csv QUOTE_MINIMAL + CRLF = byte-identical round-trip to original (verified pre-edit).
  • Cell-level diff vs HEAD: exactly 24 cells changed, 0 drift — 224 rows, 78 cols unchanged.
  • Re-run consistency: 0 OBVIOUS remain on these 6 groups.
  • UTF-8 no-BOM preserved.

NOT touched — 6 ARBITRARY (ASK to ai-01, need native/glossary)

family.fa "Raisonnement valide" (32v23) · subsubfamily.{ru,ar,fa,zh} "Raisonnement concluant"

  • "Mise à distance des idéologies" splits. These are <80% majority or 3-way ties → harmonizing
    to "majority" would be arbitrary; deferred to a terminological glossary decision.

Verified

Build 0 errors; suite 540 passed / 0 failed / 5 skipped (no regression). Fallacies had 0
such inconsistencies (already polished historically).

Honesty

No native speaker review this PR (harmonization to a >=80% majority is objective, not a
preference). The 6 ARBITRARY cases explicitly flagged for native/glossary — NOT harmonized.

🤖 Worker po-2024 · #192 passe 2 (mechanical slice) · 24 cells drift-free · 6 arbitrary flagged

…RTL/CJK consistency)

A taxonomy family/subfamily/subsubfamily field MUST have ONE translation per language
(an educational taxonomy groups nodes by shared label). Where one FR term had >1 distinct
translation but ONE variant was the clear >=80% majority, the minority variants were
data-entry drift introduced during the bulk RTL/CJK passes (#364 et al.) — harmonized to
the majority. This is the verifiable-mechanical slice of #192 pass 2 (terminological
consistency), no glossary/native judgement needed: the majority has already decided.

Harmonized (24 cells, 6 term/lang groups, all >=80% majority):
- family_fa "Échange enrichissant": 5 outliers -> majority (تبادل غنی‌ساز)
- family_zh "Échange enrichissant": 5 outliers -> majority (充实性交流)
- subfamily_ar "Effort d'objectivité": 2 outliers -> majority (جهد موضوعي)
- subfamily_fa "Effort d'objectivité": 2 outliers -> majority (کوشش برای عینیت)
- subfamily_zh "Effort d'objectivité": 2 outliers -> majority (客观性努力)
- subsubfamily_ar "Raisonnement concluant": 8 outliers -> majority (استدلال حاسم)

Method: csv.reader/writer with QUOTE_MINIMAL + CRLF (verified byte-identical round-trip to
the original). Cell-level diff vs HEAD: EXACTLY 24 cells changed, 0 drift — 224 rows and
78 columns unchanged. UTF-8 no-BOM preserved. Re-run consistency check: 0 OBVIOUS
inconsistencies remain on these 6 groups.

NOT touched (6 ARBITRARY cases, <80% majority or near-ties — need native/glossary
judgement, raised to ai-01 as ASK): family_fa "Raisonnement valide" (32v23),
subsubfamily_ru/ar/fa/zh "Raisonnement concluant" + "Mise à distance des idéologies"
splits. Left intact pending terminological glossary decision.

Context: preceded by the #192 coverage diagnostic — Virtues text content is already 100%
covered across 7 langs; this fixes structural-label consistency, not coverage. Fallacies
had 0 such inconsistencies (already polished historically).

Verified: build 0 errors, suite 540 passed / 0 failed / 5 skipped (no regression).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@clusterManager-Myia clusterManager-Myia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[NanoClaw]

ArgumentumGames/Argumentum #595fix(i18n): #192 passe 2 — harmonize 24 Virtues term-structure cells (RTL/CJK consistency) · 1 file (Cards/Fallacies/Argumentum Virtues - Taxonomy.csv) · +15/-15 · head 91c4504 ← master e789ae08 · depth: deep cell-level CSV diff (byte-exact, cross-locale completeness).

Verdict: LGTM

A textbook mechanical-harmonization PR. I parsed both base and head versions of the 78-column × 224-row CSV with a dedicated RFC4180 parser and diffed cell-by-cell; the PR does exactly what it claims and nothing more.

Placeholder integrity

None of the 24 changed cells contain any placeholder/interpolation ({x}, %s, {{x}}, ICU) — these are short taxonomy labels (≤13 chars), not template strings. No placeholder-mismatch risk exists. All other template-bearing cells are untouched.

Spot-checks (4 cells across RTL + CJK)

  1. subsubfamily_ar (rows 105–123, 8 cells) — FR "Raisonnement concluant". Three pre-existing variants merged to the 80% majority استدلال حاسم (conclusive). The displaced variants استدلال قاطع (categorical) and استدلال منتج (valid/productive) are all legitimate Arabic logic terms for the same concept; the chosen majority term is the common rendering and matches the FR source. No stray LTR marks, no bidi-control characters.

  2. family_fa (rows 181/197–200, 5 cells) — FR "Échange enrichissant"تبادل غنی‌ساز. The Persian ZWNJ (U+200C) between غنی and ساز is correctly preserved (verified at codepoint level: U+063A U+0646 U+06CC U+200C U+0633 U+0627 U+0632). This is a grammatically valid Persian compound ("enrichment-making exchange"), not mojibake.

  3. family_zh (rows 181/197–200, 5 cells) — 充实性交流 (substantive exchange). Correct noun+性+noun structure, consistent with the sibling 客观性努力. Outlier 富有成效的交流 (also valid, but longer) correctly yielded to the 89% majority.

  4. subfamily_{ar,zh,fa} (rows 179–180) — FR "Effort d'objectivité" harmonized in all three touched languages to a compact noun structure. Cross-language check on row 179 confirms structural parallelism: جهد موضوعي / 客观性努力 / کوشش برای عینیت / "Objectivity effort".

Harmonization completeness (the strongest check)

I regrouped every row by its FR source term and counted distinct translations per (level, language), before and after:

  • Before: 8 inconsistent groups in the touched languages.
  • After: exactly the 6 claimed groups are now uniform. The 2 remaining (family_fa "Raisonnement valide" at 58%, subsubfamily_ar "Mise à distance des idéologies" at 60%) are the <80% "ARBITRARY" cases the PR explicitly defers to native/glossary review — both correctly left intact. The PR's own accounting is accurate to the cell.

Verification

  • Catalog structurally valid: 224 rows × 78 cols in both base and head, 0 structural drift, balanced quotes (RFC4180 parse clean), valid UTF-8 (Buffer round-trip identical), no BOM (matches claim), CRLF preserved (224/224, 0 lone-LF), -59-byte delta consistent with shorter harmonized strings.
  • Security clean: no URLs/paths/secrets/scripts/API-keys in any of the 24 changed cells. The only URLs in the whole file are pre-existing per-row Wikipedia reference links (untouched, present in base).
  • Scope tight: changes confined to exactly 6 translation columns (*_ar, *_fa, *_zh). FR master columns, structural columns (pk/path/depth/decimal_path_padded), and all other locales (en/ru/…) are byte-identical to base. No code files touched.

Notes (non-blocking)

  • Honesty caveat acknowledged by the author: no native-speaker review for this pass. The harmonization is to a ≥80% majority, which is objective rather than preferential, but the 2 deferred <80% cases do warrant a glossary decision in a future pass.
  • The displaced outlier terms were all linguistically valid (not errors being fixed, just non-majority synonyms) — so this is genuine de-duplication, not a correction of wrong strings.

— NanoClaw (myia-ai-01)

@jsboige
jsboige merged commit bef3bc6 into master Jun 24, 2026
3 checks passed
@jsboige
jsboige deleted the fix/192-virtues-terminology-consistency-rtl-cjk branch June 24, 2026 15:15
jsboige added a commit that referenced this pull request Jun 25, 2026
…5 post-#595 (vraiment)

§3.3: 'RAPPORTÉ régén 12 juin' → ✅ régên fraîche 2026-06-25 (bef3bc6). PDFs Virtues
re-rendered post-#595 (clobber targeted, Chromium invoqué, labels familiaux localisés
vérifiés : zh 有效论证, ar حجة_معتبرة). i18n distinct anti-leak #216 OK. Régén headless.
§5.2: décision 'régén fraîche ?' → FAITE. Risque résiduel '16 commits depuis 12 juin' levé.
§7: risque 'régén 12 juin non reproduite' → LEVÉ (régén fraîche reproduit counts + #592/#595).
  + note de procédure stale-harvest honnête (oubli clobber corrigé ce cycle).
§8: reco point 3 (régén fraîche) → FAITE.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 25, 2026
…-NOTES régên date

Coherence cross-check RELEASE-NOTES + dossier §3.2 (extension relecture po-2024
qui n'avait vérifié que RELEASE-VALIDATION). 2 nombres faux vs assets réels :

- Taille Fallacies_zh.svg : '17.2 MB (18 075 919 B)' → '5.45 MB (5 451 309 B)'
  (fichier réellement commité, byte-proven régên 2026-06-25). Aucun SVG zh n'est
  à 17.2 MB (content/links = 4.0 MB, main = 5.45 MB). Corrigé dans RELEASE-NOTES
  l.23 + dossier §3.2 table zh.
- RELEASE-NOTES l.22 : 'régên 12 juin 2026' → 'régên fraîche 2026-06-25 bef3bc6
  post-#592/#595' (cohérent dossier §3.3 refresh).

CHANGELOG '21 SVGs FR/EN/RU/PT' l.47 = section historique recovery (commits
avril 2026), pas contradiction avec l.16 Added '8 languages' = laissé intact.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 27, 2026
* docs(i18n): #192 link_* coverage research track (own the URL gap)

Research/own the link_* URL coverage gap, per ai-01 deep-queue v2 dispatch
2026-06-25 (msg-…jp3hx2). Output = docs only, 0 write under Cards/
(release freeze), master stays bef3bc6.

Key findings (measured read-only on bef3bc6):
- link_<lang> cells are per-language Wikipedia article URLs
  (quasi-exclusively; link_{ru,pt,es,ar,zh,fa} = 100% wikipedia-<lang>).
- The gap is NOT translation and NOT pure human research — it is
  cross-language article resolution, semi-automatable via the MediaWiki
  `langlinks` API (no key, rate-limited). Refines memory
  i18n-coverage-gap-is-link-urls ("human, not gpt-5.5" -> "API-resolvable
  + human-validated residue").
- Fillable-candidate upper bounds (have link_en, missing link_<lang>):
  Virtues 364 cells, Fallacies 8110 cells. These are CEILINGS, not
  guarantees — langlinks returns nothing when the target article doesn't
  exist (genuine Wikipedia content gap, out of our control).
- Root cause of per-language variance: EN 95% (reference corpus), FR 45%
  (sub-families lack dedicated article), RU/PT/ES/AR/ZH/FA 6-9% (these
  Wikipedias have fewer articles for these specific fallacies).

Proposed methodology (follow-up PR, post-release): resolve via langlinks
for nodes with link_en, preserve curated non-Wikipedia links, human
spot-validate residue (~5%, critical for AR/FA/ZH homonym risk), CSV
safety QUOTE_MINIMAL+CRLF as #595.

Authoritative sources: wikipedia-<lang> (primary, via langlinks), Wikidata
sitelinks (more complete for rare terms), yourlogicalfallacyis (EN depth,
already used), Stanford Encyclopedia of Philosophy (academic depth).

Honesty: figures are upper bounds not guarantees; disambiguation/homonym
risk for RTL/CJK; no overwrite of curated links; not in #192 LLM scope.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(i18n): #192 measure link_* fill-rate via langlinks census (refine #600 ceiling → decision-grade)

Idle-fallback of dispatch v2 ("read-only measure to enrich the jsboige
decision dossier; GLM = abundant budget"). Converts #600's theoretical
ceiling (8110 Fallacies candidates) into a MEASURED number by probing
every unique candidate article against the MediaWiki langlinks API.

Method: only en.wikipedia.org/wiki/<Title> link_en URLs qualify as
langlinks-resolvable (433 non-Wikipedia URLs excluded, preserved as-is).
For each candidate article, langlinks returns the cross-language article
when it exists; a missing cell is resolvable iff the article has a
langlink to that language.

Measured (full census, 0 errors, Fallacies 741 + Virtues 88 articles):
- Fallacies: 2739 / 4823 candidate cells resolve = 57%
  (ru 46%, pt 53%, es 61%, ar 66%, fa 53%, zh 60%)
- Virtues: 180 / 322 candidate cells resolve = 56%
- Combined ~2919 confirmed resolvable; ~2770 realistically auto-fillable
  after ~5% human spot-validation attrition (AR/FA/ZH homonym risk).
AR/ZH densest (60-66%) = highest fill-pass return; RU/PT/FA mid (44-53%).

This supersedes the 8110 ceiling for prioritization: the gap is a genuine
Wikipedia content gap for the missing 43% (no article / not Wikidata-linked,
unfixable by us), not a data-entry or translation gap.

New: docs/taxonomy/192-link-coverage-langlinks-probe.py (read-only,
no API key, ~0.3s throttle, descriptive User-Agent — MediaWiki 403s the
default urllib UA). Default = full census; pass N for strided sample,
"0 virtues"/"0 fallacies" for one dataset.
Doc: +sec 5.1 (measured tables), TL;DR pt 4 updated, sec 6 effort revised,
sec 9 reproducibility references the probe.

0 write under Cards/ (release freeze). master stays bef3bc6. Docs only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 28, 2026
…ster, #601 guardrail) (#604)

Add VirtueMindMapFRFrozenCharacterizationTests — a Golden-Master suite pinning
the CURRENT behaviour that Virtues mind map node text renders French for all 8
release languages (Fallacies are localized 8-lang, pinned by
MindMapLocalizationRegressionTests).

This is an executable snapshot of the 2-layer gap traced in investigation #601
(merged 71033f8), NOT a regression test of desired behaviour:

- Layer B (config): the single Virtue-targeting MindMapLocalization entry is a
  stale stub that only rewrites the tree-root literal "Vertus", for 4 of 7
  target languages (en/ru/pt/es; ar/fa/zh absent).
- Layer A (entity): Virtue exposes no TitleAr/TitleFa/TitleZh (CSV columns from
  #590/#595 exist but are invisible to the mind map), and Virtue.Text hard-
  delegates to TitleFr.

Three guardrail sections, each GREEN today and expected to FAIL when the 3-step
fix documented in #601 lands (entity ar/fa/zh props + {item.TitleFr} expression
+ mirrored config table). At that point flip the assertions to the localized
expectation rather than deleting them — the failure is the signal the gap closed.

The test-writing process surfaced a precision my investigation missed:
nameof(VirtueMindMapDocumentConfig.TitleExpression) returns the bare "TitleExpression",
which is SHARED with the Fallacy text entry — so the entry must be located by its
unique "Vertus" conversion, not by nameof (documented in the helper XML doc).

Additive only: no production code or existing test modified. 0 write under Cards/
(release freeze). Dispatch ai-01 2026-06-27 primary (characterization test, GO).

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 28, 2026
… gate → ratifiable WE) (#598)

* docs(i18n): #192 terminology glossary/register proposal (pre-instruct gate)

Converts the #192 glossary/register gate (blocker for passes 2-4) into a
ratifiable checklist for jsboige WE 27/06, per ai-01 dispatch 2026-06-25
(msg-…o0vg3h). Output = docs only, 0 write under Cards/ (release freeze),
master stays bef3bc6.

Contents:
- docs/taxonomy/192-terminology-glossary-register.md: TL;DR + Decision A
  (register/capitalization convention) + Decision B (per-case lexical
  picks with RATIFY columns). Audit found 14 multi-variant groups:
  1 OBVIOUS + 13 ARBITRARY — Virtues 6 (residue of #595), Scenarii PT 8,
  Fallacies 0, Rules N/A. Coverage appendix confirms text = 100% across
  7 langs x 4 datasets (gate is harmonization, not coverage); only gap
  is link_* URLs (human research, out of #192 LLM scope).
- docs/taxonomy/192-terminology-audit.py + 192-coverage-report.py:
  read-only reproducibility tools (re-run anytime from repo root).

Honesty: corrected my earlier #595-prep audit, which under-reported by
building non-EN target cols as <FR-stem>_<lang> (silently skipped
Scenarii/Fallacies non-EN). Uniform <en-base>_<lang> rule surfaces the
Scenarii PT residue that was always present. Fallacies re-verified 0 both
ways. Same cross-verify rigor as release dossier #591.

After ratification: po-2024 executes bounded mechanical harmonization
(~65 Virtues + ~30 Scenarii PT cells, QUOTE_MINIMAL+CRLF drift-free as
#595) -> #192 passes 2-4 DONE.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(taxonomy): #598 enrich Decision B with factual PT rationale + RTL/CJK status

Secondary task of ai-01 deep-queue v2 dispatch (msg-…jp3hx2): pre-digest
the 4 genuinely-jsboige PT cases (S2/S3/S6/S7) so WE 27/06 ratification
takes minutes, not an open translation debate. Docs only, 0 write under
Cards/ (release freeze), master stays bef3bc6.

Enrichment:
- PT rationale (S1-S8) is FACTUAL, not opinion: dialect/register split
  (BR-PT Paquera vs PT-PT engate = audience decision), semantic fork
  (relacoes NO trabalho interpersonal vs DE trabalho legal-HR), number
  concordance (FR "religions" plural -> Religioes, contra majority),
  scope expansion (FR "contes" tales-only vs PT "Contos e literatura").
  Each case now carries a defensible recommendation + the decision knob.
- RTL/CJK (V1-V6): clarified native-REQUIRED, gpt-5.5 = candidate assist
  only (/v1/responses + reasoning.effort=low per memory), never ratified
  as authoritative. Majority is a safe interim default if no native at WE.

Net: PT decisions ratifiable from rationale (~5 min); RTL/CJK defers to
native (non-blocking v0.9.0). Either way #192 passes 2-4 become actionable.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 28, 2026
…drift-free) (#602)

Counterpart to 192-terminology-audit.py. Applies the ratified #192
harmonization (Decision A register + Decision B lexical picks) as a
drift-free CSV edit (QUOTE_MINIMAL + CRLF, same dialect as #595), gated
behind ratification: default DRY-RUN (0 write under Cards/), --apply
writes only POST-ratification, --verify-drift proves the writer dialect
is byte-identical on a zero-change round-trip (run FIRST on any machine).

Detection mirrors the audit at runtime (FR source label -> Counter of
translations, winner = majority) — no hardcoded variant strings, stays
correct as the CSVs evolve. An EXCEPTION TABLE encodes the
glossary-register doc's per-case guidance:
  - OVERRIDE : apply a contra-majority recommended value (S6 meaning,
               S7 number — the doc argues these beat the majority)
  - FLAG     : do NOT auto-apply; list as PENDING. Covers S2 scope,
               S3 regional (jsboige judgment) and V1-V6 RTL/CJK near-ties
               (native-required, per the doc's LOW confidence).
Decision A (sentence-case) is not reimplemented as a caser (casing
Portuguese proper nouns like "Idade Media" is error-prone); the cap-only
cases resolve by picking the winner variant AS-IS, so the ratified register
surfaces through the chosen variant string, not a transform.

plan_changes scopes every change to rows sharing the group's FR source
label, so cross-family rows are never dragged into another group's winner
(the bug class that initially over-reported 1108+955 cells; fixed ->
matches the audit's 14-group census exactly).

Verified 2026-06-26 against master bef3bc6 (release-frozen, 0 write Cards/):
- --verify-drift : all 3 CSVs byte-identical on zero-change round-trip
- dry-run       : Virtues 0 auto-apply / 6 FLAG (V1-V6 RTL/CJK
                   native-required); Scenarii 39 auto-apply (PT overrides) /
                   2 FLAG (S2 scope, S3 regional); Fallacies 0/0 — matches
                   the audit (14 groups: 6 Virtues + 8 Scenarii)
- --apply       : proven end-to-end on a scratchpad copy — writes exactly
                   the 39 dry-run cells (34 lines; 5 rows carry 2 cell
                   changes each), S2/S3 flagged groups untouched

Post-ratification (WE 27/06) = 1 command:
    python docs/taxonomy/192-terminology-apply.py --apply
then re-run --verify-drift to confirm the dialect is intact. Converts the
ratification gate into instant execution.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 29, 2026
…607) (#608)

Turnkey ✅-checklist (format #598) for the 8 multi-variant groups
correctly left PENDING by #607 — exactly jsboige's ratification
"RTL/CJK + BR/PT deferred to native". Ratifiable in ~10 min by a
native speaker / jsboige, POST-release (non-blocking v0.9.0 — the
tag ships with the majority variant on each).

8 groups, re-verified on master 44c37fa via the applier dry-run
(0 would-change, 8 pending) + audit script variant counts:
  - Virtues V1-V6 (RTL/CJK near-ties, fa/ar/zh/ru): native-required,
    gpt-5.5 assist optional (/v1/responses + reasoning.effort=low)
    but LLM pick must NOT be ratified as authoritative
  - Scenarii PT S2 (contes scope: tales-only vs tales+literature)
  - Scenarii PT S3 (regional: BR 'Paquera' vs PT-PT 'engate')

Each row: 1-line context + options + recommended default + confidence.
Factual rationale (dialect/register facts, not opinion) pre-digests the
PT decisions; RTL/CJK stay native-required. Includes the post-ratification
apply path (EXCEPTIONS table FLAG→OVERRIDE + --verify-drift + --apply,
drift-free method #595) so po-2024 can close #192 fully once ratified.

Scope: docs only. 0 write under Cards/. master stays 44c37fa. Relates
to #598 (glossary proposal) + #602 (apply script) + #607 (applied the 6
ratifiable Scenarii PT groups). Closes the deferred-residue pass of #192.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jun 29, 2026
…or, 0 Cards/) (#610)

* docs(taxonomy): #141 non-card census + AIF cross-ref enrichment proposal

Grounds #141 in the CURRENT taxonomy state (issue text is ~2 years old;
DatasetUpdater + translation PRs have moved the state considerably).

Census finding (read-only script on ba8e4a6, reproducible like #600/#606):
  - Non-card nodes: Fallacies 1232, Virtues 110 (Scenarii/Rules out of scope).
  - TEXT enrichment (desc/example/title x 8 langs) = 100% on both datasets
    => original #141 scope item 3 (descriptions + examples + translations)
    is DONE. No text-enrichment gap to script.
  - The "GPT-4 enrichment script" = already adapted: modern DatasetUpdater
    (PR #210, OpenAI SDK v2.10.0, gpt-5.5). Not a rewrite.

Reframes #141: the genuinely-open residue is the AIF cross-reference graph
(crossLink_* 8 relationship cols + AIF_skos* 4 SKOS mappings). Schema is
READY (columns exist), content ~0% on non-cards:
  - Fallacies: crossLink 0.2%, AIF_skos 0.5% (barely started)
  - Virtues: crossLink_Opposes + skosDirectRef/MappingType 100% via #498
    pilots, but 7 OTHER crossLink verbs = 0%
  - ~16700 empty cross-ref cells where schema is waiting.

Proposal: 4-stage method (ground -> gpt-5.5 candidate via /v1/responses
effort=low -> drift-free sidecar dry-run -> expert/jsboige ratification
gate -> OWL/2sxc export). NOT auto-write: AIF/Walton mappings are
specialised (memory: "Anti-Fab Validator: Walton scheme = WARN"), and
Cards/ is frozen pre-tag. "Adapt the script" = add a crossLink prompt +
task config to DatasetUpdater, GATED post-release on jsboige GO + schema
decision (single decimal_path vs structured {target,note}).

Overlaps #498 (AIF scale-up pilots produced the few filled cells).
Feeds dependencies #141 declares: #130 (OWL/SKOS) + #136 (2sxc).

Scope: docs + read-only script only. 0 write Cards/, 0 AssetConverter code
change (pre-tag safe). Base ba8e4a6.

Relates to #130, #136, #498.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

* tools(i18n): link_* langlinks resolver (sidecar candidate-URL generator, 0 Cards/)

SECONDAIRE of ai-01 deep-queue dispatch. Extends the #600/#606 coverage
track from MEASUREMENT to RESOLUTION.

The existing probe (docs/taxonomy/192-link-coverage-langlinks-probe.py)
measures the resolvable ceiling (~2919 cells, 57%, #600 §5.1) but
discards the target title (returns only lang codes). This tool captures
the target-language title via the MediaWiki langlinks API and emits the
candidate fill URLs as a sidecar report — step 1 of #600 §6 methodology.

For every node with an en.wikipedia.org/wiki/<Title> link_en missing
link_<lang>, queries langlinks, captures the target title, emits:
  dataset,key,link_lang,resolved_url
to stdout or --out <path>. NEVER writes under Cards/.

Verified on a strided sample of 10 fallacies: 0 errors, 37 candidate
fills (~57% rate, consistent with the probe ceiling), URLs correctly
URL-encoded (Cyrillic %D0, PT accented %C3). Sample also surfaced the
homonym risk #600 §6.4 warns about (English "Engagement" for a fallacy
node) -> confirms human spot-validation is non-optional for AR/FA/ZH.

Safety:
  - 0 write Cards/ (sidecar only, pre-tag freeze).
  - Public MediaWiki API, no key, 0.3s throttle, descriptive UA
    (default urllib UA is 403-forbidden).
  - Resolves FROM link_en (Wikipedia only); 433 non-Wikipedia curated
    sources excluded + preserved as-is (#600 §6.2).

Next (post-release, gated): follow-up PR consumes the sidecar -> apply
candidate URLs cell-by-cell (drift-free QUOTE_MINIMAL+CRLF, method #595),
skip non-empty cells, human spot-validate ~5% residue (AR/FA/ZH, ~150).

Relates to #600, #606.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

---------

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 1, 2026
… (0 errors, 0 Cards/) (#618)

PRIMAIRE of ai-01 deep-queue #2 (dispatch msg-...92kynx). Produces the
candidate-URL output behind the #600 §5.1 probe ceiling.

Ran tools/link-langlinks-resolve.py (PR #610, merged) over full Fallacies +
Virtues taxonomy: 829 articles probed -> 2934 candidate link_<lang> URLs
(fallacies 2754, virtues 180), 0 errors. Materializes the ~57% probe ceiling
(~2919 cells) into concrete reviewable URLs. Step 1 of #600 §6 methodology.

Per-lang: ru 371, pt 430, es 513, ar 603, fa 480, zh 537. Resolves FROM
link_en (Wikipedia only); 433 non-Wikipedia curated sources excluded +
preserved (#600 §6.2).

Quality verified:
  - URL encoding correct per script (Cyrillic %D0, Arabic %D8, Persian %D8,
    CJK %E8, PT/ES Latin unencoded).
  - Homonym scan (#600 §6.4) of 1620 AR/FA/ZH candidates: 0 real leaks.
    AR 0/603, FA 0/480 Latin-path; ZH 3/537 (0.6%) = legit loanwords/
    acronyms (FUD x2, Creepypasta) — not errors. The "Engagement" homonym
    from PR #610 sample does NOT recur at scale (isolated case).
  - Asymmetry: virtues ru/pt = 4 each (sparse abstract-concept coverage)
    vs ar/fa/zh 46-56 — coverage reality, flagged for spot-validation.

Sidecars (reviewable snapshot, UTF-8 no-BOM LF):
  - docs/taxonomy/600-link-resolve-fallacies.csv (2754 rows)
  - docs/taxonomy/600-link-resolve-virtues.csv (180 rows)
  schema: dataset,key,link_lang,resolved_url

Robustness: try/except + continue (skip on network error, never abort) ->
0 errors/829. Public MediaWiki API, no key, 0.3s throttle, descriptive UA.

Next (gated post-release): apply PR consumes sidecar -> cell-by-cell drift-free
(QUOTE_MINIMAL+CRLF, method #595), skip non-empty, human spot-validate ~5%
AR/FA/ZH residue (~150 cells), re-run probe to confirm gap closed.

Scope: docs/taxonomy/ only. 0 write Cards/, 0 AssetConverter code change
(pre-tag safe). Base 18b4d02.

Relates to #600, #606, #610.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 1, 2026
…free, gated post-release (#622)

SECONDAIRE of ai-01 deep-queue supersede #3 (msg-...v95b6l): the apply
harness for the #600 step "ratify -> apply" (§6), built from the #618
sidecar (2934 candidate link_<lang> URLs). Dry-run only — 0 write Cards/.

What it delivers:
  - drift-free write path (#595: QUOTE_MINIMAL + quotechar " + CRLF +
    UTF-8 no-BOM), verified to match the on-disk dialect of both CSVs.
  - skip-non-empty — refuses to overwrite any filled link_<lang> cell.
  - spot-validation ~5% of the AR/FA/ZH residue (§6.4 homonym risk):
    80 candidates inspected.

Dry-run headline (master d0856aa):
  - cands=2934, would-apply=2934, skip-nonempty=0 -> zero clobber.
  - 0 orphan-PK, 0 col-missing, 0 duplicate-(key,lang), 0 homonym.
  - Both CSVs #595 drift-safe (Fallacies: 1409 records = 1409 CRLF +
    144 intra-cell-LF benign; Virtues: 224 = 224 CRLF, 0 intra-cell-LF).
  - All 2754 fallacies URLs host-match their declared language (0 mismatch).
  - Spot-sample decoded correct: ar احتكام إلى الجهل (ignorance),
    zh 合成謬誤 (composition), zh 定錨效應 (anchoring) — no English leak.

The --apply path is wired but NOT exercised (freeze forbids Cards/
write). Post-release: `python docs/taxonomy/600-link-apply.py --apply`.

PK detection is case-insensitive (Fallacies='PK', Virtues='pk').

Scope: docs/taxonomy/600-link-apply.py + 600-link-apply-report.md only.
0 write Cards/, 0 AssetConverter code change. Base d0856aa.

Relates to #600, #618, #595, #192.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 1, 2026
…ANGELOG SVGs 8-lang (#591)

* docs(release): v0.9.0 validation dossier + fix CHANGELOG SVGs 8-lang

Add the v0.9.0 release validation dossier (go-live GATE for #134):
- scope (8 languages, jsboige decisions), verified asset inventory
  (CSV 8-lang, MindMap SVGs 8-lang via #565, OWL 5.4MB, PDFs reported
  from 12-Jun regen 64/64), toolchain state (zero-warning CS+NU,
  533/0/5), residual risks, and the decision points that need jsboige
  before the tag (pixel RTL/CJK eyeball, fresh regen, DNN coupling,
  tag v0.9.0).

Fix CHANGELOG line 16: it stated "ES/AR/FA/ZH SVG regeneration
pending" but #565 (merged 2fff84c) shipped those 12 SVGs. Wording
now reflects the delivered 8-language set including RTL + CJK.

Dossier built on committed assets + the reported 12-Jun release
regen (64/64 PDFs, 9834 images, exit 0); no fresh regen run this
session (bin/ empty post-crash, no RDP GO).

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(release): #591 address NanoClaw review notes on v0.9.0 dossier

- Byte-verify asset sizes: OWL 5.4MB→5.13MB (5,378,765 bytes),
  zh SVG 18MB→17.2MB (18,075,919 bytes).
- Add §3.1 source pointer for the 1408/223/167 taxonomy counts
  (issues #335/#499 spec / Scenarii commits, not re-derived this day).
- Clarify §5 point 5: CHANGELOG already fixed in this PR;
  docs/RELEASE-NOTES-v0.9.0.md does NOT exist on master
  (verified 22eb5f3) — only CHANGELOG.md documents the release.

Non-blocking review notes from ai-01/NanoClaw; no content change
to the 6 gate decisions.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(release): #591 update coupling reco + draft v0.9.0 release notes

Per ai-01 jalon dispatch (review-ready WE):
- Update dossier §5 point #3 with po-2023 recommendation: DE-COUPLE
  (tag assets-only; DNN upgrade as post-release ops milestone), with
  effort estimate pointer to UPGRADE-ASSESSMENT §10 (#593).
- Add docs/RELEASE-NOTES-v0.9.0.md (draft, did not exist on master) —
  GitHub Release body auto-built from CHANGELOG + dossier, for jsboige
  to validate/paste. Covers 8-language scope, assets, pipeline infra,
  recovery fixes, quality (533 tests, zero-warning build), and a
  security note (DNN CVEs → post-release, target 10.3.2).

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(release): #591 refresh dossier — #592 OWL merged + MindMap byte-proven + Virtues FR-gap

§3.2: + reproductibilité byte-proven (régên 8-lang 2026-06-24, Fallacies_zh.svg
cmp IDENTICAL) + gap structurel honnête (Virtues .content.svg FR-figés, non-bloquant).
§3.4: #499 Phase 2 OWL — 'en cours po-2024' → MERGED (#592, 8d5d275).
§5.5: RELEASE-NOTES-v0.9.0.md créé dans cette PR (n'existait pas sur master).
§5.6: #499 Phase 2 OWL — livré avant ce dossier, conservé pour traçabilité.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(release): #591 refresh §3.3/§5.2/§7/§8 — régên fraîche 2026-06-25 post-#595 (vraiment)

§3.3: 'RAPPORTÉ régén 12 juin' → ✅ régên fraîche 2026-06-25 (bef3bc6). PDFs Virtues
re-rendered post-#595 (clobber targeted, Chromium invoqué, labels familiaux localisés
vérifiés : zh 有效论证, ar حجة_معتبرة). i18n distinct anti-leak #216 OK. Régén headless.
§5.2: décision 'régén fraîche ?' → FAITE. Risque résiduel '16 commits depuis 12 juin' levé.
§7: risque 'régén 12 juin non reproduite' → LEVÉ (régén fraîche reproduit counts + #592/#595).
  + note de procédure stale-harvest honnête (oubli clobber corrigé ce cycle).
§8: reco point 3 (régén fraîche) → FAITE.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(release): #591 cohérence — header note de méthode + §4 tests actualisés (relecture)

Header note de méthode : 'RAPPORTÉS non re-vérifiés ce jour' → régén fraîche 2026-06-25
re-vérifiée ce jour (résout contradiction header vs §3.3 refresh tick préc.).
§4 tests : 533 pass (22eb5f3) → 540 pass (bef3bc6, test run Release 22s),
+7 = tests OWL Virtues #592. 0 fail, 5 skip (GUI/Freeplane).

Relecture cohérence read-only (pas de polish) : 1 contradiction header/corps
résiduelle trouvée + facts actualisés vs bef3bc6.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(release): #591 fix FINDING relecture po-2024 — header ligne 6 stale + mineurs

FINDING principal (po-2024 cross-verify, gate go-live) :
- ligne 6 (Master de référence) : 22eb5f3/533 → bef3bc6/540 (j'avais corrigé ligne 14
  mais manqué la 6 — contradiction header/corps résiduelle).

Mineurs :
- ligne 3 (date) : 2026-06-23 → 2026-06-25 (refresh).
- ligne 95 (attribution) : '+7 = tests OWL Virtues #592' → précisé +5 VirtueOwlGeneration
  +ContractTests + +2 VirtueClassMapRegressionTests (#592 = 8d5d275). +7 confirmé correct
  vs doute po-2024 (message commit #592 liste explicitement +5 et +2).

Dossier cohérent bout-en-bout après cross-verify indépendant po-2024.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(release): #591 fix taille SVG zh fausse (17.2→5.45 MB) + RELEASE-NOTES régên date

Coherence cross-check RELEASE-NOTES + dossier §3.2 (extension relecture po-2024
qui n'avait vérifié que RELEASE-VALIDATION). 2 nombres faux vs assets réels :

- Taille Fallacies_zh.svg : '17.2 MB (18 075 919 B)' → '5.45 MB (5 451 309 B)'
  (fichier réellement commité, byte-proven régên 2026-06-25). Aucun SVG zh n'est
  à 17.2 MB (content/links = 4.0 MB, main = 5.45 MB). Corrigé dans RELEASE-NOTES
  l.23 + dossier §3.2 table zh.
- RELEASE-NOTES l.22 : 'régên 12 juin 2026' → 'régên fraîche 2026-06-25 bef3bc6
  post-#592/#595' (cohérent dossier §3.3 refresh).

CHANGELOG '21 SVGs FR/EN/RU/PT' l.47 = section historique recovery (commits
avril 2026), pas contradiction avec l.16 Added '8 languages' = laissé intact.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(release): #591 refresh dossier v0.9.0 — régén 2026-07-01 + verdict visuel PASS ai-01

The validation dossier was last refreshed 2026-06-25 and flagged the PDF counts as
"RAPPORTÉS" and the visual verdict as "remaining (ai-01, règle HARD)". Both are now
resolved as of the 2026-07-01 cycle:

- Régén Release 8-lang executed 2026-07-01 on 18b4d02 + #614 (serial,
  EnableParallelism=false, isolated worktree) — 0 échec, 64/64 PDFs, 1229 img/lang ×8,
  "Generation finished.", bundle on GDrive review-v0.9.0-2026-06-28/ with sha256 manifest.
- Verdict visuel ai-01 = PASS représentatif (2026-07-01 08:50): spot-check
  Fallacies_Web_Thumbnails p1 across zh/ar/fa/ru/es covering CJK + RTL + cyrillic + latin;
  bug #216 (FR contamination) held, structure complete 8 types × 8 languages.
- DNN migration full-IIS now CLOSED (dnn.argumentum.myia.io LIVE) — coupling no longer an
  assets blocker.

Refreshed sections: header (date/status/master-ref), §1 (méthode), §3.3 (PDFs counts +
verdict PASS), §4 (tests 549/0/5), §5 (decisions: verdict FAIT/PASS, régén 07-01, DNN LIVE,
new §5.7 params Debug-vs-Release), §7 (risks lifted), §8 (recommendations).

content-only, 0 write Cards/, 0 runtime change. CHANGELOG + RELEASE-NOTES untouched.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 4, 2026
…oposal + sample, gated) (#673)

Secondaire dispatch l0wt63 (ai-01). Additive proposal for #141's remaining
engineering artifact: the crossLink DatasetUpdater task that lets the Stage-3
expert-gate output flow into Cards/ drift-free.

State recap (code=truth): #141 text-enrichment DONE (100% x8 langs), AIF
cross-ref DELIVERED through Stage-3 (#626: 1232/1232, 0 fab), 'adapt the GPT-4
script' DONE (DatasetUpdater = modern counterpart, SDK v2.10.0/gpt-5.5). Closure-rec
says remaining work is expert-gate judgment, not engineering. This stages the ONE
engineering artifact that follows the gate.

Proposal contents (text in doc, NOT materialized files):
- Task config skeleton (Enabled=false, mirrors the 7 existing configs)
  - UseFunctionCalling=true (closed-set anti-fab, the #626-validated design)
  - SkipNonEmpty=true (preserves the 12 expert-adjudicated existing-AIF nodes)
  - ChunkSize=1 (per-node relationship judgement)
- Prompt design (PromptCrossLinkSystem sample framing, closed verb+path set)
- Sample (echantillon) drawn from committed sidecar 141-aif-candidates-sample.csv
- Staged flow gated post-tag + post-Stage-3-ratification

Gated: 0 Cards/ write, 0 AssetConverter C# change (pre-tag freeze honored per
closure-rec). Same method as scale-ups #497-499 (proposal doc + sample first,
prod-write separate gated step).

Relates #141 #626 #130 #136 #498 #595. Base a41cbda.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants