From e99088a69833836147d60d4df197303a5369b6ac Mon Sep 17 00:00:00 2001 From: Your Date: Thu, 23 Jul 2026 05:47:05 +0200 Subject: [PATCH] =?UTF-8?q?docs(taxonomy):=20#847=20AIF=20audit=20Layer=20?= =?UTF-8?q?(c)=20label-distortion=20flag=20(staged,=20read-only)=20?= =?UTF-8?q?=E2=80=94=20MEASURE=20not=20verdict?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Layer-(c) second installment dispatched by ai-01 (msg-20260723T024907-2l19wg, tick 89). Completes the §2.1 half (internal-node labelling constraint); companion to #853 (§2.2 reference-reachability, merged 9f3caaef). Read-only, 0 run live, 0 prod CSV. §2.1 has two bullets; only one is mechanically computable: - (2) lexical availability (borrowed_from_lang): YES — grounded in the taxonomy's own Latin column + morphology. Finds Ad hominem (raw Latin) among the 39 grouping nodes. Axis-C lexical pressure on the grouping layer = LOW (1/39 raw-Latin). - (1) distorted scope (stretched/narrowed): NO — semantic. Two proxies tested & non-discriminatory (documented, NOT used): concept-dispersion SATURATED (median 0.94-1.00, each fallacy is its own Wikipedia article); EN-grouping 1:1 everywhere (EN columns are a FR translation, not an independent tradition tree -> no axis-C structural signal in the grouping layer). label_fit: 1 borrowed (Ad hominem) + 38 unflagged (no mechanical flag, NOT a verified scope match). Cross-ref acompte -> combined_reading: Ad hominem = borrowed x A-candidate = PRIORITY arbitration candidate -> tradition-divergence (bridge-node, do NOT bill as defect), corroborates #853 (Obstruction base 69% wiki-anchored). Resolves the c-DEFER on Ad hominem. The other 6 A-candidates carry no mechanical label signal -> genuine axis-A or human semantic read; layer (c) bills none of them as tradition artifacts. Governance: MEASURE not verdict. 0 prod CSV. 0 reorganisation wording. borrowed is lexical fact (conservative); unflagged = absence of flag, not a positive verdict. Synthesis + verdict = ai-01/jsboige. Artefact: docs/taxonomy/847-layerc-label-distortion.{md,csv} (39 nodes). Relates #847 #498 #853 #850. Base master 82a1e027. Co-Authored-By: Claude-Code --- docs/taxonomy/847-layerc-label-distortion.csv | 43 ++++++ docs/taxonomy/847-layerc-label-distortion.md | 146 ++++++++++++++++++ 2 files changed, 189 insertions(+) create mode 100644 docs/taxonomy/847-layerc-label-distortion.csv create mode 100644 docs/taxonomy/847-layerc-label-distortion.md diff --git a/docs/taxonomy/847-layerc-label-distortion.csv b/docs/taxonomy/847-layerc-label-distortion.csv new file mode 100644 index 00000000..3ee431a0 --- /dev/null +++ b/docs/taxonomy/847-layerc-label-distortion.csv @@ -0,0 +1,43 @@ +# 847 Layer-(c) label-distortion flag — per internal node. MEASURE not verdict. INPUT for ai-01 synthesis. GATED jsboige ratification. 0 prod CSV. +"# label_fit in {borrowed, unflagged}. borrowed=Latin/anglicism (taxonomy Latin col = ground truth + morphology, §2.1 b2). unflagged=NO mechanical flag (NOT verified scope match). stretched/narrowed (§2.1 b1) are SEMANTIC — two proxies tested & non-discriminatory: dispersion_fr SATURATED (median 0.94-1.00), n_en_groupings=1 everywhere (EN=FR translation, no axis-C structural signal in grouping layer)." +# combined_reading = label_fit x acompte decision_rule_tag. borrowed x A-candidate = PRIORITY (tradition-divergence -> bridge-node); borrowed x coherent = VINDICATED. +level,node,n_leaves_full,n_modelled,label_fit,borrowed_lang,dispersion_fr,n_en_groupings,acompte_tag,combined_reading +family(d1),Abus de langage,89,57,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +family(d1),Erreur de raisonnement,102,14,unflagged,-,0.793,1,A-candidate,"tree-tension, NO mechanical label-distortion signal -> genuine axis-A mix OR human semantic read needed (stretched/narrowed not mechanically detectable)" +family(d1),Erreur mathématique,102,12,unflagged,-,0.968,1,B+A,"material-resistance + type-mix, no label signal" +family(d1),Influence,420,12,unflagged,-,0.885,1,coherent,"mechanism-coherent, no label signal" +family(d1),Insuffisance,174,16,unflagged,-,0.935,1,B+A,"material-resistance + type-mix, no label signal" +family(d1),Obstruction,126,12,unflagged,-,0.95,1,B+A,"material-resistance + type-mix, no label signal" +family(d1),Tricherie,394,22,unflagged,-,0.942,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Ad hominem,46,4,borrowed,latin,0.875,1,A-candidate,"PRIORITY: borrowed x tree-tension -> tradition-divergence (bridge-node, do NOT bill as defect; corroborates #853 wiki-anchor)" +sub-family(d2),Ambiguïté,41,31,unflagged,-,1.0,1,B,"material-resistance, no label signal" +sub-family(d2),Appel à l'émotion,57,4,unflagged,-,1.0,1,coherent,"mechanism-coherent, no label signal" +sub-family(d2),Argument bâclé,68,7,unflagged,-,0.917,1,A-candidate,"tree-tension, NO mechanical label-distortion signal -> genuine axis-A mix OR human semantic read needed (stretched/narrowed not mechanically detectable)" +sub-family(d2),Arranger les faits,85,4,unflagged,-,0.939,1,B,"material-resistance, no label signal" +sub-family(d2),Causalité douteuse,29,5,unflagged,-,0.929,1,A-candidate,"tree-tension, NO mechanical label-distortion signal -> genuine axis-A mix OR human semantic read needed (stretched/narrowed not mechanically detectable)" +sub-family(d2),Changement de cap,51,5,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Comparaison fallacieuse,13,10,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Définition inexacte,34,16,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Généralisation abusive,37,4,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Manipulation mentale,239,5,unflagged,-,0.818,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Mauvaise composition,32,4,unflagged,-,1.0,1,coherent,"mechanism-coherent, no label signal" +sub-family(d2),Mauvaise déduction,40,4,unflagged,-,0.875,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Mauvaise interprétation,34,4,unflagged,-,0.923,1,A-candidate,"tree-tension, NO mechanical label-distortion signal -> genuine axis-A mix OR human semantic read needed (stretched/narrowed not mechanically detectable)" +sub-family(d2),Pensée biaisée,257,13,unflagged,-,0.946,1,A-candidate,"tree-tension, NO mechanical label-distortion signal -> genuine axis-A mix OR human semantic read needed (stretched/narrowed not mechanically detectable)" +sub-family(d2),Procédé rhétorique,123,3,unflagged,-,0.975,1,B,"material-resistance, no label signal" +sub-family(d2),Préjugé,63,4,unflagged,-,0.9,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Refus du débat,31,4,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Résultat invalide,30,4,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Saboter le débat,48,3,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-family(d2),Surinterprétation,42,4,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-sub(d3),Acception arbitraire,22,6,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-sub(d3),Acception vague,4,4,unflagged,-,0.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-sub(d3),Ambiguïté narrative,11,11,unflagged,-,1.0,1,B,"material-resistance, no label signal" +sub-sub(d3),Amphibologie,8,6,unflagged,-,1.0,1,B,"material-resistance, no label signal" +sub-sub(d3),Biais culturels,68,3,unflagged,-,1.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-sub(d3),Biais naturels,150,8,unflagged,-,0.896,1,A-candidate,"tree-tension, NO mechanical label-distortion signal -> genuine axis-A mix OR human semantic read needed (stretched/narrowed not mechanically detectable)" +sub-sub(d3),Comparaison abusive,5,3,unflagged,-,0.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-sub(d3),Définition inconsistante,7,5,unflagged,-,0.0,1,B,"material-resistance, no label signal" +sub-sub(d3),Equivoque,21,13,unflagged,-,1.0,1,B,"material-resistance, no label signal" +sub-sub(d3),Fausse analogie,5,4,unflagged,-,0.0,1,B+A,"material-resistance + type-mix, no label signal" +sub-sub(d3),Justification triviale,22,3,unflagged,-,1.0,1,coherent,"mechanism-coherent, no label signal" diff --git a/docs/taxonomy/847-layerc-label-distortion.md b/docs/taxonomy/847-layerc-label-distortion.md new file mode 100644 index 00000000..80a5a87b --- /dev/null +++ b/docs/taxonomy/847-layerc-label-distortion.md @@ -0,0 +1,146 @@ +# #847 AIF structural audit — Layer (c) label-distortion flag (read-only, 0 run) + +> **Provenance.** Layer-(c) **second installment**, dispatched by ai-01 +> (`msg-20260723T024907-2l19wg`, tick 89, PRIMARY). **Read-only — 0 run live, 0 +> prod-CSV.** Companion to [#853](847-layerc-reference-reachability.md) (Layer-(c) +> reference-reachability, the §2.2 half). This completes the §2.1 half: the +> **internal-node labelling constraint**. Method doc +> [`aif-structural-audit-method.md`](aif-structural-audit-method.md) §2.1, §4, §5. +> **MEASURE, not verdict. INPUT for ai-01 synthesis. GATED jsboige ratification.** +> Scope: the **39 internal nodes** in [`847-acompte.csv`](847-acompte.csv) (the +> (a)+(b) layers were measured on the 145 fully-modeled subset; the +> `combined_reading` join only resolves there). + +--- + +## TL;DR + +`label_fit` is computed for the **39 grouping nodes** of the acompte. The §2.1 +labelling constraint has **two bullets**; only one is mechanically computable: + +| §2.1 bullet | What | Mechanically computable? | This instrument | +|-------------|------|--------------------------|-----------------| +| **(2) lexical availability** | a term is **borrowed** from another tradition where no FR term exists | **Yes** (lexical facts) | ✅ `borrowed_from_lang` | +| **(1) distorted scope** | label **stretched / narrowed** vs the cluster | **No** (semantic) | ⛔ proxies tested, non-discriminatory (§3) | + +**Headline:** of 39 grouping nodes, **1 is `borrowed`** (`Ad hominem`, raw Latin — +confirmed against the taxonomy's own `Latin` column); **38 are `unflagged`** (no +mechanical distortion signal). `unflagged` means *no mechanical flag raised*, **not** +a verified scope match. + +**Axis-C lexical pressure on the grouping layer is LOW** (1/39 raw-Latin): the +taxonomy is authored in FR descriptive phrases (or FR calques), importing raw Latin +**only** where no FR single-term equivalent exists — *Ad hominem* (the FR calque +*Attaque personnelle* sits at d3, but the d2 grouping keeps the Latin). This is a +real structural finding about the tree. + +## The one priority candidate (distorted × heterogeneous) + +| Node | Level | label_fit | acompte tag | Combined reading | +|------|-------|-----------|-------------|------------------| +| **Ad hominem** | sub-family (d2) | **borrowed (latin)** | **A-candidate** | **PRIORITY arbitration candidate** → **tradition-divergence → bridge-node, do NOT bill as grouping defect** | + +This **resolves the `c-DEFER` on Ad hominem** from the acompte. Combined with #853 +(Ad hominem's Obstruction base is **69% Wikipedia-anchored**, well-evidenced), the +layer-(c) reading is consistent: *Ad hominem*'s mechanism-heterogeneity is the +**tradition's** (a Latin scholastic term spanning cultures), **not** a jsboige +grouping error. The bridge holds. + +There are **no `borrowed × B+A`** nodes (Ad hominem is the only borrowed node), so +no second-tier "label may have forced a bad merge under material-resistance" cases +arise from this mechanical pass. + +## A-candidate resolution (the 7 tree-tension nodes) + +| Node | Level | label_fit | Layer-(c) labelling read | +|------|-------|-----------|--------------------------| +| **Ad hominem** | d2 | **borrowed** | tradition-divergence → **bridge-node** | +| Erreur de raisonnement | d1 | unflagged | no label signal → genuine axis-A mix **or** human semantic read | +| Argument bâclé | d2 | unflagged | " | +| Causalité douteuse | d2 | unflagged | " | +| Mauvaise interprétation | d2 | unflagged | " | +| Pensée biaisée | d2 | unflagged | " | +| Biais naturels | d3 | unflagged | " | + +**Reading for ai-01/jsboige:** of the 7 tree-tension (A-candidate) nodes, **only +Ad hominem carries a mechanical axis-C label signal** (borrowed Latin). The other 6 +have **no mechanical labelling distortion** → their heterogeneity is either a +genuine axis-A grouping choice **or** a stretched/narrowed label that this +first-pass cannot detect mechanically (semantic; §3). Layer (c) therefore does +**not** bill any of the 6 as tradition artifacts — it leaves them as axis-A +candidates for jsboige, exactly per the method doc's "check axis C first; if low, +the grouping is genuinely his to arbitrate." + +## Why stretched/narrowed are NOT mechanically derived (§3 — saves future work) + +Two obvious mechanical proxies for §2.1 bullet (1) (distorted scope) were tested +and found **non-discriminatory**. They are reported in the CSV as contextual +evidence (`dispersion_fr`, `n_en_groupings`) but **not** used to derive `label_fit`. + +1. **Concept-dispersion (within-language scope sprawl).** `dispersion_fr` = + distinct FR Wikipedia slugs among a node's FR-referencing leaves / FR-referencing + leaves. **Saturated**: median **0.94 (d1) / 1.00 (d2) / 1.00 (d3)** — i.e. at + sub-family and below, essentially **every leaf is its own Wikipedia article**, so + distinct-slugs ≈ leaves everywhere. It cannot distinguish a tight cluster from a + sprawled one. *(Note: `dispersion_fr = 0.0` on a few small d3 nodes means their + leaves carry **no FR Wikipedia reference** (dict-only / empty), not "perfectly + clustered" — a coverage gap, flagged, not a distortion signal.)* + +2. **EN-grouping divergence (cross-tradition structural signal).** For each FR + grouping node, count distinct EN grouping labels (`Family`/`Subfamily`/`Subsubfamily`) + among its leaves. **Result: 1:1 everywhere** (97% EN coverage; every FR node maps + to exactly one EN node). The EN grouping columns are a **FR translation/calque, + not an independent EN-tradition tree**. So no axis-C structural divergence is + encoded at the grouping layer — the cross-linguistic divergence the method doc + describes (§2 axis C) lives at the **leaf-reference** level (different Wikipedia + articles per language, already measured in #853), not in the grouping labels. + +**Consequence:** a defensible `stretched`/`narrowed` flag requires **human semantic +read** (compare each grouping label's consecrated scope to its members) — it is not +mechanizable from the taxonomy + #853 data without an LLM-assisted pass (heavier, +gated). This first-pass delivers the **mechanically-grounded** part (`borrowed`) and +documents the dead-end on the rest, so no future worker re-tries these proxies. + +## Governance + +- **MEASURE, not verdict.** 0 reorganisation wording. Synthesis + verdict = ai-01 / + jsboige. Same governance as #850 / #853. +- **0 prod-CSV write** (T&A freeze). Docs-only artefact. DatasetUpdater untouched, + `Enabled=false`. +- **`borrowed_from_lang` is grounded in lexical fact**, not semantic claim: a node + is `borrowed` only if its label matches the taxonomy's own `Latin` column or Latin + morphology (conservative). `unflagged` is explicitly *not* a positive "exact" + verdict — it is the absence of a mechanical flag (mirrors #853's "dead = upper + bound" honesty; a no-match is not a content claim). +- Calques (e.g. *Pétition de principe* ← *petitio principii*) are a **gray zone not + flagged** by this conservative pass: the term was translated into FR (FR words), + so by §2.1 bullet (2) it is not a raw borrow. Flagged for human read if a calque's + scope is later found distorted. + +## Caveats + +1. **Coverage of `borrowed` is deliberately narrow** (raw Latin/anglicism only). + A broader "compromise-label" notion (vernacular terms pressed into taxonomy + service, e.g. *Tricherie*, *Humour*) is real but not mechanically detectable + without a French frequency lexicon — out of scope, flagged. +2. **`unflagged` ≠ "label fits".** 38/39 nodes are unflagged; this instrument makes + no positive scope-match claim for any of them. It only surfaces the 1 node where + a mechanical distortion signal exists. +3. The join resolves **only** the 39 acompte nodes. Full-tree extension (all ~92 + grouping nodes) is the same heuristic but yields no `combined_reading` without + the acompte (a)/(b) layers — out of scope here. + +## Artefacts + +- [`847-layerc-label-distortion.csv`](847-layerc-label-distortion.csv) — per + internal node (39 rows): `level,node,n_leaves_full,n_modelled,label_fit,borrowed_lang, + dispersion_fr,n_en_groupings,acompte_tag,combined_reading`. +- Compute script (scratchpad): `compute-847-layerc-label-distortion.py`. + +## Refs + +- Method: [`aif-structural-audit-method.md`](aif-structural-audit-method.md) §2.1, + §4, §5. Companion Layer-(c) half: [`847-layerc-reference-reachability.md`](847-layerc-reference-reachability.md) + (#853). Acompte (layers a+b): [`847-acompte-homogeneity.md`](847-acompte-homogeneity.md) + (#850). Tracking: #847. Chantier: #498. Dispatch ai-01 `msg-20260723T024907-2l19wg` + (tick 89). Base master `82a1e027`.