docs(taxonomy): #141 AIF Stage-1 full-scale results — 1232 nodes, 0 fabrication - #626
Merged
Conversation
…abrication PRIMAIRE of ai-01 deep-queue supersede #3 (msg-...v95b6l): the full-scale AIF Stage-1 candidate sidecar + coverage/WARN report. The generator was shipped in #623; this is the run output (1232/1232 non-card nodes, gpt-5.5 assist, dry-run, 0 write Cards/). Headline — anti-fab closed-set HOLDS at full scale: - 1232 nodes covered (100% of Fallacies non-cards, the #609 gap). - 3850 crossLinks (3.12/node, 100% nodes) + 2310 AIF refs (1.88/node, 96% nodes). - confidence mean 0.73, median 0.72 (1301 high >=0.8 / 2460 mid / 89 low). - FABRICATION = 0. invalid targets = 0. invalid verbs = 0. The closed-set design (targets from enumerated real nodes, AIF tokens from the 60-token observed Walton vocab) contains LLM fabrication completely across 1232 nodes. crossLink verb mix: IsRelatedTo 42% / Leverages 33% / Mirrors 15% / Allows 7% / Inverts 2% / Opposes 1% / PredatesOn 1% / Denounces <1%. Honest improvement vs pilot: Opposes/PredatesOn/Denounces were 0 in the 28-node pilot (flagged limitation) -> non-zero at scale (sample was too small to surface them); still sparse, conservative-but-incomplete. Top Walton DirectRef schemes well-represented: Bias_Inference 262, Sign_Inference 119, Ethotic_Inference 117, EvidenceToHypothesis 94. WARN = 87 nodes (7%), ALL bad_map:* (mappingType outside the observed {broad,close,narrow}Match set: relatedMatch 78, none 5, exactMatch 2, noMatch 2). These are LEGIT SKOS predicates -> a schema-extension decision for the expert gate, NOT fabrications. none/noMatch = hedge candidates to DROP during ratification. Sidecars: 141-aif-candidates-fullscale.csv (1.1 MB, flat, reviewable, committed) + report. The structured JSON (1.4 MB) is regenerable via --finalize from the checkpoint and is NOT committed (repo already 2.05 GiB per #415/#621 — limit additive weight). 1 node (6.3.1.2.3.1.4 Hypocrisie morale) failed once on an API read- timeout, recovered on a 1-node resume (0 content/parse errors). Scope: docs/taxonomy/141-aif-candidates-fullscale.csv + report only. 0 write Cards/, 0 AssetConverter code change. Base 0e39425. Next (gated post-release): Stage 2 dry-run diff, Stage 3 expert/jsboige ratification gate (decide relatedMatch/exactMatch schema extension, drop none/noMatch hedges, prioritize >0.8 candidates), Stage 4 -> #130 OWL / #136 2sxc. Relates to #141, #609, #620, #130, #136, #192. Co-Authored-By: Claude-Code <noreply@anthropic.com>
…m / 5 silent Read-only diff of the Stage-1 sidecar (#626) against the 12 non-card nodes that already carry an AIF value on the taxonomy (0.97% — the other 1220 are entirely net-new, no conflict possible). High-signal prep for the expert gate: - 2 CONFIRM: gpt-5.5 independently re-derived the exact expert token (Ignorance_Inference, VagueVerbalClassification_Inference) — closed-set produces real Walton mappings. - 9 CONFLICT: 1 field-swap/churn node (Pente glissante) + 5 frank token disagreements → gate must adjudicate, not auto-apply. - 5 SILENT: generator stayed silent on a filled field → gate must preserve existing values. Adds 141-aif-stage2-diff.py (read-only, 0 Cards/ write), 16-row diff.csv, and a Stage-2 section to the report. Same scope/honors pre-tag freeze. Co-Authored-By: Claude-Code <noreply@anthropic.com>
jsboige
added a commit
that referenced
this pull request
Jul 1, 2026
…commendation (#627) * docs(taxonomy): #141 Stage-3 expert adjudication package + closure recommendation Stage-3 (PRIMAIRE dispatch ai-01): one-pass adjudication package for the 12 non-card nodes that already carry an AIF value (0.97% of 1232). Sourced from the reproducible stage2-diff.csv (anti-fab): 2 CONFIRM (ratify), 9 CONFLICT across 6 nodes (advisory recos with Walton reasoning, expert decides), 5 SILENT (preserve existing). No auto-apply — every row is a human gate call. Surfaces a cross-cutting scheme-vs-conflict tier convention (_Inference/Scheme -> DirectRef, _Conflict -> ExceptionRef) for the gate to adopt or reject. Closure (TERTIAIRE): synthesis of Phases 1-3 (census #609 + full-scale #626 + Stage-2/3) -> recommend #141 closed as delivered-and-gated, with the 3 close criteria for the expert gate. Read-only docs, 0 Cards/ write, 0 AssetConverter change. Honors pre-tag freeze. Co-Authored-By: Claude-Code <noreply@anthropic.com> * docs(taxonomy): #141 WARN bad_map:* triage — 87 nodes (80 adopt / 7 drop) Completes the #141 expert-gate package (the schema-extension decision, alongside Stage-3 adjudication + closure already in this PR). Triages the 87 non-card nodes whose mappingType fell outside the observed {broad,close,narrow}Match set: - skos:relatedMatch 78 + skos:exactMatch 2 -> ADOPT (extend observed set; legit SKOS, enriches semantics at zero fabrication risk; exactMatch spot-verify). - none 5 + skos:noMatch 2 -> DROP (weak hedges; leave mappingType empty). 80 ADOPT / 7 DROP. Confirms these are NOT fabrications (0 across 1232 nodes) — a schema call for the gate. Sourced from the structured JSON (source of truth for WARNs — the flat CSV undercounts by ~23 nodes that have a mappingType but 0 DirectRef/ExceptionRef rows). Reproducible via 141-aif-warn-triage.py -> 87-row reviewable CSV. Read-only docs+tool, 0 Cards/ write, 0 AssetConverter change. Honors pre-tag freeze. Co-Authored-By: Claude-Code <noreply@anthropic.com> --------- Co-authored-by: Your <your.email@example.com> Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige
added a commit
that referenced
this pull request
Jul 4, 2026
…oposal + sample, gated) (#673) Secondaire dispatch l0wt63 (ai-01). Additive proposal for #141's remaining engineering artifact: the crossLink DatasetUpdater task that lets the Stage-3 expert-gate output flow into Cards/ drift-free. State recap (code=truth): #141 text-enrichment DONE (100% x8 langs), AIF cross-ref DELIVERED through Stage-3 (#626: 1232/1232, 0 fab), 'adapt the GPT-4 script' DONE (DatasetUpdater = modern counterpart, SDK v2.10.0/gpt-5.5). Closure-rec says remaining work is expert-gate judgment, not engineering. This stages the ONE engineering artifact that follows the gate. Proposal contents (text in doc, NOT materialized files): - Task config skeleton (Enabled=false, mirrors the 7 existing configs) - UseFunctionCalling=true (closed-set anti-fab, the #626-validated design) - SkipNonEmpty=true (preserves the 12 expert-adjudicated existing-AIF nodes) - ChunkSize=1 (per-node relationship judgement) - Prompt design (PromptCrossLinkSystem sample framing, closed verb+path set) - Sample (echantillon) drawn from committed sidecar 141-aif-candidates-sample.csv - Staged flow gated post-tag + post-Stage-3-ratification Gated: 0 Cards/ write, 0 AssetConverter C# change (pre-tag freeze honored per closure-rec). Same method as scale-ups #497-499 (proposal doc + sample first, prod-write separate gated step). Relates #141 #626 #130 #136 #498 #595. Base a41cbda. Co-authored-by: Your <your.email@example.com> Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige
pushed a commit
that referenced
this pull request
Jul 4, 2026
update) Idle tick follow-up on #677. Adds a "Target-existence validation" section confirming the second axis of #626 0-fabrication claim: every one of the 3850 crossLinks targets actually exists in the prod taxonomy (0/3850 orphans, 0/1232 nodes), and all #626 source_dps are present in prod (0/1232 missing). Independent code=truth verification of the closed-set design quality bar. Includes a format-conversion anti-false-finding note (naive comma->dot replace produces a spurious 74% orphan rate; correct conversion splits each digit). Co-Authored-By: Claude-Code <noreply@anthropic.com>
jsboige
added a commit
that referenced
this pull request
Jul 4, 2026
…ated from #626) (#677) * docs(taxonomy): #141 crossLink enrichment sample run (real output curated from #626) Primaire of dispatch `lofjtd` (ai-01). Concrete before/after sample for the #141 crossLink enrichment staged by proposal #673. The enrichment was already run in #626 (fullscale gpt-5.5 closed-set pass over 1232 non-card nodes, 3850 candidate links, 0 fabrication warnings). Rather than re-spend gpt-5.5 credits to reproduce existing output, this doc curates a 7-node representative sample from #626's real output and shows before/after, quality, and cost. - Before (prod taxonomy CSV): 11242/11264 crossLink_* cells empty (99.8%) - Sample: 7 non-card nodes, depth 0-7, confidence 0.38-0.98 - Quality (21 links): mean conf 0.70, 11/21 ratifiable (>=0.70), 0 fab warnings - Surfaces the multi-target serialization question for the Stage-3 expert gate (5/7 nodes have >=1 verb with multiple candidates; cell convention = 1 target) 0 prod Cards/ write, 0 AssetConverter C# change, 0 fresh gpt-5.5 call, 0 auto-ratification. Gated post-tag + post-Stage-3. Co-Authored-By: Claude-Code <noreply@anthropic.com> * docs(taxonomy): #141 add target-existence validation to sample-run (#677 update) Idle tick follow-up on #677. Adds a "Target-existence validation" section confirming the second axis of #626 0-fabrication claim: every one of the 3850 crossLinks targets actually exists in the prod taxonomy (0/3850 orphans, 0/1232 nodes), and all #626 source_dps are present in prod (0/1232 missing). Independent code=truth verification of the closed-set design quality bar. Includes a format-conversion anti-false-finding note (naive comma->dot replace produces a spurious 74% orphan rate; correct conversion splits each digit). Co-Authored-By: Claude-Code <noreply@anthropic.com> --------- Co-authored-by: Your <your.email@example.com> Co-authored-by: Claude-Code <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
PRIMAIRE of ai-01 deep-queue supersede #3 (
msg-…v95b6l): the full-scale AIF Stage-1 candidate sidecar + coverage/WARN report. The generator shipped in #623; this is the run output — 1232/1232 non-card nodes, gpt-5.5 assist, dry-run, 0 writeCards/.Headline — anti-fab closed-set HOLDS at full scale
Across 1232 nodes, there are zero fabricated scheme names, zero invented decimal_paths, zero out-of-vocab verbs. The closed-set design (targets ∈ enumerated real nodes, AIF tokens ∈ the 60-token observed Walton vocab) contains LLM fabrication completely.
crossLink verb distribution
IsRelatedTo42 % ·Leverages33 % ·Mirrors15 % ·Allows7 % ·Inverts2 % ·Opposes1 % ·PredatesOn1 % ·Denounces<1 %.Honest improvement vs pilot:
Opposes/PredatesOn/Denounceswere 0 in the 28-node pilot (flagged limitation) → non-zero at scale (38/33/1) — the sample was too small to surface them; still sparse, conservative-but-incomplete.Top Walton DirectRef schemes well-represented:
Bias_Inference262,Sign_Inference119,Ethotic_Inference117.WARN — schema decision for the expert gate (not fabrications)
87 nodes (7 %), all
bad_map:*— mappingTypes outside the observed{broad,close,narrow}Matchset:relatedMatch78,none5,exactMatch2,noMatch2. These are legitimate SKOS predicates; the validator is strict-by-construction. The gate decides whether to extend the observed set (relatedMatch/exactMatch) or down-grade them tocloseMatch.none/noMatch= hedges to drop during ratification.Sidecars
141-aif-candidates-fullscale.csv141-aif-candidates-fullscale.json--finalizeThe structured JSON is regenerable on demand from the checkpoint; omitted from the commit to limit repo weight (#415/#621 — repo is 2.05 GiB, 100 % historical).
Safety / scope
docs/taxonomy/only — sidecar CSV + report. 0 writeCards/, 0 AssetConverter code change (pre-tag safe). Base0e394255.6.3.1.2.3.1.4) failed once on an API read-timeout, recovered on a 1-node resume (0 content/parse errors).Next (gated post-release)
Stage 2 dry-run diff · Stage 3 expert/jsboige ratification gate (decide
relatedMatch/exactMatchschema extension, dropnone/noMatch, prioritize >0.8 candidates) · Stage 4 → #130 OWL / #136 2sxc.Relates to #141, #609, #620, #130, #136, #192.