Skip to content

docs(taxonomy): #141 AIF Stage-1 full-scale results — 1232 nodes, 0 fabrication - #626

Merged
jsboige merged 2 commits into
masterfrom
docs/141-aif-fullscale-results
Jul 1, 2026
Merged

docs(taxonomy): #141 AIF Stage-1 full-scale results — 1232 nodes, 0 fabrication#626
jsboige merged 2 commits into
masterfrom
docs/141-aif-fullscale-results

Conversation

@jsboige

@jsboige jsboige commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

Summary

PRIMAIRE of ai-01 deep-queue supersede #3 (msg-…v95b6l): the full-scale AIF Stage-1 candidate sidecar + coverage/WARN report. The generator shipped in #623; this is the run output — 1232/1232 non-card nodes, gpt-5.5 assist, dry-run, 0 write Cards/.

Headline — anti-fab closed-set HOLDS at full scale

Metric #620 pilot (28) Full-scale (1232)
Non-card nodes covered 28 1232 (100 %)
crossLink candidates 93 3850 (3.12/node · 100 % nodes)
AIF scheme refs 55 2310 (1.88/node · 96 % nodes)
Mean confidence 0.75 0.73 (median 0.72)
Fabricated tokens 0 0
Invalid targets / verbs 0 / 0 0 / 0

Across 1232 nodes, there are zero fabricated scheme names, zero invented decimal_paths, zero out-of-vocab verbs. The closed-set design (targets ∈ enumerated real nodes, AIF tokens ∈ the 60-token observed Walton vocab) contains LLM fabrication completely.

crossLink verb distribution

IsRelatedTo 42 % · Leverages 33 % · Mirrors 15 % · Allows 7 % · Inverts 2 % · Opposes 1 % · PredatesOn 1 % · Denounces <1 %.

Honest improvement vs pilot: Opposes/PredatesOn/Denounces were 0 in the 28-node pilot (flagged limitation) → non-zero at scale (38/33/1) — the sample was too small to surface them; still sparse, conservative-but-incomplete.

Top Walton DirectRef schemes well-represented: Bias_Inference 262, Sign_Inference 119, Ethotic_Inference 117.

WARN — schema decision for the expert gate (not fabrications)

87 nodes (7 %), all bad_map:* — mappingTypes outside the observed {broad,close,narrow}Match set: relatedMatch 78, none 5, exactMatch 2, noMatch 2. These are legitimate SKOS predicates; the validator is strict-by-construction. The gate decides whether to extend the observed set (relatedMatch/exactMatch) or down-grade them to closeMatch. none/noMatch = hedges to drop during ratification.

Sidecars

File Size Committed?
141-aif-candidates-fullscale.csv 1.1 MB ✅ flat, reviewable
141-aif-candidates-fullscale.json 1.4 MB ❌ tmp/, regenerable via --finalize

The structured JSON is regenerable on demand from the checkpoint; omitted from the commit to limit repo weight (#415/#621 — repo is 2.05 GiB, 100 % historical).

Safety / scope

  • docs/taxonomy/ only — sidecar CSV + report. 0 write Cards/, 0 AssetConverter code change (pre-tag safe). Base 0e394255.
  • ✅ 1 node (6.3.1.2.3.1.4) failed once on an API read-timeout, recovered on a 1-node resume (0 content/parse errors).

Next (gated post-release)

Stage 2 dry-run diff · Stage 3 expert/jsboige ratification gate (decide relatedMatch/exactMatch schema extension, drop none/noMatch, prioritize >0.8 candidates) · Stage 4#130 OWL / #136 2sxc.

Relates to #141, #609, #620, #130, #136, #192.

Your and others added 2 commits July 1, 2026 10:41
…abrication

PRIMAIRE of ai-01 deep-queue supersede #3 (msg-...v95b6l): the full-scale
AIF Stage-1 candidate sidecar + coverage/WARN report. The generator was
shipped in #623; this is the run output (1232/1232 non-card nodes,
gpt-5.5 assist, dry-run, 0 write Cards/).

Headline — anti-fab closed-set HOLDS at full scale:
  - 1232 nodes covered (100% of Fallacies non-cards, the #609 gap).
  - 3850 crossLinks (3.12/node, 100% nodes) + 2310 AIF refs (1.88/node,
    96% nodes).
  - confidence mean 0.73, median 0.72 (1301 high >=0.8 / 2460 mid / 89 low).
  - FABRICATION = 0. invalid targets = 0. invalid verbs = 0.
    The closed-set design (targets from enumerated real nodes, AIF tokens
    from the 60-token observed Walton vocab) contains LLM fabrication
    completely across 1232 nodes.

crossLink verb mix: IsRelatedTo 42% / Leverages 33% / Mirrors 15% /
Allows 7% / Inverts 2% / Opposes 1% / PredatesOn 1% / Denounces <1%.
Honest improvement vs pilot: Opposes/PredatesOn/Denounces were 0 in the
28-node pilot (flagged limitation) -> non-zero at scale (sample was too
small to surface them); still sparse, conservative-but-incomplete.

Top Walton DirectRef schemes well-represented: Bias_Inference 262,
Sign_Inference 119, Ethotic_Inference 117, EvidenceToHypothesis 94.

WARN = 87 nodes (7%), ALL bad_map:* (mappingType outside the observed
{broad,close,narrow}Match set: relatedMatch 78, none 5, exactMatch 2,
noMatch 2). These are LEGIT SKOS predicates -> a schema-extension
decision for the expert gate, NOT fabrications. none/noMatch = hedge
candidates to DROP during ratification.

Sidecars: 141-aif-candidates-fullscale.csv (1.1 MB, flat, reviewable,
committed) + report. The structured JSON (1.4 MB) is regenerable via
--finalize from the checkpoint and is NOT committed (repo already
2.05 GiB per #415/#621 — limit additive weight).

1 node (6.3.1.2.3.1.4 Hypocrisie morale) failed once on an API read-
timeout, recovered on a 1-node resume (0 content/parse errors).

Scope: docs/taxonomy/141-aif-candidates-fullscale.csv + report only.
0 write Cards/, 0 AssetConverter code change. Base 0e39425.

Next (gated post-release): Stage 2 dry-run diff, Stage 3 expert/jsboige
ratification gate (decide relatedMatch/exactMatch schema extension,
drop none/noMatch hedges, prioritize >0.8 candidates), Stage 4 ->
#130 OWL / #136 2sxc.

Relates to #141, #609, #620, #130, #136, #192.

Co-Authored-By: Claude-Code <noreply@anthropic.com>
…m / 5 silent

Read-only diff of the Stage-1 sidecar (#626) against the 12 non-card nodes
that already carry an AIF value on the taxonomy (0.97% — the other 1220 are
entirely net-new, no conflict possible). High-signal prep for the expert gate:
- 2 CONFIRM: gpt-5.5 independently re-derived the exact expert token
  (Ignorance_Inference, VagueVerbalClassification_Inference) — closed-set
  produces real Walton mappings.
- 9 CONFLICT: 1 field-swap/churn node (Pente glissante) + 5 frank token
  disagreements → gate must adjudicate, not auto-apply.
- 5 SILENT: generator stayed silent on a filled field → gate must preserve
  existing values.

Adds 141-aif-stage2-diff.py (read-only, 0 Cards/ write), 16-row diff.csv, and
a Stage-2 section to the report. Same scope/honors pre-tag freeze.

Co-Authored-By: Claude-Code <noreply@anthropic.com>
@jsboige
jsboige merged commit 01e7fdc into master Jul 1, 2026
3 checks passed
@jsboige
jsboige deleted the docs/141-aif-fullscale-results branch July 1, 2026 10:44
jsboige added a commit that referenced this pull request Jul 1, 2026
…commendation (#627)

* docs(taxonomy): #141 Stage-3 expert adjudication package + closure recommendation

Stage-3 (PRIMAIRE dispatch ai-01): one-pass adjudication package for the 12
non-card nodes that already carry an AIF value (0.97% of 1232). Sourced from
the reproducible stage2-diff.csv (anti-fab): 2 CONFIRM (ratify), 9 CONFLICT
across 6 nodes (advisory recos with Walton reasoning, expert decides), 5 SILENT
(preserve existing). No auto-apply — every row is a human gate call. Surfaces a
cross-cutting scheme-vs-conflict tier convention (_Inference/Scheme -> DirectRef,
_Conflict -> ExceptionRef) for the gate to adopt or reject.

Closure (TERTIAIRE): synthesis of Phases 1-3 (census #609 + full-scale #626 +
Stage-2/3) -> recommend #141 closed as delivered-and-gated, with the 3 close
criteria for the expert gate.

Read-only docs, 0 Cards/ write, 0 AssetConverter change. Honors pre-tag freeze.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

* docs(taxonomy): #141 WARN bad_map:* triage — 87 nodes (80 adopt / 7 drop)

Completes the #141 expert-gate package (the schema-extension decision, alongside
Stage-3 adjudication + closure already in this PR). Triages the 87 non-card
nodes whose mappingType fell outside the observed {broad,close,narrow}Match set:

- skos:relatedMatch 78 + skos:exactMatch 2 -> ADOPT (extend observed set; legit
  SKOS, enriches semantics at zero fabrication risk; exactMatch spot-verify).
- none 5 + skos:noMatch 2 -> DROP (weak hedges; leave mappingType empty).

80 ADOPT / 7 DROP. Confirms these are NOT fabrications (0 across 1232 nodes) —
a schema call for the gate. Sourced from the structured JSON (source of truth
for WARNs — the flat CSV undercounts by ~23 nodes that have a mappingType but
0 DirectRef/ExceptionRef rows). Reproducible via 141-aif-warn-triage.py ->
87-row reviewable CSV.

Read-only docs+tool, 0 Cards/ write, 0 AssetConverter change. Honors pre-tag freeze.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

---------

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 4, 2026
…oposal + sample, gated) (#673)

Secondaire dispatch l0wt63 (ai-01). Additive proposal for #141's remaining
engineering artifact: the crossLink DatasetUpdater task that lets the Stage-3
expert-gate output flow into Cards/ drift-free.

State recap (code=truth): #141 text-enrichment DONE (100% x8 langs), AIF
cross-ref DELIVERED through Stage-3 (#626: 1232/1232, 0 fab), 'adapt the GPT-4
script' DONE (DatasetUpdater = modern counterpart, SDK v2.10.0/gpt-5.5). Closure-rec
says remaining work is expert-gate judgment, not engineering. This stages the ONE
engineering artifact that follows the gate.

Proposal contents (text in doc, NOT materialized files):
- Task config skeleton (Enabled=false, mirrors the 7 existing configs)
  - UseFunctionCalling=true (closed-set anti-fab, the #626-validated design)
  - SkipNonEmpty=true (preserves the 12 expert-adjudicated existing-AIF nodes)
  - ChunkSize=1 (per-node relationship judgement)
- Prompt design (PromptCrossLinkSystem sample framing, closed verb+path set)
- Sample (echantillon) drawn from committed sidecar 141-aif-candidates-sample.csv
- Staged flow gated post-tag + post-Stage-3-ratification

Gated: 0 Cards/ write, 0 AssetConverter C# change (pre-tag freeze honored per
closure-rec). Same method as scale-ups #497-499 (proposal doc + sample first,
prod-write separate gated step).

Relates #141 #626 #130 #136 #498 #595. Base a41cbda.

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
jsboige pushed a commit that referenced this pull request Jul 4, 2026
 update)

Idle tick follow-up on #677. Adds a "Target-existence validation" section
confirming the second axis of #626 0-fabrication claim: every one of the 3850
crossLinks targets actually exists in the prod taxonomy (0/3850 orphans, 0/1232
nodes), and all #626 source_dps are present in prod (0/1232 missing).

Independent code=truth verification of the closed-set design quality bar.
Includes a format-conversion anti-false-finding note (naive comma->dot replace
produces a spurious 74% orphan rate; correct conversion splits each digit).

Co-Authored-By: Claude-Code <noreply@anthropic.com>
jsboige added a commit that referenced this pull request Jul 4, 2026
…ated from #626) (#677)

* docs(taxonomy): #141 crossLink enrichment sample run (real output curated from #626)

Primaire of dispatch `lofjtd` (ai-01). Concrete before/after sample for the
#141 crossLink enrichment staged by proposal #673.

The enrichment was already run in #626 (fullscale gpt-5.5 closed-set pass over
1232 non-card nodes, 3850 candidate links, 0 fabrication warnings). Rather than
re-spend gpt-5.5 credits to reproduce existing output, this doc curates a
7-node representative sample from #626's real output and shows before/after,
quality, and cost.

- Before (prod taxonomy CSV): 11242/11264 crossLink_* cells empty (99.8%)
- Sample: 7 non-card nodes, depth 0-7, confidence 0.38-0.98
- Quality (21 links): mean conf 0.70, 11/21 ratifiable (>=0.70), 0 fab warnings
- Surfaces the multi-target serialization question for the Stage-3 expert gate
  (5/7 nodes have >=1 verb with multiple candidates; cell convention = 1 target)

0 prod Cards/ write, 0 AssetConverter C# change, 0 fresh gpt-5.5 call, 0
auto-ratification. Gated post-tag + post-Stage-3.

Co-Authored-By: Claude-Code <noreply@anthropic.com>

* docs(taxonomy): #141 add target-existence validation to sample-run (#677 update)

Idle tick follow-up on #677. Adds a "Target-existence validation" section
confirming the second axis of #626 0-fabrication claim: every one of the 3850
crossLinks targets actually exists in the prod taxonomy (0/3850 orphans, 0/1232
nodes), and all #626 source_dps are present in prod (0/1232 missing).

Independent code=truth verification of the closed-set design quality bar.
Includes a format-conversion anti-false-finding note (naive comma->dot replace
produces a spurious 74% orphan rate; correct conversion splits each digit).

Co-Authored-By: Claude-Code <noreply@anthropic.com>

---------

Co-authored-by: Your <your.email@example.com>
Co-authored-by: Claude-Code <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant