Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
121 changes: 121 additions & 0 deletions docs/decisions/2026-10-04-ns2604-verified-e2e.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,121 @@
# NS2604 verified foundation E2E publication

Date: 2026-10-04. Scope: sanitized historical repository evidence.

## Context and decision

The user asked to publish the verified new-host E2E so future sessions can recover
its status and fixes, without raw conversations, credentials, local user identities
or concrete home/profile paths. The machine label NativeStack2604 is retained
consistently. This bounded job leaves changes in the working tree and a commit
message for the coordinator; manifests/evidence.json remains coordinator-owned.

Publish the 15-unit, 80-slot status chain as dated evidence. Use the independent
reviews and adjudications rather than treating executor labels or installation
logs as final readiness. The north-star action is to make the native foundation's
remaining work recoverable for complex systems, US-equities research and
historical simulation.

The aggregation source is the supplied e2e_merge.py and review-prompt.txt, with
source-file hashes retained in the
[receipt](../../evidence/receipts/ns2604-e2e-20261004.json).
Repository format and claim boundaries follow
[the acceptance-evidence policy](../acceptance-evidence-policy.md) and the recent
[RTK qualification receipt](../../evidence/receipts/rtk-051-qualification-20261004.json).
The publication base is native-agent-stack revision
5358bb564fc42edb54a1edf7b848bf07d7f0997a.

## Method and result

Claude Sonnet executors recorded native host observations. Independent GPT Sol
reviewers judged one unit each; 22 Claude Opus adjudications resolved all 22
executor/reviewer disagreements: 14 kept the reviewer's status and 8 took the
executor's. The merge uses the schema-checked executor
labels, preserves each independent review, and lets adjudication override the
final label. Agreement still describes executor versus reviewer.

Recorded executor model metadata names claude-sonnet-5-5. All 15 reviewer result
records name cx/gpt-6.1-sol-max at requested_effort=max. Workflow runs
wf_8bf450a9-d1e (17 agents) and wf_cdbf70a0-02c (5 agents) identify the adjudicator
as claude-opus-5-5, with effort max and agent type evidence-reviewer requested by
the workflow script. Their starts were 2026-10-04T20:06:45Z and
2026-10-04T20:08:29Z; completions were 2026-10-04T20:20:30Z and
2026-10-04T20:19:31Z, respectively. The prompt template as run ships in this PR. The
[method](../../evidence/artifacts/ns2604-e2e-20261004/method.md) distinguishes these
provenance levels and dates.

| Final status | Slots |
| --- | ---: |
| READY | 18 |
| PARTIAL | 31 |
| FAIL | 2 |
| BY_DESIGN | 12 |
| INTERIM | 17 |
| UNJUDGED | 0 |
| Total | 80 |

**(READY + BY_DESIGN) / 80 = (18 + 12) / 80 = 37.5%, rounded half up to 38%.**
Every slot remains in the denominator. BY_DESIGN counts an intentional exclusion;
INTERIM is held by design and contributes zero to readiness. The
[summary](../../evidence/artifacts/ns2604-e2e-20261004/summary.json) retains the
per-layer table, and [slots](../../evidence/artifacts/ns2604-e2e-20261004/slots.json)
retain the individual status chains, bounded reasons, fixes and citations.

This is native executor observation, independent model review and adjudication,
not an upstream benchmark or a fresh run by the publisher. Host state is as of
**2026-10-04 19:10–20:10Z**; the final merge is stamped 20:26:16Z. Fixes are pending.

## Fix plan and handoff

The current fix plan is **25 repository-plan and 8 host follow-ups** over the
33 open PARTIAL/FAIL slots after adjudication. The earlier 30 plan / 11 host
figure covered 41 reviewer-stage slots before adjudication.
Keep the 17 INTERIM slots held until their
existing gates are satisfied, then run their native acceptance and session-use
checks. Repair plan-backed installation or wiring gaps, apply the host follow-ups,
and retain actual failed and passing outcomes for independent review before
changing a slot's status.

The [fix list](../../evidence/artifacts/ns2604-e2e-20261004/fixes.json) ships in this
PR. Four INTERIM token slots are held for [PR #684](https://github.com/seathatflowsinourveins/native-agent-stack/pull/684):
context-supply, command-output, output-compression and code-index. PR #684 has not
merged. The two open token slots have separate fixes: ccusage (FAIL, plan) and
statusline (PARTIAL, host).

The [adjudication prompt record](../../evidence/artifacts/ns2604-e2e-20261004/adjudication-prompt.md)
contains the template as run, with run-specific placeholders. The receipt retains
the adjudicator model, run ids, timestamps and private record hashes. Registry
integration for all eight published files, with the hashes and mirrored fields
this bounded repair changed, is the PR's final registry-only commit, and
validate.py passes at that head.

This repair corrects publication contradictions against adjudication.json,
fixes.json, both workflow run records and the final 80-slot merge. The completeness
critic covers the reviewer's/executor's 14/8 split, the four held token slots,
the two separate open token fixes and the published prompt/fix-list handoff.
It preserves technical strings while masking concrete home/profile paths and
local user identities. Source inspection and structural checks supply the repair
evidence; fresh host replication and upstream benchmark acceptance remain absent.

## Alternatives and overturn condition

The publication brief reports an earlier log-based readiness estimate of **62%**
with weaker criteria. Its underlying log corpus is not supplied here, so that
number is retained as reported comparison context. Installation or invocation
logs can support part of a claim without proving upstream acceptance, native
wiring and fresh client-session use. The present 38% uses the stricter retained
reviews and final adjudications.

Alternatives were to preserve that 62% log estimate as final readiness, publish
executor labels without review, or publish only aggregate counts. The status
chain was selected because it preserves the stricter evidence and the actual
remaining work without exposing raw sessions.

A new dated qualification would overturn the current readiness result when it
supplies native passing acceptance, required wiring and fresh session use for
the same 80 slots, followed by independent review and adjudication. Recovering
original contradictory evidence could also justify a source-cited correction.
Merging #684 or applying fixes alone does not change this historical receipt.
A changed denominator needs an explicitly different scope and comparison.
Provenance corrections update publication completeness without changing historical
host readiness. Preserve this record and append later evidence.
22 changes: 22 additions & 0 deletions evidence/artifacts/ns2604-e2e-20261004/adjudication-prompt.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Adjudication prompt (Opus, per disputed slot)

Model: `opus` (Claude Opus 5.5) at effort `max`, agent type `evidence-reviewer` (read-only), one agent per disputed slot,
in two Claude Code workflow runs on 2026-10-04 (17 slots, then 5 more after the executor statuses were corrected from the
workflow's schema-checked returns). Schema-forced return: slot_id, final_status, side, reason, fix, fix_kind, needs_user,
sources. The template below is the prompt as run, with run-specific values shown as <placeholders>.

```
Adjudicate one slot of the NativeStack2604 live E2E (date 2026-10-04). Read-only; never run anything on any WSL distribution; never read credential files or values.

Slot: <slot_id> (unit <unit>). The Sonnet executor proposed <executor_status>; the independent GPT Sol reviewer judged <reviewed_status>.

Status definitions (from the review brief): READY = present at version, its upstream acceptance passes, wired natively, and used in a fresh headless session if the slot is client-facing. PARTIAL = some of these hold. FAIL = absent or broken. BY_DESIGN = nothing to install by design. INTERIM = an explicitly held interim the plan records (for example a token tool awaiting the token-layer PR #684); an INTERIM must not hide a missing install the current plan says should exist.

Read: the executor evidence <e2e dir>/<unit>.json (this slot's record: commands, exit codes, output excerpts), the GPT review <e2e dir>/review-<unit>.json (this slot), the slot entry in <e2e dir>/units.json, and the install plan row in <host path> (origin/main).

Executor summary: <executor_summary>
Reviewer reason: <review_reason>
Reviewer fix: <fix>

Decide the status the evidence supports today, on the definitions above. Side with whichever is right, or neither. Then give the fix that would make it READY (or confirm none): exact upstream command or plan change, whether it is a repository plan change or host state, and anything that needs the user (a sign-in, an identity, a policy choice), else an empty string. Cite file:line in the evidence or plan and any upstream doc the reviewer cited.
```
Loading
Loading