feat(ava): establish controlled Hermes runtime foundation - #1
Conversation
M4-SHADOW-PREFLIGHT review — BLOCKED correctly; authorize M4-NETWORK-GATE onlyClosure
Bubblewrap alone cannot simultaneously provide OpenAI/Telegram egress and prevent access to host loopback, LAN and arbitrary Internet destinations. Do not launch the shadow with host networking and do not weaken this gate. Authorized operation:
|
M4-NETWORK-GATE review — authorize M4-NETWORK-GATE-V2 onlyLe premier design est Une voie sans installation ni privilège doit maintenant être testée : Autorisation limitéeExécuter uniquement Contraintes :
Preuve demandée :
Rapport requis : Ne pas affaiblir l’isolation ni partager le réseau hôte en cas d’échec. Aucun shadow réel n’est autorisé. Stop après le rapport. |
M4-NETWORK-GATE-V2 closure — PASS; authorize M4-PROXY-COMPAT-PREFLIGHT onlyM4-NETWORK-GATE-V2 is accepted as PASS. The validated architecture provides loopback-only networking inside bubblewrap, a Unix-socket egress bridge, a loopback TCP relay, strict CONNECT allowlisting, blocked direct DNS/network/host-loopback access, and complete cleanup with no live or candidate mutation. Authorized single operation:
|
M4-PROXY-COMPAT-PREFLIGHT review — PARTIAL PASS; authorize M4-CLIENT-PROXY-HARNESS onlyThe inspection closes the static compatibility question for Codex inference and OpenAI OAuth refresh. Telegram also supports an explicit proxy, but two runtime facts remain open: the messaging extra is not installed in the candidate environment, and Telegram contains direct DoH/IP fallback paths that must be proven inert inside the V2 network gate. Authorize M4-CLIENT-PROXY-HARNESS only, under Luna medium, without subagent. Disposable environment
Synthetic client harnessReuse the proven V2 architecture: Synthetic authorities only: Use dummy credentials only. Redirect endpoint constants in memory or in the disposable harness; do not edit candidate source. Run exactly these paths:
Required environment closure: Required assertions:
Return only: No shadow launch, real endpoint test, source patch, commit, merge, service action, or promotion is authorized. |
M4-CLIENT-PROXY-HARNESS review — network gate preserved; authorize M4-CODEX-PROTOCOL-TRACE onlyThe observed Authorize M4-CODEX-PROTOCOL-TRACE only, under Luna medium, no sub-agent. ScopeRecreate a fresh locked disposable environment and the already-proven M4 network gate V2. Test only the OpenAI SDK Codex client against one synthetic TLS endpoint. Do not run OAuth or Telegram in this phase. Required trace
ForbiddenReportOAuth and Telegram remain suspended. Do not relaunch the combined client harness. |
M4-CODEX-PROTOCOL-TRACE closure — PASS; authorize M4-OAUTH-PROTOCOL-TRACE onlyThe prior Codex failure is classified as a synthetic-fixture defect. Do not rerun the Codex client. Authorized single operation: M4-OAUTH-PROTOCOL-TRACERun under Luna medium, without sub-agent. Use the already proven network gate V2 and a fresh disposable environment. Exercise only Required boundaries:
Required proofs: The harness must prove that a successful synthetic refresh updates only the two disposable auth states in the expected schema and that a rejected/malformed synthetic response fails closed without corrupting either file. Maximum three bounded response variants. No Telegram retest and no shadow launch are authorized. |
M4-OAUTH-PROTOCOL-TRACE interruption review — BLOCKED safely; authorize M4-GATE-V2-RECOVERY onlyThe interruption does not invalidate any prior proof. Do not rerun M2, M3, the Codex client, OAuth, or Telegram in this operation. Authorized single operation: M4-GATE-V2-RECOVERYRun under GPT-5.6 Luna medium, without sub-agent.
Required report: No OAuth, Telegram, model/provider call, gateway, cron, systemd, Docker, source change, merge, promotion, or upstream rebase is authorized. |
M4-GATE-V2-RECOVERY closure — PASS; authorize M4-OAUTH-PROTOCOL-TRACE onlyThe previous interruption is fully repaired. Earlier M2/M3/M4 proofs remain valid. Do not rerun Codex or prior gates. Authorized single operation: M4-OAUTH-PROTOCOL-TRACERun under GPT-5.6 Luna medium, without sub-agent. Required setup:
Exercise only Required bounded variants, maximum three:
Required report: Forbidden:
Stop after the report. |
M4-OAUTH-PROTOCOL-TRACE interruption review — BLOCKED safely; authorize M4-OAUTH-TLS-FIXTURE onlyThis is a missing synthetic protocol fixture, not a network-confinement failure. Do not modify or replace the validated gate V2 core. Authorized single operation: M4-OAUTH-TLS-FIXTURERun under GPT-5.6 Luna medium, without sub-agent. Create a separate persistent, non-secret OAuth TLS fixture beside the gate, for example: Required boundaries:
Required self-test:
Report: No OAuth protocol trace or shadow launch is authorized yet. |
Sol High integration sweep authorization — finish synthetic closure, prepare shadow, stop before real stateThis replaces the overly granular one-micro-phase cadence. The operator may complete the remaining non-destructive synthetic validation chain in one continuous integration sweep, while preserving evidence after each internal gate. Mission A — compose the durable harnesses
Mission B — close OAuth synthetic protocolAgainst
Mission C — close Telegram synthetic protocolUsing a disposable locked
Mission D — offline shadow readinessAfter A-C pass:
Optional audit assistanceSol High may use at most two short read-only counterchecks at depth 1 only when a named contradiction remains. No autonomous implementation branch, broad repository refactor, or speculative source change. Hard stop boundaryThe sweep MUST stop before any of the following:
Final reportReturn one consolidated report with: Stop after that report. Real-state execution remains subject to SE7H/ChatGPT audit. |
M4-SOL-HIGH-INTEGRATION-SWEEP closure — PASS; authorize unattended M5 shadow sequence under bounded auto-reviewOperator/model decision
Permission decisionForbidden on the living Minisforum host: Required unattended posture: approval_policy = "on-request"
approvals_reviewer = "auto_review"
sandbox_mode = "workspace-write"Use a dedicated trusted shadow workspace and grant only the additional directories explicitly needed. Do not broaden to the user's full HOME. Keep the external bubblewrap/network-gate boundary authoritative. Authorized operation: M5-UNATTENDED-SHADOW-SEQUENCERun as one coherent Sol High sequence. Stop immediately and report if any invariant fails; do not improvise broader access. Stage 0 — autonomy and boundary proof
Stage 1 — isolated real provider shadow
Stage 2 — isolated real Telegram shadowOnly after Stage 1 passes:
Stage 3 — disposable cron/state/restart proofOnly after Stage 2 passes:
Mandatory final state
Return one consolidated report: No subagent execution lane, no dangerous full access, and no promotion are authorized. |
M5 unattended sequence review — BLOCKED safely; authorize one-click M5A provider shadow onlyThe block is expected and is not a runtime failure. Authorized operation: M5A-PROVIDER-SHADOW-ONECLICKOperator: GPT-5.6 Sol High. No sub-agent. Prepare a persistent non-secret runner and manifest outside Git, for example: Before asking SE7H for approval, complete all non-sensitive work:
After that single human approval, the runner may autonomously:
Forbidden in M5A:
Report: If Codex requests approval for any second command after the top-level runner starts, fail closed, clean up and report the exact reason. Stop after M5A. |
M5A review — fail-closed before providerThe candidate is not defective. The M5A runner used a non-canonical managed identity. A shadow AEON run must keep Authorized next operation: update only the non-secret M5A runner and manifest, then perform an offline binding preflight proving all resolved writable paths remain inside a fresh disposable shadow root. Do not access any external endpoint or living runtime state in this operation. Return the binding report and stop. |
M5A canonical identity binding — PASS; authorize provider shadow rerun onlyThe prior fail-closed result was caused only by an invalid composite identity value. The corrected runner now proves that every runtime path resolves inside the disposable shadow root and that no live AEON path is reachable. Authorized single operation: rerun Requirements:
If a second approval is requested or any path escapes the disposable root, fail closed, clean up, and stop. |
M5A provider shadow — PASS; authorize final shadow goal sequenceAuthorized goal: finish the remaining AEON Minisforum shadow validation in one continuous bounded run. Do not stop after each successful stage. Required stages:
Execution policy:
Return one consolidated report only: Do not request review between successful stages. Stop only at final completion or after a fail-closed cleanup. |
M5 final shadow closure — PASS with documented Telegram polling exceptionThe omitted The hosted AEON Minisforum validation lane is closed. No live deployment, merge, promotion, or ref movement is authorized by this closure. Authorized next operation: one Goal-mode final audit and current-upstream rebase preparation in isolated worktrees only. It may inspect current upstream |
M6 review — local rebase used stale PR base; fresh-main rebase requiredThe reported Authorized next operation: Required gate:
If upstream advances after the frozen SHA, report the new head separately; do not restart repeatedly unless a relevant file changed. |
M6R semantic conflict resolution — preserve current main test topology, transplant only resume invariantsDecision for conflict in
Reason: current upstream main intentionally has a much smaller No push or runtime mutation is authorized. |
M6R closure — fresh-main candidate accepted; controlled upstream push authorizedThe reduced current-main test topology is authoritative. The historical 126-test count is not a required gate because those tests no longer exist in that topology and were intentionally not restored. Accepted evidence: 6 dedicated resume tests, 31 current resume/session/CWD tests, compile PASS, Ruff PASS, diff-check PASS, clean worktree. Authorized operation: verify the remote branch still equals the exact old head above, then push only the final candidate to After push, add one concise comment to upstream PR NousResearch#74397 recording the frozen/current main SHAs, semantic conflict resolution, final head, and current test results. Do not mark ready, merge, deploy, or mutate any runtime. Stop after reporting the pushed SHA and initial CI run identifiers. |
…own (NousResearch#74136) Fix-up for the cherry-picked cooldown persistence: the PR's tests mocked the DB (SimpleNamespace(_db=MagicMock())), which cannot prove the cooldown survives a restart. Replace with the production shape — a real SessionDB on disk behind the real AsyncSessionDB facade — and add a restart regression: fail a hygiene compression on runner #1, tear it down, build a fresh GatewayRunner on the SAME database, and assert the cooldown is still honored (no compression agent instantiated). Also updates the timeout test to assert the DB-backed record_compression_failure_cooldown write instead of the removed in-memory dict. Sabotage-verified: reverting gateway/run.py to the in-memory dict makes the restart test fail.
Users following abbreviated links guess /docs/quickstart and /docs/installation and hit raw GitHub-Pages 404s — the real pages live under /docs/getting-started/. Add client redirects for both. Consumer-onboarding audit finding #1, Aug 2026.
The #1 patch failure class in production (state.db mining, 250k-window) is a re-send of an edit that already landed: 'old_string and new_string are identical' (299 occurrences) plus a share of hunk-not-found errors where the new text is already in the file. These errored, sending models into re-read/re-patch loops. New tools/fuzzy_match.is_already_applied(content, old, new) — a conservative check requiring (1) non-trivial new_string (>=8 chars), (2) EXACT presence of new_string, (3) old_string gone (unless identical). Wired into three sites: - patch_replace (replace mode): returns success + no_change: true + an explicit note instead of the identical-strings / no-match error. - V4A validation phase: an already-applied hunk validates as a no-op so multi-hunk patches no longer fail wholesale when one hunk landed in a prior call. - V4A apply phase: mirrors the same skip so the two phases agree. Genuine no-matches (new text absent) and half-applied renames (old text still present) keep their error behavior — covered by tests.
process(action='wait') hitting its window returned status='timeout' with a terse note — models read it as an error and re-issued identical waits (process is the #1 exact-duplicate tool call in production: 511 dupes in a 400k-msg window; wait is 57% of all process actions). The timeout result now carries: - process_running: true — machine-readable 'this is a status, not a failure' - an explicit note: 'Wait window of Ns elapsed — the process is still running. This is not an error. Uptime: Ms.' plus the right next step: when notify_on_complete is set, 'you will be notified on exit — do more work instead of waiting again'; otherwise a pointer to notify_on_complete for next time. - the clamp note (requested > max) now composes with the status note instead of replacing it. Exited/interrupted results are unchanged.
…e-review #1) revoke_commit_admission() used to invoke the holder-qualified lease release unconditionally — including while an admitted commit was still mutating SessionDB — letting a second compressor acquire the durable lock mid-commit and interleave with the first commit's writes. The admission_revoked flag store stays lock-free, but the lease-release decision now coordinates with the fence lock: - revoke acquires the fence lock non-blocking; on success no commit can be in flight (an admitted commit retains the lock until finish_commit) and the release runs immediately, still under the lock so a racing begin_commit cannot slip between the check and the release. - on failure the release is deferred: finish_commit() re-checks _admission_revoked and performs it AFTER the commit completes (prompt even if the worker thread is later parked), and the begin_commit refusal path does the same for a revoke that lost the race to a transient lock-setup/cancel boundary. All paths are idempotent with the worker's own outer cleanup (DB release is holder-qualified). Invariant encoded + tested: no second compressor can acquire the durable lock while an admitted commit is still mutating; after a post-revoke commit finishes the lease is released promptly. Both regressions (revoke-during-commit deferral, revoke-before-commit immediate release + refused begin_commit) are sabotage-verified.
M6R upstream CI gate — awaiting maintainer approval, not a code failureInterpretation: the workflow did not execute any job. The run is awaiting approval required for a pull request workflow originating from an external fork. This is an upstream repository governance boundary, not a test or code failure. No rerun, code change, rebase, push, or CI repair is authorized. The next valid event is maintainer approval of the workflow in NousResearch/hermes-agent. After approval, inspect the newly executed jobs and logs before any further action. |
Operator/runtime type split + frozen live-update policyCanonical topology
AVAEON may retain operator metadata, isolated worktrees, harnesses, manifests, reports, and disposable shadow roots. It must never reuse a live AVA/AEON state root or present an operator instance as Required control-plane repairRefactor the PR so that:
Live Hermes update policyThe living AEON Hermes runtime on the Minisforum is frozen at its exact current working revision until an explicit controlled promotion. Do not assume the live revision equals the validated shadow candidate. Inventory and record the exact live SHA before any future promotion. All future Hermes updates must use this pipeline: Upstream changes are intake material, never authority to mutate the live vessel automatically. Authority separation
Execution authorizationUse GPT-5.6 Luna medium, Goal mode, no sub-agent. Modify only PR #1 code/tests/docs and isolated worktrees. Run the relevant control-plane and schema tests. A normal fast-forward push to No live runtime change, no service restart, no Return one consolidated report:
Include modified files, tests, resulting SHA, proof that AVAEON is no longer a runtime entity, proof that |
M7 postflight — structure accepted; three closure repairs required
Before M7 is fully closed, apply only these repairs:
Accuracy correction for update policy:
Therefore document the exact boundary as: The residual external-fleet migration remains deferred: inventory it read-only before any future promotion, but do not touch live AEON configuration now. No runtime restart, promotion, Unchained work, AVAORUS work, or broad Hermes CLI test run is authorized. Push one fast-forward cleanup commit only after the focused checks pass, then STOP. |
M7 final closure — operator/runtime split and frozen update policy sealedPostflight confirmed:
Deferred boundaries remain explicit: read-only inventory and migration of external fleet configurations before any promotion; AEON Unchained later; AVAORUS validation when the host becomes reachable. M7 is closed. AVAEON Codex should remain stopped until a new authorized goal exists. |
Purpose
Establish a controlled Hermes runtime and fleet control plane for the two
runtime entities
avaandaeon. AVAEON Codex is the portable deployment andaudit operator; it is not a Hermes runtime and is never a fleet member.
Canonical identity model
AVA_ENTITYis always the selected runtime target (avaoraeon).operator_id=avaeon-codexis separate metadata. AVAEON has no runtime,Telegram, gateway, service,
state.db, permanent snapshot, quorum, orpromotion obligation. It may retain isolated worktrees, harnesses, manifests,
reports, and disposable shadow roots only.
Delivered foundation
HERMES_HOME, workspace and session isolation;TERMINAL_CWDhandling;Current M1–M7 evidence
tests closed.
no real OpenAI endpoint, credential or live provider was used.
persistence, duplicate-consumer and rollback evidence. AEON Unchained remains
deferred.
isolated worktrees; no live runtime mutation.
dcc9538145a4d58438bbc7137a859009166fba9b; 77 AVA-runtime tests,compilation, Ruff and diff-check passed. The postflight cleanup removes the
residual test processes and corrects the documentation/policy boundary.
AEON hosted shadow is closed. AVA/AVAORUS validation is explicitly deferred
while that host is unavailable. No live promotion, stable movement, service
restart, or live state mutation has been performed.
Frozen update policy
The managed control plane rejects
auto_update: true, non-40-character refs,and moving branches as live sources.
hermes updateis operationallyforbidden. The upstream binary is not globally intercepted outside the managed
control plane; it must not be invoked directly for live updates.
Every future version must use an isolated pinned candidate, semantic audit,
focused and repository-compatible tests, disposable shadow, state snapshot,
explicit authorization, controlled promotion and rollback. Upstream is intake
material only. AEON Core may provide read-only diagnostics and smokes, but may
not deploy, approve or mutate its own live runtime.
Deferred boundary
External fleet configurations have not been migrated. They must first be
inventoried read-only and then migrated to the two-runtime/operator model before
any future promotion. No live AEON configuration was changed.
The PR remains draft; no promotion to staging or stable is performed here.