fix(omnio): preserve skill instructions and worker results - #112
Merged
Merged
Conversation
Adapt NousResearch#98736 (2fce577) to the fork without its upstream-only repeat-view cache. Preserve linked-file selectors and recover complete sections through both per-result and aggregate budgets. Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com>
Adapt the current-prompt budget correction from upstream NousResearch#103486, cache-path mapping from NousResearch#103667, and tasks-only schema from NousResearch#96424. Retain the legacy call interface, preserve shared batch context, and transfer only active-profile delegation artifacts into the paired Toolbox using existing file APIs.
Use the existing file transport or bounded shell chunks, publish atomically, and retain dispatched results through delivery retries. Fail explicitly after exhausted delivery instead of executing the same request again.
ppazosp
marked this pull request as ready for review
September 18, 2026 08:49
ppazosp
added a commit
that referenced
this pull request
Sep 21, 2026
… Toolbox paths Every host-side producer that hands the model a file (stored web pages, browser snapshots, worker artifacts, execute_code output) used to need its own bridge to the paired Toolbox, and several still published host paths that read_file cannot open there. Adopt upstream's single mechanism instead: the cache subdirectories in credential_files._CACHE_DIRS are synced to the Brand's private /tmp/.omnio-session/cache before every command and before any read under it, and to_agent_visible_cache_path() renders the Toolbox path for the sprites backend at every footer that publishes one (web_extract, browser snapshots, delegate summaries). from_agent_visible_cache_path() maps back so host-side media reads keep their fast path. browser_vision keeps publishing the host screenshot path: the omnio_paths plugin owns that field and translates it itself once it detects the projection. The delegation-only sync from #112 folds into the generic projection: same redaction of text artifacts, same refusal of paths outside the projected root, plus a 50 MiB per-file ceiling (logged once) so media never stalls the per-command sync. Text files above the 2 MiB JSON read cap are read and paged through the bounded raw stream (5 MiB) instead of failing with "File too large". execute_code stdout recovery (upstream NousResearch#97043/NousResearch#97048) rides on it: when stdout is truncated to its head/tail window, the sanitized full output (up to 5 MB, complete lines) is saved under cache/exec and the result names the Toolbox path, so the middle can be paged instead of re-running the script. A failed transfer surfaces as a read error, never as a host path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
ppazosp
added a commit
that referenced
this pull request
Sep 21, 2026
#115) * feat(sprites): project the harness cache onto the Toolbox and publish Toolbox paths Every host-side producer that hands the model a file (stored web pages, browser snapshots, worker artifacts, execute_code output) used to need its own bridge to the paired Toolbox, and several still published host paths that read_file cannot open there. Adopt upstream's single mechanism instead: the cache subdirectories in credential_files._CACHE_DIRS are synced to the Brand's private /tmp/.omnio-session/cache before every command and before any read under it, and to_agent_visible_cache_path() renders the Toolbox path for the sprites backend at every footer that publishes one (web_extract, browser snapshots, delegate summaries). from_agent_visible_cache_path() maps back so host-side media reads keep their fast path. browser_vision keeps publishing the host screenshot path: the omnio_paths plugin owns that field and translates it itself once it detects the projection. The delegation-only sync from #112 folds into the generic projection: same redaction of text artifacts, same refusal of paths outside the projected root, plus a 50 MiB per-file ceiling (logged once) so media never stalls the per-command sync. Text files above the 2 MiB JSON read cap are read and paged through the bounded raw stream (5 MiB) instead of failing with "File too large". execute_code stdout recovery (upstream NousResearch#97043/NousResearch#97048) rides on it: when stdout is truncated to its head/tail window, the sanitized full output (up to 5 MB, complete lines) is saved under cache/exec and the result names the Toolbox path, so the middle can be paged instead of re-running the script. A failed transfer surfaces as a read error, never as a host path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(sprites): deliverable-safe cache root, publish-time push and streaming pages Review of the cache projection found three gaps: - The projected root was a dot directory. The Omnio proxy refuses hidden path segments in sandbox: links, so a screenshot or artifact published under it could never be handed over. The root is now /tmp/omnio-session. - Paths were published before the file reached the Toolbox, so consumers that read the Toolbox directly (the proxy's deliverable warm-up, the video wrapper) could miss a fresh artifact until the next Hermes command or read. publish_cache_path() now pushes the projection before it hands out a path; the Sprites environment registers its live instances for that flush, and a failed push still falls back to the read-time sync. - Above the 5 MiB whole-read ceiling the model was told to page with offset/limit, but paging loaded the whole file and failed the same way. Pages are now computed from the raw stream one chunk at a time, so any size pages and the recipe always works. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(sprites): publish cold cache and attachments before remote reads * fix(skills): align runtime paths and adopt upstream lookup repairs * fix(skills): keep runtime path lookup from replacing live mounts * fix(sprites): refresh skill helpers before execution and reads * test(skills): make config cache invalidation timing deterministic * test(gateway): verify finalize timeout without wall-clock races * fix(execute_code): preserve remote recovery when consolidating #114 Retain the direct remote stdout publication from #114 (930b3c0), adapted from upstream NousResearch#97043 and NousResearch#97048, while Sprites keeps #115's shared cache projection. Cover once-only remote execution, recovery of middle output, and failed publication without a host path. Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com> * test(execute_code): guard the POSIX remote filesystem fixture --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Omnio could report a skill as loaded after losing most of its instructions, discard shared worker context, cut worker results to 2,000 characters, and return artifact paths outside the paired Toolbox. A separate RPC delivery failure could execute the same tool repeatedly. This PR fixes those five mechanical failures in the fork.
Fixes and upstream adaptations
Design
Instruction delivery remains subject to context budgets. An omitted skill body now says that its instructions were not delivered and exposes bounded navigation that can recover every part. Small skills retain their existing complete response.
Worker summaries use current context occupancy, with the existing static cap when occupancy is unknown. Complete summaries and logs stay in the active profile; only regular delegation artifacts cross into the brand-scoped Toolbox. Transfers reject symlinks and invalid destinations, redact through the existing log redactor, and refresh on commands/reads. Direct artifact reads surface transfer failure instead of claiming a stale copy is current. Remote logs are snapshots refreshed on each read, not a continuously updating tail.
RPC tool execution and response delivery are separate steps. Retrying delivery cannot redispatch a completed request during the RPC lifetime. A transport outage can still fail the script, but it no longer turns into repeated tool execution.
flowchart LR H[Hermes tool result] --> B{Context budget} B -->|fits| M[Parent model] B -->|skill omitted| I[Explicit incomplete receipt + section recovery] I --> M W[Worker] --> C[Current-prompt summary budget] C --> M W --> F[Profile delegation artifacts] F -->|existing file API| T[Brand Toolbox]Verification
CI confirmed green at
74600002d7: all 8 Python groups and all required checks passed. Run 35325635727. Previously failing group 3 completed with 6,825 passed, 0 failed, including all 60 MCP execution tests.CI follow-up
74600002d7: reproduced the two failing MCP transport tests (58 passed, 2 failed), then replaced shell-command-order assertions with actual byte-transfer checks under the command limit. Covers large Unicode content, quoted paths, stale staging, empty files and staging cleanup. All six execute-code suites now pass locally: 229 passed, retries disabled. Production code is unchanged by this follow-up.Adjacent real-browser scenarios
omnio.runtime.persistence.real-user-tool-history-cold-restoreandomnio.runtime.persistence.real-user-compaction-cold-restore: both passed, zero retries (5.0 minutes total). Intentional restore compaction remains intact.Five behavioral reproductions fail on the unmodified base
e851c3f8bd(5 failed, 0 passed). They cover incomplete skills, current-prompt budgeting, worker artifact paths, shared context, and oversized RPC delivery.Canonical
scripts/run_tests.shrunner: 900 tests passed across 21 relevant/adjacent files. Later focused checks passed too: 212 skill/delegation/RPC tests and 163 delegation tests after the final wording correction.Full
ruff check .andgit diff --checkpass.Real authenticated browser scenario
omnio.runtime.personalization.real-user-skill-delivery: two successful runs, zero retries, all six checks passed (five failures plus durable UI completion).Separate real-Sprite probe
execute-code-result-delivery: 100,000 bytes, exact SHA-256, one inner terminal call.Fault tests cover failed publication, failed request removal, permanent transport outage, partial writes, large raw uploads, and artifact refresh failure.
Live runtime:
omnio-e2e-omnio-b17f8144bf21156f, Toolboxomnio-e2e-toolbox-7159b4969d468aa0; remote Git HEAD verified asc69058b1125fb2be68b5fc504fc53b76cc4f3d15. The final0fc6ddc1fdcommit changes transcript-refresh guidance and a comment only; its delegation tests passed. Browser inference was scripted while tools, transport, persistence and UI were real. These checks establish mechanics, not Luna's editorial quality. The full probe library has not been run.The initial browser harness failures are retained: oversized fixture argv, decorated-module loading, unsupported inner tool, and a trailing-newline expectation. Each was corrected before the successful runs; no automatic retry concealed a failure.
Backward compatibility
Backward compatible. Old Omnia/proxy/Toolbox with new Hermes uses the existing file endpoints and unchanged wire contracts. The companion Omnia change works with old Hermes (its new regression correctly fails until this fix is present). Legacy single-task and shared-context calls still execute, including persisted calls; only the model-facing schema prefers tasks[]. Existing roles, tool permissions, persisted events and intentional cold-restore compaction are unchanged. No DB migration, capability floor, or simultaneous rollout is required.
Rollout
Companion regression coverage: useomnia/omnia#4690.
Merge this fork PR before selecting its commit for new/reprovisioned Sprites. Existing production Sprites keep their current Hermes until reprovisioned. This work does not change a production pin or deploy customer runtimes.
Evidence
The attached reel summarizes the verified mechanism and the successful browser path. Authenticated traces remain local; only synthetic, sanitized media is attached.
Reel specification
{ "specVersion": "1.0.0", "scenario": "omnio.runtime.personalization.real-user-skill-delivery", "tag": "INSTRUCTION DELIVERY", "date": "18 September 2026", "footer": { "left": "Real browser + Omnio/Toolbox pair \u00b7 scripted inference", "right": "5 baseline failures \u00b7 2 repaired runs \u00b7 zero retries" }, "scenes": [ { "type": "title", "duration": 8, "kicker": "FIVE MECHANICAL FAILURES", "headline": "Instructions and results must <em>reach the model.</em>", "lede": "Skills could look loaded while most instructions were missing. Workers lost shared context, shortened useful results, and advertised inaccessible files. Large RPC responses could replay tools.", "bullets": [ [ "Before", "All five reproductions fail on the original Hermes source." ], [ "After", "All five pass through the real browser and paired Sprites." ] ] }, { "type": "footage", "duration": 10, "beat": "delivery-completed", "state": "after", "step": "1", "headline": "The follow-up reads complete worker artifacts.", "subhead": "Normal authenticated Omnio chat. Inference is scripted; tools, transport, persistence and UI are real.", "still": "conformance-repeat-green/playwright/omnio.runtime.personalizat-ced75-on-real-user-skill-delivery/instruction-delivery.png", "rail": [ { "label": "Skill recovery", "value": "Final instruction received", "tone": "good" }, { "label": "Worker context", "value": "Both instructions delivered", "tone": "good" }, { "label": "Large RPC result", "value": "100,000 bytes \u00b7 1 call", "tone": "good" }, { "label": "Worker artifact", "value": "Full-content hash matches", "tone": "good" } ] }, { "type": "closing", "duration": 10, "kicker": "FORK ADAPTATIONS VERIFIED", "headline": "Keep limits. Make delivery <em>complete and explicit.</em>", "points": [ "Oversized skills disclose omitted instructions and expose recoverable sections.", "Worker budgets use the current prompt; shared context reaches every child.", "Toolbox receives worker artifacts; RPC retries never redispatch completed tools.", "These checks prove mechanics. They do not measure Luna\u2019s editorial quality." ], "closer": "Backward compatible. Production rollout requires merge and reprovision." } ] }skill-delivery-verification.mp4
Sanitized real-Sprite verification results
{ "scenario": "omnio.runtime.personalization.real-user-skill-delivery", "status": "passed", "runId": "c63e8e27884946e2bfb3eb24cf6e6dfa", "conversationId": "b1bdb9d4-62ba-4c1d-b180-ca12d4348b4b", "runtimeId": "omnio-e2e-omnio-b17f8144bf21156f", "toolboxRuntimeId": "omnio-e2e-toolbox-7159b4969d468aa0", "checks": { "realBoundary": "passed", "rpcDelivery": "passed", "skillRecovery": "passed", "workerArtifacts": "passed", "workerBudget": "passed", "workerContext": "passed" }, "resultSha256": "390b61cac602b76033ea56bd5be61032b2a949cddfc742ad552e63ae6f5b77a9", "rpcProbe": { "bytes": 100000, "innerCalls": 1, "sha256": "1d909a774cd0c7b3bcfaab8fef0a4039ee30319e88ff01a91387470bbe323a69" } }