Skip to content

fix(omnio): preserve skill instructions and worker results - #112

Merged
ppazosp merged 5 commits into
mainfrom
ppp/omnio-skill-delivery-reliability
Sep 18, 2026
Merged

ppazosp merged 5 commits into
mainfrom
ppp/omnio-skill-delivery-reliability

Conversation

@ppazosp

@ppazosp ppazosp commented Sep 18, 2026 •

Copy link
Copy Markdown

Omnio could report a skill as loaded after losing most of its instructions, discard shared worker context, cut worker results to 2,000 characters, and return artifact paths outside the paired Toolbox. A separate RPC delivery failure could execute the same tool repeatedly. This PR fixes those five mechanical failures in the fork.

Fixes and upstream adaptations

Failure Upstream source Fork adaptation
A large skill becomes a misleading preview #98736, closed without merge Explicit incomplete receipt plus complete section/part recovery. Preserve linked-file selectors, duplicate headings, preambles and aggregate-budget retries. Omit the upstream-only repeat-view cache.
Worker summaries shrink as cumulative billing grows #103486, merged Measure the current parent prompt separately from lifetime billing and adviser usage. Keep the fork's existing asynchronous delegation lifecycle.
Worker artifact paths cannot be read from Toolbox #103667, merged Map paths and actually transfer active-profile delegation files into Toolbox using existing file endpoints; upstream mapping alone is insufficient for paired Sprites.
Shared context is silently ignored beside a batch #96424, merged Present tasks[] to the model, preserve legacy call inputs, and merge shared context into every child's context. Preserve existing role/depth policy.
Large execute_code responses repeatedly dispatch the same tool Fork-specific correction; no matching upstream fix identified Use existing file transfer or bounded chunks, publish atomically, retain completed results through delivery retries, and return an explicit transport error after exhaustion.

Design

Instruction delivery remains subject to context budgets. An omitted skill body now says that its instructions were not delivered and exposes bounded navigation that can recover every part. Small skills retain their existing complete response.

Worker summaries use current context occupancy, with the existing static cap when occupancy is unknown. Complete summaries and logs stay in the active profile; only regular delegation artifacts cross into the brand-scoped Toolbox. Transfers reject symlinks and invalid destinations, redact through the existing log redactor, and refresh on commands/reads. Direct artifact reads surface transfer failure instead of claiming a stale copy is current. Remote logs are snapshots refreshed on each read, not a continuously updating tail.

RPC tool execution and response delivery are separate steps. Retrying delivery cannot redispatch a completed request during the RPC lifetime. A transport outage can still fail the script, but it no longer turns into repeated tool execution.

flowchart LR
  H[Hermes tool result] --> B{Context budget}
  B -->|fits| M[Parent model]
  B -->|skill omitted| I[Explicit incomplete receipt + section recovery]
  I --> M
  W[Worker] --> C[Current-prompt summary budget]
  C --> M
  W --> F[Profile delegation artifacts]
  F -->|existing file API| T[Brand Toolbox]
Loading

Verification

CI confirmed green at 74600002d7: all 8 Python groups and all required checks passed. Run 35325635727. Previously failing group 3 completed with 6,825 passed, 0 failed, including all 60 MCP execution tests.

CI follow-up 74600002d7: reproduced the two failing MCP transport tests (58 passed, 2 failed), then replaced shell-command-order assertions with actual byte-transfer checks under the command limit. Covers large Unicode content, quoted paths, stale staging, empty files and staging cleanup. All six execute-code suites now pass locally: 229 passed, retries disabled. Production code is unchanged by this follow-up.

  • Adjacent real-browser scenarios omnio.runtime.persistence.real-user-tool-history-cold-restore and omnio.runtime.persistence.real-user-compaction-cold-restore: both passed, zero retries (5.0 minutes total). Intentional restore compaction remains intact.

  • Five behavioral reproductions fail on the unmodified base e851c3f8bd (5 failed, 0 passed). They cover incomplete skills, current-prompt budgeting, worker artifact paths, shared context, and oversized RPC delivery.

  • Canonical scripts/run_tests.sh runner: 900 tests passed across 21 relevant/adjacent files. Later focused checks passed too: 212 skill/delegation/RPC tests and 163 delegation tests after the final wording correction.

  • Full ruff check . and git diff --check pass.

  • Real authenticated browser scenario omnio.runtime.personalization.real-user-skill-delivery: two successful runs, zero retries, all six checks passed (five failures plus durable UI completion).

  • Separate real-Sprite probe execute-code-result-delivery: 100,000 bytes, exact SHA-256, one inner terminal call.

  • Fault tests cover failed publication, failed request removal, permanent transport outage, partial writes, large raw uploads, and artifact refresh failure.

Live runtime: omnio-e2e-omnio-b17f8144bf21156f, Toolbox omnio-e2e-toolbox-7159b4969d468aa0; remote Git HEAD verified as c69058b1125fb2be68b5fc504fc53b76cc4f3d15. The final 0fc6ddc1fd commit changes transcript-refresh guidance and a comment only; its delegation tests passed. Browser inference was scripted while tools, transport, persistence and UI were real. These checks establish mechanics, not Luna's editorial quality. The full probe library has not been run.

The initial browser harness failures are retained: oversized fixture argv, decorated-module loading, unsupported inner tool, and a trailing-newline expectation. Each was corrected before the successful runs; no automatic retry concealed a failure.

Backward compatibility

Backward compatible. Old Omnia/proxy/Toolbox with new Hermes uses the existing file endpoints and unchanged wire contracts. The companion Omnia change works with old Hermes (its new regression correctly fails until this fix is present). Legacy single-task and shared-context calls still execute, including persisted calls; only the model-facing schema prefers tasks[]. Existing roles, tool permissions, persisted events and intentional cold-restore compaction are unchanged. No DB migration, capability floor, or simultaneous rollout is required.

Rollout

Companion regression coverage: useomnia/omnia#4690.

Merge this fork PR before selecting its commit for new/reprovisioned Sprites. Existing production Sprites keep their current Hermes until reprovisioned. This work does not change a production pin or deploy customer runtimes.

Evidence

The attached reel summarizes the verified mechanism and the successful browser path. Authenticated traces remain local; only synthetic, sanitized media is attached.

Reel specification
{
  "specVersion": "1.0.0",
  "scenario": "omnio.runtime.personalization.real-user-skill-delivery",
  "tag": "INSTRUCTION DELIVERY",
  "date": "18 September 2026",
  "footer": {
    "left": "Real browser + Omnio/Toolbox pair \u00b7 scripted inference",
    "right": "5 baseline failures \u00b7 2 repaired runs \u00b7 zero retries"
  },
  "scenes": [
    {
      "type": "title",
      "duration": 8,
      "kicker": "FIVE MECHANICAL FAILURES",
      "headline": "Instructions and results must <em>reach the model.</em>",
      "lede": "Skills could look loaded while most instructions were missing. Workers lost shared context, shortened useful results, and advertised inaccessible files. Large RPC responses could replay tools.",
      "bullets": [
        [
          "Before",
          "All five reproductions fail on the original Hermes source."
        ],
        [
          "After",
          "All five pass through the real browser and paired Sprites."
        ]
      ]
    },
    {
      "type": "footage",
      "duration": 10,
      "beat": "delivery-completed",
      "state": "after",
      "step": "1",
      "headline": "The follow-up reads complete worker artifacts.",
      "subhead": "Normal authenticated Omnio chat. Inference is scripted; tools, transport, persistence and UI are real.",
      "still": "conformance-repeat-green/playwright/omnio.runtime.personalizat-ced75-on-real-user-skill-delivery/instruction-delivery.png",
      "rail": [
        {
          "label": "Skill recovery",
          "value": "Final instruction received",
          "tone": "good"
        },
        {
          "label": "Worker context",
          "value": "Both instructions delivered",
          "tone": "good"
        },
        {
          "label": "Large RPC result",
          "value": "100,000 bytes \u00b7 1 call",
          "tone": "good"
        },
        {
          "label": "Worker artifact",
          "value": "Full-content hash matches",
          "tone": "good"
        }
      ]
    },
    {
      "type": "closing",
      "duration": 10,
      "kicker": "FORK ADAPTATIONS VERIFIED",
      "headline": "Keep limits. Make delivery <em>complete and explicit.</em>",
      "points": [
        "Oversized skills disclose omitted instructions and expose recoverable sections.",
        "Worker budgets use the current prompt; shared context reaches every child.",
        "Toolbox receives worker artifacts; RPC retries never redispatch completed tools.",
        "These checks prove mechanics. They do not measure Luna\u2019s editorial quality."
      ],
      "closer": "Backward compatible. Production rollout requires merge and reprovision."
    }
  ]
}
skill-delivery-verification.mp4
Sanitized real-Sprite verification results
{
  "scenario": "omnio.runtime.personalization.real-user-skill-delivery",
  "status": "passed",
  "runId": "c63e8e27884946e2bfb3eb24cf6e6dfa",
  "conversationId": "b1bdb9d4-62ba-4c1d-b180-ca12d4348b4b",
  "runtimeId": "omnio-e2e-omnio-b17f8144bf21156f",
  "toolboxRuntimeId": "omnio-e2e-toolbox-7159b4969d468aa0",
  "checks": {
    "realBoundary": "passed",
    "rpcDelivery": "passed",
    "skillRecovery": "passed",
    "workerArtifacts": "passed",
    "workerBudget": "passed",
    "workerContext": "passed"
  },
  "resultSha256": "390b61cac602b76033ea56bd5be61032b2a949cddfc742ad552e63ae6f5b77a9",
  "rpcProbe": {
    "bytes": 100000,
    "innerCalls": 1,
    "sha256": "1d909a774cd0c7b3bcfaab8fef0a4039ee30319e88ff01a91387470bbe323a69"
  }
}

ppazosp and others added 5 commits September 18, 2026 09:49
Adapt NousResearch#98736 (2fce577) to the fork without its upstream-only repeat-view cache. Preserve linked-file selectors and recover complete sections through both per-result and aggregate budgets.

Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com>
Adapt the current-prompt budget correction from upstream NousResearch#103486, cache-path mapping from NousResearch#103667, and tasks-only schema from NousResearch#96424. Retain the legacy call interface, preserve shared batch context, and transfer only active-profile delegation artifacts into the paired Toolbox using existing file APIs.
Use the existing file transport or bounded shell chunks, publish atomically, and retain dispatched results through delivery retries. Fail explicitly after exhausted delivery instead of executing the same request again.
@ppazosp
ppazosp marked this pull request as ready for review September 18, 2026 08:49
@ppazosp
ppazosp merged commit 3b03efe into main Sep 18, 2026
37 checks passed
ppazosp added a commit that referenced this pull request Sep 21, 2026
… Toolbox paths

Every host-side producer that hands the model a file (stored web pages,
browser snapshots, worker artifacts, execute_code output) used to need
its own bridge to the paired Toolbox, and several still published host
paths that read_file cannot open there.

Adopt upstream's single mechanism instead: the cache subdirectories in
credential_files._CACHE_DIRS are synced to the Brand's private
/tmp/.omnio-session/cache before every command and before any read under
it, and to_agent_visible_cache_path() renders the Toolbox path for the
sprites backend at every footer that publishes one (web_extract, browser
snapshots, delegate summaries). from_agent_visible_cache_path() maps back
so host-side media reads keep their fast path. browser_vision keeps
publishing the host screenshot path: the omnio_paths plugin owns that
field and translates it itself once it detects the projection.

The delegation-only sync from #112 folds into the generic projection:
same redaction of text artifacts, same refusal of paths outside the
projected root, plus a 50 MiB per-file ceiling (logged once) so media
never stalls the per-command sync. Text files above the 2 MiB JSON read
cap are read and paged through the bounded raw stream (5 MiB) instead
of failing with "File too large".

execute_code stdout recovery (upstream NousResearch#97043/NousResearch#97048) rides on it: when
stdout is truncated to its head/tail window, the sanitized full output
(up to 5 MB, complete lines) is saved under cache/exec and the result
names the Toolbox path, so the middle can be paged instead of re-running
the script. A failed transfer surfaces as a read error, never as a host
path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
ppazosp added a commit that referenced this pull request Sep 21, 2026
#115)

* feat(sprites): project the harness cache onto the Toolbox and publish Toolbox paths

Every host-side producer that hands the model a file (stored web pages,
browser snapshots, worker artifacts, execute_code output) used to need
its own bridge to the paired Toolbox, and several still published host
paths that read_file cannot open there.

Adopt upstream's single mechanism instead: the cache subdirectories in
credential_files._CACHE_DIRS are synced to the Brand's private
/tmp/.omnio-session/cache before every command and before any read under
it, and to_agent_visible_cache_path() renders the Toolbox path for the
sprites backend at every footer that publishes one (web_extract, browser
snapshots, delegate summaries). from_agent_visible_cache_path() maps back
so host-side media reads keep their fast path. browser_vision keeps
publishing the host screenshot path: the omnio_paths plugin owns that
field and translates it itself once it detects the projection.

The delegation-only sync from #112 folds into the generic projection:
same redaction of text artifacts, same refusal of paths outside the
projected root, plus a 50 MiB per-file ceiling (logged once) so media
never stalls the per-command sync. Text files above the 2 MiB JSON read
cap are read and paged through the bounded raw stream (5 MiB) instead
of failing with "File too large".

execute_code stdout recovery (upstream NousResearch#97043/NousResearch#97048) rides on it: when
stdout is truncated to its head/tail window, the sanitized full output
(up to 5 MB, complete lines) is saved under cache/exec and the result
names the Toolbox path, so the middle can be paged instead of re-running
the script. A failed transfer surfaces as a read error, never as a host
path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(sprites): deliverable-safe cache root, publish-time push and streaming pages

Review of the cache projection found three gaps:

- The projected root was a dot directory. The Omnio proxy refuses hidden
  path segments in sandbox: links, so a screenshot or artifact published
  under it could never be handed over. The root is now /tmp/omnio-session.
- Paths were published before the file reached the Toolbox, so consumers
  that read the Toolbox directly (the proxy's deliverable warm-up, the
  video wrapper) could miss a fresh artifact until the next Hermes command
  or read. publish_cache_path() now pushes the projection before it hands
  out a path; the Sprites environment registers its live instances for
  that flush, and a failed push still falls back to the read-time sync.
- Above the 5 MiB whole-read ceiling the model was told to page with
  offset/limit, but paging loaded the whole file and failed the same way.
  Pages are now computed from the raw stream one chunk at a time, so any
  size pages and the recipe always works.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(sprites): publish cold cache and attachments before remote reads

* fix(skills): align runtime paths and adopt upstream lookup repairs

* fix(skills): keep runtime path lookup from replacing live mounts

* fix(sprites): refresh skill helpers before execution and reads

* test(skills): make config cache invalidation timing deterministic

* test(gateway): verify finalize timeout without wall-clock races

* fix(execute_code): preserve remote recovery when consolidating #114

Retain the direct remote stdout publication from #114 (930b3c0), adapted from upstream NousResearch#97043 and NousResearch#97048, while Sprites keeps #115's shared cache projection. Cover once-only remote execution, recovery of middle output, and failed publication without a host path.

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>

* test(execute_code): guard the POSIX remote filesystem fixture

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant