Skip to content

feat: add opt-in Jev timing policy for native omp compaction - #30

Merged
MrGTV-love merged 2 commits into
mainfrom
fm/fm-compact-jev-omp-pipeline
Oct 6, 2026
Merged

MrGTV-love merged 2 commits into
mainfrom
fm/fm-compact-jev-omp-pipeline

Conversation

@MrGTV-love

Copy link
Copy Markdown
Owner

Intent

"in terms of cutting over to kun's compact-adviser vs using omp's system, compare and contrast the three options. What are the most relevant decision factors (some assessment of compaction quality would chief among them I would suspect)?" Then: "go for the compaction bakeoff".

This task implements option 3 for the approved three-way omp compaction comparison: Jev integrated into omp's compaction pipeline, rather than replacing omp's compactor.

What Changed

  • Add an experimental omp extension that judges settled turns with Jev and requests native compaction, leaving summary generation, retention, and persistence to omp.
  • Allow bounded deferral of automatic threshold and idle compaction while preserving manual and recovery precedence; invalidate stale judgments on lifecycle and identity changes.
  • Document isolated bake-off setup, timing-only limitations, and optional metrics; add lifecycle and recovery regression coverage and wire it into the test runner.

Risk Assessment

✅ Low: The change is explicitly opt-in and confined to native-compaction timing, with bounded deferral and stale-result guards, and no material source defect or intent contradiction was substantiated.

Testing

Both targeted regression invocations passed, and isolated real-omp RPC, tool-loop and TUI scenarios demonstrated opt-in timing policy, native persistence, recovery precedence and cancellation guards. Product transcripts and persisted entries were retained; all disposable resources were removed. No full suite, lint or static analysis ran. Fixture endpoints do not establish model quality or paid usage; no UI layout/copy changed, so terminal evidence rather than screenshots was captured.

  • Live validation: ✅ go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
An explicitly enabled, qualifying settled checkpoint invokes omp's native compactor ✅ pass live positive.json: qualifying judgment, checkpoint request and persisted soft compaction from 50030 to 31291 tokens with fromExtension=false.
Loading the entry without opt-in or acknowledged auto consent makes no Jev request ✅ pass live optin-off.json and consent-off.json: completed real product turns with no Jev traffic or eager compaction.
An unfinished judgment defers once within an agent loop, then permits threshold recovery ✅ pass live bounded-tool-loop.json and bounded-tool-loop-product.txt: two successful real read calls; one cancellation after agent_start, followed by native persisted compaction.
A timed-out Jev request does not block automatic native compaction ✅ pass live timeout-native.json: judge-unavailable after the request timeout, no native-deferred event, and subsequent native threshold persistence.
New user input invalidates a pending qualifying judgment ✅ pass live pending-new-input.json: pending-invalidated after a new prompt and no checkpoint compaction after the delayed positive response.
Native model selection invalidates a pending qualifying judgment ✅ pass live pending-model-switch.json: real set_model to fixture-b, pending-invalidated, and no stale checkpoint compaction.
Near the context-window limit, unfinished advice cannot veto native recovery ✅ pass live near-window-precedence.json: unfinished judgment at 260000 context tokens, no advisory deferral, and native persisted compaction.
A healthy unfinished judgment cannot block overflow or incomplete-turn recovery ✅ pass live healthy-overflow.json and incomplete-native.json: actual host-classified overflow/incomplete events, native-precedence records and persisted native soft compactions.
A failed current turn cannot borrow an older successful answer as a checkpoint ✅ pass live failed-terminal.json and failed-terminal-product.txt: actual terminal assistant error despite older successful history, with no Jev request or checkpoint compaction.
Structured secret-bearing context cannot authorize eager compaction ✅ pass live redacted-coverage.json: escaped and nested synthetic saved-key fields produced coverage-ineligible, with no Jev traffic.
The genuine TUI idle timer respects unfinished advice while explicit compact remains native ✅ pass live tui-idle.json and tui-idle-product.txt: idle deferral after 60 seconds, visible maintenance cancellation, then manual native persistence from 50030 to 31330 tokens.
Evidence: Validation index, commands, isolation and evidence limitations

Source: Validation index, commands, isolation and evidence limitations

{
  "targetCommit": "dd801cffe63eaa42cb063f3213becb36a9234979",
  "productVersion": "omp/18.6.3",
  "scope": "Timing-only Jev integration in native omp compaction; not a three-arm compaction-quality assessment",
  "baselineCommands": [
    "bash bin/fm-test-run.sh tests/fm-omp-jev-pipeline.test.sh",
    "FM_JEV_ADVISER_DIR=\"$PWD/.fm-adviser-root/node_modules/compact-adviser\" bash bin/fm-test-run.sh tests/fm-omp-jev-pipeline.test.sh"
  ],
  "baselineObserved": "Both targeted invocations passed; first skips optional real-helper redaction, second exercises it. No full suite, lint, formatter or static analysis run.",
  "dependencyArtifactSHA256": "d19924fa542a777179c2143ef023ddd6f08a207a2216fb06a79a6ac8f3a39cf3",
  "liveSetup": {
    "primaryCommand": "omp --no-extensions --no-rules --no-skills --no-title --no-tools --no-lsp --no-pty --model lab/fixture-a --thinking off --config \"$LAB/overlay.yml\" --resume \"$LAB/session.jsonl\" -e \"$PWD/.fm-adviser-root/omp-jev-entry.mjs\" --mode rpc --no-ui",
    "tuiDifference": "omit --mode rpc --no-ui; private tmux primary is 180x50; idle threshold 40000, idle timeout 60 seconds",
    "toolLoopDifference": "replace --no-tools with --tools read",
    "isolation": "Marked worktree-local lab created by bin/fm-lab-home.sh create, private socket from tmux-dir, inherited path/gate overrides unset, FM_HOME/HOME/PI_CODING_AGENT_DIR assigned only to lab",
    "fixtureEndpoints": "Loopback HTTP model and TypeSafe wire-compatible responses; actual compiled omp, actual static extension and actual adapted helper package; no real credentials/network model calls",
    "cleanup": "For each primary, same private socket tmux -L fm-lab kill-server; bin/fm-lab-home.sh teardown; remove lab. Stop HTTP endpoint and remove .fm-adviser-root materialized only for this phase."
  },
  "reports": [
    {
      "scenario": "near-window-precedence",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/near-window-precedence.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/near-window-precedence-product.txt"
    },
    {
      "scenario": "positive",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/positive.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/positive-product.txt"
    },
    {
      "scenario": "consent-off",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/consent-off.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/consent-off-product.txt"
    },
    {
      "scenario": "bounded-tool-loop",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/bounded-tool-loop.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/bounded-tool-loop-product.txt"
    },
    {
      "scenario": "tui-idle",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/tui-idle.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/tui-idle-product.txt"
    },
    {
      "scenario": "failed-terminal",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/failed-terminal.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/failed-terminal-product.txt"
    },
    {
      "scenario": "pending-new-input",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/pending-new-input.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/pending-new-input-product.txt"
    },
    {
      "scenario": "timeout-native",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/timeout-native.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/timeout-native-product.txt"
    },
    {
      "scenario": "optin-off",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/optin-off.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/optin-off-product.txt"
    },
    {
      "scenario": "incomplete-native",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/incomplete-native.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/incomplete-native-product.txt"
    },
    {
      "scenario": "pending-model-switch",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/pending-model-switch.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/pending-model-switch-product.txt"
    },
    {
      "scenario": "redacted-coverage",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/redacted-coverage.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/redacted-coverage-product.txt"
    },
    {
      "scenario": "healthy-overflow",
      "report": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/healthy-overflow.json",
      "productTranscript": "~/.no-mistakes/evidence/01M49C3HQSKS56S7YCKMMW32FM/healthy-overflow-product.txt"
    }
  ],
  "limitations": [
    "Fixture judgments and summary model responses prove integration/policy mechanics, not real Jev accuracy or native summary fidelity/downstream task performance. Numeric fixture usage is not paid usage.",
    "No UI layout/copy was changed; genuine TUI behavior is retained as terminal transcript/metrics/persisted native entries, not presented as visual screenshot evidence."
  ]
}
Evidence: Qualifying judgment and persisted native compaction

Source: Qualifying judgment and persisted native compaction

{
  "product": "installed omp/18.6.3",
  "isolation": "marked disposable FM_HOME; private fm-lab tmux socket; isolated HOME and PI_CODING_AGENT_DIR; real product and native compactor, fixture model and Jev HTTP endpoints only",
  "name": "positive",
  "observations": {
    "expected": "qualifying Jev judgment requests native soft compaction; persisted result is not supplied by extension",
    "result": "pass",
    "setupCorrection": "Initial fixture classified ordinary chat as summarization and advertised only 1000 tokens; corrected endpoint classification and re-drove."
  },
  "metrics": [
    {
      "event": "pipeline-loaded",
      "at": 1791316596797
    },
    {
      "event": "settled-idle",
      "at": 1791316603664,
      "waitMs": 4
    },
    {
      "event": "checkpoint-ineligible",
      "at": 1791316603666,
      "reason": "context",
      "contextTokens": 1000
    },
    {
      "event": "settled-idle",
      "at": 1791316671119,
      "waitMs": 1
    },
    {
      "event": "judge-start",
      "at": 1791316671123
    },
    {
      "event": "judge-result",
      "at": 1791316671125,
      "apiMs": 2,
      "inputTokens": 1234,
      "outputTokens": 20,
      "estimatedCostUsd": 5.1828000000000006e-05,
      "inputUsdPerMillion": 0.042,
      "model": "jev-1.13.0"
    },
    {
      "event": "judgment",
      "at": 1791316671125,
      "finished": true
    },
    {
      "event": "checkpoint-compact",
      "at": 1791316671126,
      "contextTokens": 50000
    },
    {
      "event": "native-precedence",
      "at": 1791316671130,
      "reason": "manual"
    },
    {
      "event": "native-persisted",
      "at": 1791316671139,
      "tokensBefore": 50030,
      "tokensAfter": 31291,
      "fromExtension": false,
      "method": "soft"
    },
    {
      "event": "checkpoint-complete",
      "at": 1791316671140,
      "elapsedMs": 14
    }
  ],
  "providerTraffic": [
    {
      "surface": "model fixture",
      "at": 1791316603.653335,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 83
    },
    {
      "surface": "model fixture",
      "at": 1791316671.1148102,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": false,
      "messageCount": 85
    },
    {
      "surface": "Jev fixture",
      "at": 1791316671.124557,
      "coverage": {
        "omittedUserMessages": 0,
        "olderMessagesOmitted": 21,
        "recentTextTruncated": false,
        "hasImages": false,
        "redacted": false,
        "unknownContext": false,
        "transcriptRecoverable": true
      },
      "requestBytes": 6165,
      "secretLeaked": false,
      "mode": "positive"
    },
    {
      "surface": "model fixture",
      "at": 1791316671.132679,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 2
    },
    {
      "surface": "model fixture",
      "at": 1791316671.1376178,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 2
    }
  ],
  "persistedCompactions": [
    {
      "type": "compaction",
      "id": "d4baabed",
      "parentId": "84f9c80f",
      "timestamp": "2026-10-06T19:57:51.138Z",
      "summary": "No prior history.\n\n---\n\n**Turn Context (split turn):**\n\n## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.",
      "shortSummary": "## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.",
      "firstKeptEntryId": "dda2b919",
      "tokensBefore": 50030,
      "tokensAfter": 31291,
      "method": "soft",
      "details": {
        "readFiles": [],
        "modifiedFiles": []
      },
      "fromExtension": false
    }
  ]
}
Evidence: Bounded deferral during two real read-tool calls

Source: Bounded deferral during two real read-tool calls

{
  "product": "installed omp/18.6.3",
  "isolation": "marked disposable FM_HOME; private fm-lab tmux socket; isolated HOME and PI_CODING_AGENT_DIR; real product and native compactor, fixture model and Jev HTTP endpoints only",
  "name": "bounded-tool-loop",
  "observations": {
    "expected": "after one advisory deferral in the active agent loop, a subsequent real tool boundary must allow native threshold recovery",
    "result": "pass",
    "scope": "two actual built-in read calls in one agent loop; pre-agent_start threshold attempt and first in-loop attempt are both cancelled, next in-loop attempt persists native result"
  },
  "metrics": [
    {
      "event": "pipeline-loaded",
      "at": 1791317067141
    },
    {
      "event": "settled-idle",
      "at": 1791317067380,
      "waitMs": 1
    },
    {
      "event": "judge-start",
      "at": 1791317067383
    },
    {
      "event": "judge-result",
      "at": 1791317067385,
      "apiMs": 2,
      "inputTokens": 1234,
      "outputTokens": 20,
      "estimatedCostUsd": 5.1828000000000006e-05,
      "inputUsdPerMillion": 0.042,
      "model": "jev-1.13.0"
    },
    {
      "event": "judgment",
      "at": 1791317067385,
      "finished": false
    },
    {
      "event": "native-deferred",
      "at": 1791317067686,
      "reason": "threshold",
      "midUnit": true
    },
    {
      "event": "native-deferred",
      "at": 1791317067725,
      "reason": "threshold",
      "midUnit": true
    },
    {
      "event": "native-persisted",
      "at": 1791317067812,
      "tokensBefore": 50030,
      "tokensAfter": 31718,
      "fromExtension": false,
      "method": "soft"
    },
    {
      "event": "native-persisted",
      "at": 1791317067898,
      "tokensBefore": 50030,
      "tokensAfter": 20667,
      "fromExtension": false,
      "method": "soft"
    },
    {
      "event": "settled-idle",
      "at": 1791317067901,
      "waitMs": 2
    }
  ],
  "providerTraffic": [
    {
      "surface": "model fixture",
      "at": 1791317067.375343,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": false,
      "messageCount": 83
    },
    {
      "surface": "Jev fixture",
      "at": 1791317067.384746,
      "coverage": {
        "omittedUserMessages": 0,
        "olderMessagesOmitted": 19,
        "recentTextTruncated": false,
        "hasImages": false,
        "redacted": false,
        "unknownContext": false,
        "transcriptRecoverable": true
      },
      "requestBytes": 6017,
      "secretLeaked": false,
      "mode": "negative"
    },
    {
      "surface": "model fixture",
      "at": 1791317067.713004,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "toolRound": 2
    },
    {
      "surface": "model fixture",
      "at": 1791317067.7389581,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "toolRound": 1
    },
    {
      "surface": "model fixture",
      "at": 1791317067.797151,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 2
    },
    {
      "surface": "model fixture",
      "at": 1791317067.8072472,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 2
    },
    {
      "surface": "model fixture",
      "at": 1791317067.873649,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": false,
      "messageCount": 82
    },
    {
      "surface": "model fixture",
      "at": 1791317067.893216,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 2
    },
    {
      "surface": "model fixture",
      "at": 1791317067.89692,
      "path": "/v1/chat/completions",
      "model": "fixture-a",
      "summaryRequest": true,
      "messageCount": 2
    }
  ],
  "persistedCompactions": [
    {
      "type": "compaction",
      "id": "1f96c77d",
      "parentId": "8087e02c",
      "timestamp": "2026-10-06T20:04:27.811Z",
      "summary": "No prior history.\n\n---\n\n**Turn Context (split turn):**\n\n## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.",
      "shortSummary": "## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.",
      "firstKeptEntryId": "a4edcc72",
      "tokensBefore": 50030,
      "tokensAfter": 31718,
      "method": "soft",
      "details": {
        "readFiles": [],
        "modifiedFiles": []
      },
      "fromExtension": false
    },
    {
      "type": "compaction",
      "id": "4a2918fe",
      "parentId": "95d174db",
      "timestamp": "2026-10-06T20:04:27.898Z",
      "summary": "## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.",
      "shortSummary": "## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.",
      "firstKeptEntryId": "5c1de9d4",
      "tokensBefore": 50030,
      "tokensAfter": 20667,
      "method": "soft",
      "details": {
        "readFiles": [],
        "modifiedFiles": []
      },
      "fromExtension": false
    }
  ]
}
Evidence: Actual read-tool output and native recovery transcript

Source: Actual read-tool output and native recovery transcript

{"type":"ready","protocolVersion":1,"supportedProtocolVersions":[1,2],"maxFrameBytes":1048576,"maxReassembledFrameBytes":67108864}
{"type":"advisor_cost_changed"}
{"type":"available_commands_update","commands":[{"name":"security","description":"Plan, run, inspect, import, and compare OMP-native security scans","input":{"hint":"<plan|scan|status|cancel|scans|show|import|export|validate|compare|disposition>"},"subcommands":[{"name":"plan","description":"Create an immutable security scan plan"},{"name":"scan","description":"Start a planned or newly planned native scan"},{"name":"status","description":"Show native scan operation status"},{"name":"cancel","description":"Cancel a running native scan"},{"name":"scans","description":"List stored project security scans"},{"name":"show","description":"Render a scan or security:// resource"},{"name":"import","description":"Import SARIF or a Codex Security bundle"},{"name":"export","description":"Export a canonical bundle, SARIF, or report"},{"name":"validate","description":"Validate one finding with OMP-native tools"},{"name":"compare","description":"Compare finding lineage across two scans"},{"name":"disposition","description":"Set a finding disposition with rationale"}],"source":"builtin"},{"name":"model","aliases":["models"],"description":"Show current model selection","source":"builtin"},{"name":"switch","description":"Switch model for this session only","input":{"hint":"[model]"},"source":"builtin"},{"name":"fast","description":"Toggle fast mode","input":{"hint":"[on|ultra|off|status]"},"subcommands":[{"name":"on","description":"Enable fast mode (priority tier)"},{"name":"ultra","description":"Enable Ultrafast (OpenAI API, or Codex models that offer it)"},{"name":"off","description":"Disable fast mode"},{"name":"status","description":"Show fast mode status"}],"source":"builtin"},{"name":"slow","description":"Toggle slow mode","input":{"hint":"[on|off|status]"},"subcommands":[{"name":"on","description":"Flex tier, or Anthropic low priority at the session limit (auto)"},{"name":"off","description":"Standard service; stop Anthropic low priority"},{"name":"status","description":"Show slow mode status"}],"source":"builtin"},{"name":"skillful","description":"Toggle skill listing","input":{"hint":"[on|off|status]"},"subcommands":[{"name":"on","description":"List skills in the prompt for this session"},{"name":"off","description":"Omit the skills listing for this session"},{"name":"status","description":"Show skill listing status"}],"source":"builtin"},{"name":"extended-context","description":"Toggle extended context","input":{"hint":"[on|off|status]"},"subcommands":[{"name":"on","description":"Enable larger context windows"},{"name":"off","description":"Use default or standard-pricing context windows"},{"name":"status","description":"Show extended context status"}],"source":"builtin"},{"name":"computer","description":"Toggle computer use","input":{"hint":"[on|off|status]"},"subcommands":[{"name":"on","description":"Enable computer use for this session"},{"name":"off","description":"Disable computer use for this session"},{"name":"status","description":"Show computer use status"}],"source":"builtin"},{"name":"ratchet","description":"Build (or reuse) an eval for an LLM flow, then hillclimb it unattended","input":{"hint":"[flow and goal]"},"source":"builtin"},{"name":"prewalk","description":"Arm or restart prewalk","input":{"hint":"[restart]"},"subcommands":[{"name":"restart","description":"Return to @default and re-arm the handoff to @smol"}],"source":"builtin"},{"name":"modelpreset","description":"Manage model presets","input":{"hint":"[list|save|switch|delete] [name]"},"subcommands":[{"name":"list","description":"List saved presets"},{"name":"save","description":"Save the current role models and thinking level","usage":"<name>"},{"name":"switch","description":"Apply a saved preset","usage":"<name>"},{"name":"delete","description":"Delete a saved preset","usage":"<name>"}],"source":"builtin"},{"name":"effort","description":"Set or show reasoning effort (thinking level, intelligence)","input":{"hint":"[level]"},"subcommands":[{"name":"off","description":"No reasoning"},{"name":"minimal","description":"Very brief reasoning (~1k tokens)"},{"name":"low","description":"Light reasoning (~2k tokens)"},{"name":"medium","description":"Moderate reasoning (~8k tokens)"},{"name":"high","description":"Deep reasoning (~16k tokens)"},{"name":"xhigh","description":"Extended reasoning (~32k tokens)"},{"name":"max","description":"Maximum reasoning the model supports"},{"name":"auto","description":"Auto-detect per prompt"}],"source":"builtin"},{"name":"advisor","description":"Toggle advisor","input":{"hint":"[on|off|status|dump [raw]|configure]"},"subcommands":[{"name":"on","description":"Enable the advisor"},{"name":"off","description":"Disable the advisor"},{"name":"status","description":"Show advisor status"},{"name":"dump","description":"Copy the advisor's transcript to clipboard","usage":"[raw]"},{"name":"configure","description":"Open the advisor configuration editor (TUI)"}],"source":"builtin"},{"name":"export","description":"Export session to HTML file","input":{"hint":"[--themes] [path]"},"source":"builtin"},{"name":"trace","description":"Open this session's trace in the stats dashboard","source":"builtin"},{"name":"dump","description":"Return full transcript as plain text, with LLM request JSON path","input":{"hint":"[all]"},"subcommands":[{"name":"all","description":"Write a zip with the main transcript, LLM request JSON, and one file per subagent"}],"source":"builtin"},{"name":"share","description":"Share session via an encrypted link (share server or secret gist)","source":"builtin"},{"name":"browser","description":"Toggle browser eval-prelude headless vs visible mode","input":{"hint":"[headless|visible]"},"subcommands":[{"name":"headless","description":"Switch to headless mode"},{"name":"visible","description":"Switch to visible mode"}],"source":"builtin"},{"name":"todo","description":"Manage todos","input":{"hint":"<subcommand>"},"subcommands":[{"name":"edit","description":"Open todos in $EDITOR (Markdown round-trip)"},{"name":"copy","description":"Copy todos as Markdown to clipboard"},{"name":"expand","description":"Show every phase and task in the HUD"},{"name":"collapse","description":"Restore the bounded HUD preview"},{"name":"export","description":"Write todos as Markdown to a file (default: TODO.md)","usage":"[<path>]"},{"name":"import","description":"Replace todos from a Markdown file (default: TODO.md)","usage":"[<path>]"},{"name":"append","description":"Append a task; phase fuzzy-matched or auto-created","usage":"[<phase>] <task...>"},{"name":"start","description":"Mark task in_progress (fuzzy-matched)","usage":"<task>"},{"name":"done","description":"Mark task/phase/all completed (fuzzy-matched)","usage":"[<task|phase>]"},{"name":"drop","description":"Mark task/phase/all abandoned (fuzzy-matched)","usage":"[<task|phase>]"},{"name":"rm","description":"Remove task/phase/all (fuzzy-matched)","usage":"[<task|phase>]"}],"source":"builtin"},{"name":"session","description":"Show or configure the current session","input":{"hint":"[info|delete|pin [account]]"},"subcommands":[{"name":"info","description":"Show session info and stats"},{"name":"delete","description":"Delete current session and return to selector"},{"name":"pin","description":"Pin the current provider to a stored OAuth account","usage":"[account]"}],"source":"builtin"},{"name":"jobs","description":"Show background jobs","input":{"hint":"[full]"},"subcommands":[{"name":"full","description":"Show full, untruncated command lines"}],"source":"builtin"},{"name":"usage","description":"Show token usage","input":{"hint":"[show|reset [provider/credential-id|provider/active]]"},"subcommands":[{"name":"show","description":"Show provider usage and limits"},{"name":"reset","description":"Spend a saved provider rate-limit reset","usage":"[provider/credential-id|provider/active]"}],"source":"builtin"},{"name":"stats","description":"Launch the local stats dashboard","input":{"hint":"[--port <port>] [--host <host>]"},"source":"builtin"},{"name":"changelog

... [36824 bytes truncated] ...

xture"},"messageId":"msg-8"}
{"type":"message_update","assistantMessageEvent":{"type":"text_start","contentIndex":0,"partial":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture"}},"message":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture"},"messageId":"msg-8"}
{"type":"message_update","assistantMessageEvent":{"type":"text_delta","contentIndex":0,"delta":"Checkpoint complete. All assigned work is finished and reported.","partial":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture"}},"message":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture"},"messageId":"msg-8"}
{"type":"message_update","assistantMessageEvent":{"type":"text_end","contentIndex":0,"content":"Checkpoint complete. All assigned work is finished and reported.","partial":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture"}},"message":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"totalTokens":0,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture"},"messageId":"msg-8"}
{"type":"message_end","message":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":50000,"output":30,"cacheRead":0,"cacheWrite":0,"totalTokens":50030,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture","duration":11.160415999999714,"ttft":10.731915999999728,"completedAt":1791317067877},"messageId":"msg-8"}
{"type":"turn_end","message":{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":50000,"output":30,"cacheRead":0,"cacheWrite":0,"totalTokens":50030,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture","duration":11.160415999999714,"ttft":10.731915999999728},"toolResults":[]}
{"type":"auto_compaction_start","reason":"threshold","action":"context-full"}
{"type":"auto_compaction_end","action":"context-full","result":{"summary":"## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.","shortSummary":"## Goal\nComplete isolated test checkpoint.\n## Progress\nCheckpoint completed. Preserve this summary and recent native tail.\n## Next Steps\nWait for the user.","firstKeptEntryId":"5c1de9d4","tokensBefore":50030,"details":{"readFiles":[],"modifiedFiles":[]}},"aborted":false,"willRetry":false}
{"type":"agent_end","messages":[{"role":"user","content":[{"type":"text","text":"Continue the unfinished unit with two reads of the disposable checkpoint."}],"attribution":"user","timestamp":1791317067679},{"role":"assistant","content":[{"type":"toolCall","id":"checkpoint-2","name":"read","arguments":{"path":"~/.no-mistakes/worktrees/32d18ed9638d/01M49C3HQSKS56S7YCKMMW32FM/.fm-adviser-root/lab-bounded-tool-loop/data/checkpoint.txt"},"partialArgs":"{\"path\": \"~/.no-mistakes/worktrees/32d18ed9638d/01M49C3HQSKS56S7YCKMMW32FM/.fm-adviser-root/lab-bounded-tool-loop/data/checkpoint.txt\"}","streamIndex":0}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":50000,"output":30,"cacheRead":0,"cacheWrite":0,"totalTokens":50030,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1791317067709,"responseId":"tool","duration":4.408249999999953,"ttft":3.6677920000001905},{"role":"toolResult","toolCallId":"checkpoint-2","toolName":"read","content":[{"type":"text","text":"[Superseded by a newer read of this file]"}],"details":{"totalLines":1,"displayContent":{"text":"Disposable checkpoint record for actual omp read tool.","startLine":1,"lineNumbers":[1]},"fileSize":55,"meta":{"source":{"type":"path","value":"~/.no-mistakes/worktrees/32d18ed9638d/01M49C3HQSKS56S7YCKMMW32FM/.fm-adviser-root/lab-bounded-tool-loop/data/checkpoint.txt"}}},"isError":false,"timestamp":1791317067720,"prunedAt":1791317067879},{"role":"assistant","content":[{"type":"toolCall","id":"checkpoint-1","name":"read","arguments":{"path":"~/.no-mistakes/worktrees/32d18ed9638d/01M49C3HQSKS56S7YCKMMW32FM/.fm-adviser-root/lab-bounded-tool-loop/data/checkpoint.txt"},"partialArgs":"{\"path\": \"~/.no-mistakes/worktrees/32d18ed9638d/01M49C3HQSKS56S7YCKMMW32FM/.fm-adviser-root/lab-bounded-tool-loop/data/checkpoint.txt\"}","streamIndex":0}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":50000,"output":30,"cacheRead":0,"cacheWrite":0,"totalTokens":50030,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1791317067736,"responseId":"tool","duration":2.907832999999755,"ttft":2.728541999999834},{"role":"toolResult","toolCallId":"checkpoint-1","toolName":"read","content":[{"type":"text","text":"Disposable checkpoint record for actual omp read tool."}],"details":{"totalLines":1,"displayContent":{"text":"Disposable checkpoint record for actual omp read tool.","startLine":1,"lineNumbers":[1]},"fileSize":55,"meta":{"source":{"type":"path","value":"~/.no-mistakes/worktrees/32d18ed9638d/01M49C3HQSKS56S7YCKMMW32FM/.fm-adviser-root/lab-bounded-tool-loop/data/checkpoint.txt"}}},"isError":false,"timestamp":1791317067744},{"role":"assistant","content":[{"type":"text","text":"Checkpoint complete. All assigned work is finished and reported."}],"api":"openai-completions","provider":"lab","model":"fixture-a","usage":{"input":50000,"output":30,"cacheRead":0,"cacheWrite":0,"totalTokens":50030,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"stop","timestamp":1791317067865,"responseId":"fixture","duration":11.160415999999714,"ttft":10.731915999999728}],"isTerminal":true,"yielded":true}
{"type":"prompt_result","id":"6bc11d94","agentInvoked":true,"status":"completed","sessionSettled":true}
{"type":"session_settled"}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 11 of 11 scenarios driven live against the product
Scenario Result Live Evidence
An explicitly enabled, qualifying settled checkpoint invokes omp's native compactor ✅ pass live positive.json: qualifying judgment, checkpoint request and persisted soft compaction from 50030 to 31291 tokens with fromExtension=false.
Loading the entry without opt-in or acknowledged auto consent makes no Jev request ✅ pass live optin-off.json and consent-off.json: completed real product turns with no Jev traffic or eager compaction.
An unfinished judgment defers once within an agent loop, then permits threshold recovery ✅ pass live bounded-tool-loop.json and bounded-tool-loop-product.txt: two successful real read calls; one cancellation after agent_start, followed by native persisted compaction.
A timed-out Jev request does not block automatic native compaction ✅ pass live timeout-native.json: judge-unavailable after the request timeout, no native-deferred event, and subsequent native threshold persistence.
New user input invalidates a pending qualifying judgment ✅ pass live pending-new-input.json: pending-invalidated after a new prompt and no checkpoint compaction after the delayed positive response.
Native model selection invalidates a pending qualifying judgment ✅ pass live pending-model-switch.json: real set_model to fixture-b, pending-invalidated, and no stale checkpoint compaction.
Near the context-window limit, unfinished advice cannot veto native recovery ✅ pass live near-window-precedence.json: unfinished judgment at 260000 context tokens, no advisory deferral, and native persisted compaction.
A healthy unfinished judgment cannot block overflow or incomplete-turn recovery ✅ pass live healthy-overflow.json and incomplete-native.json: actual host-classified overflow/incomplete events, native-precedence records and persisted native soft compactions.
A failed current turn cannot borrow an older successful answer as a checkpoint ✅ pass live failed-terminal.json and failed-terminal-product.txt: actual terminal assistant error despite older successful history, with no Jev request or checkpoint compaction.
Structured secret-bearing context cannot authorize eager compaction ✅ pass live redacted-coverage.json: escaped and nested synthetic saved-key fields produced coverage-ineligible, with no Jev traffic.
The genuine TUI idle timer respects unfinished advice while explicit compact remains native ✅ pass live tui-idle.json and tui-idle-product.txt: idle deferral after 60 seconds, visible maintenance cancellation, then manual native persistence from 50030 to 31330 tokens.
  • bash bin/fm-test-run.sh tests/fm-omp-jev-pipeline.test.sh
  • FM_JEV_ADVISER_DIR="$PWD/.fm-adviser-root/node_modules/compact-adviser" bash bin/fm-test-run.sh tests/fm-omp-jev-pipeline.test.sh
  • Verified the authorized adviser archive SHA-256 and installed it with npm install --prefix .fm-adviser-root ~/firstmate/data/fm-compact-adviser-omp/compact-adviser-0.1.11.tgz; copied the static entry into that isolated dependency root.
  • Created disposable marked homes with bin/fm-lab-home.sh create and private sockets with bin/fm-lab-home.sh tmux-dir; launched real omp through private tmux -L fm-lab sessions with isolated HOME, FM_HOME and PI_CODING_AGENT_DIR.
  • Drove the documented static extension through real omp RPC using prompt, set_auto_compaction, compact and set_model requests against disposable loopback model/Jev endpoints.
  • Executed two successful built-in read calls within one actual agent loop and observed native threshold compaction after the in-loop advisory deferral.
  • Drove actual provider timeout, context-length error, incomplete length response, permanent failed turn, consent/opt-in guards, structured-secret context, and pending-request invalidation scenarios.
  • Launched the genuine interactive TUI in a 180x50 tmux grid, observed its 60-second idle-compaction attempt, then executed /compact Preserve the isolated checkpoint and recent native history.
  • Retained product transcripts, controller metrics and persisted native compaction entries; stopped private tmux servers and the fixture endpoint, removed all phase-created dependencies/labs/caches, and verified evidence JSON readability and fixture removal.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@MrGTV-love
MrGTV-love merged commit ea4e981 into main Oct 6, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant