Conversation
added 6 commits
September 24, 2026 12:50
… expose locked skips
…ator documentation
Author
|
Hi! This PR comes from a fork, so the CI and Require no-mistakes workflows are waiting for maintainer approval before they can run. When you have a moment, could you approve the workflow runs? Thanks! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Linear HHE-1793 (operator-approved): reduce full-context root inspection for repetitive free-text worker events with one small opt-in JEV pilot. Choose one observed recurring event class. Exact quota, trust, CI and process facts stay deterministic. Classify the remaining semantic uncertainty into a closed action/attention set that includes unknown, reusing the existing event entry point and code policy. Begin shadow-only with representative normal, uncertain and misleading events, and record the frontier decisions that could have been avoided.
Constraints: no new scheduler or timer; no suppression of actionable wakes during shadow; no automatic merge, relaunch or credential grants; no fabricated completion. Cache only identical state/schema/model and invalidate on evidence changes. Independent questions may be batched, with real request costs measured.
Done: a working opt-in shadow adapter, reproducible replay/live sanitized evidence, confusion and error examples, and an explicit recommendation on whether any low-risk behaviour qualifies for activation. The shadow result is not approval for autonomous action. Keep it one PR-sized pilot.
What Changed
declared_wait,inspect, andunknownoutcomes; normal wake handling and acknowledgement remain unchanged.Risk Assessment
✅ Low: The opt-in pilot remains shadow-only, preserves actionable wakes, and keeps replay evidence distinct from live-call provenance; no material source-backed issue was found.
Testing
The focused interface test exercised disabled and enabled drains, wake preservation, batching, evidence bounds, confidence abstention, error handling, and journal safeguards using a fake API. An offline replay through the real CLI reproduced the recorded response: five abstentions, no scored errors, two hypothetical frontier candidates, and zero actual decisions avoided. No new paid request was made; the replay retained recorded token counts. The API-classification behavior was not driven live against Typesafe in this run.
Evidence: Offline shadow replay and confusion report
Source: Offline shadow replay and confusion report
Evidence: Shadow-only annotation
Source: Shadow-only annotation
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed ✅
bin/fm-event-shadow.sh:86- When the last eight status lines exceed 4096 bytes,head -ckeeps the oldest bytes and drops the newest declaration. For example, a longpaused: waitingline followed byblocked: need assistancecan reach the model as only a wait; a high-confidencedeclared_waitresponse then records a false frontier candidate without error. Keep the newest bounded evidence, and mark a clipped declaration as unknown rather than treating it as complete.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-event-shadow.test.shFM_STATE_OVERRIDE=<isolated evidence state> bin/fm-event-shadow-replay.sh --response tests/fixtures/event-shadow/live-response.json✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.