Repository navigation
Secret-storage practice: per-host 0600 store, agent read guard hook + deny rules, managed-profile install, names-only checker - #182
Merged
Conversation
…ker, agent and commit guards
- docs/secret-storage.md runbook and docs/decisions/2026-09-24-secret-storage.md
record: per-provider 0600 files in ${XDG_CONFIG_HOME:-$HOME/.config}/native-agent-stack
(0700, outside every worktree), pointer-only shell variables for pickup,
native sign-ins left native, honest threat model (deny rules and hook stop
accidents, not a same-uid agent), WSL2/proc/output-store boundaries,
rotation and incident steps.
- adoption/credential-inventory.json: names, class, lane, status, path
template and loaders; no values. Validated by scripts/validate.py.
- scripts/credential_status.py: lstat/Git-index/env-name checks only; never
opens a credential file; exit 1 only for an unsafe required entry.
- .claude/settings.json deny rules plus scripts/hooks/secret_path_guard.py
PreToolUse Bash hook; scripts/git-hooks/pre-commit gitleaks gate (fails
closed); .gitignore credential-shaped names; env templates.
- Tests use synthetic sentinel values and assert they never appear in output.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…profile deployment Fix the independent review's medium findings and ship the guards to new hosts. - Codex: document [shell_environment_policy] inherit = "none" plus explicit non-secret set entries as the primary snippet, citing the gap-wave2 canary receipt (write_receipts.py:181-187) and the official Codex config reference; "core" is recorded as unmeasured. --client-guards counts only inherit == "none". State that the policy controls environment inheritance, not file reads. - secret_path_guard.py: block reader/search commands (grep, rg, ag, ack, git grep, find -exec with a reader, awk, sed, ...) on secret variable names, *.env files, pointer variables and the XDG store spelling; stdin redirection from a pointer variable; shell tracing or verbose mode combined with sourcing a credential file; env/printenv/export -p/ declare -p/-x and inline environ dumps after sourcing; /proc environ in any spelling. Write targets (cp destination, tee, > redirections) are not treated as reads. Tests cover each case plus a 66-command safe corpus. - credential_status.py --client-guards: booleans for the Claude Code OpenTelemetry content flags (key names and truthiness only), the user secret-guard hook, and Codex inherit=none. Measured on this host 2026-09-24: telemetry on, all five content flags false; the earlier doc claim that they were enabled is withdrawn. Pasted keys must be rotated. - Claude user profile: install_claude_profile.py installs the secret guard to ~/.claude/hooks/ (sha256-pinned; SHA256SUMS paths are relative, so sha256sum -c works), and the settings template carries the deny rules and a PreToolUse Bash hook that is inert until the guard is installed. Project .claude/settings.json is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tory secret name - Block `git credential fill`, `git credential-<helper> get`, `git-credential-* get` and `gh auth git-credential` (reason native_token_print). After the documented `gh auth setup-git` each of these prints the GitHub token. Matching is by parsed command words, so a commit message that mentions the form still passes. - Add deny rules `Bash(git credential fill*)` and `Bash(gh auth git-credential *)` to the project settings, the user-profile template and the manual merge list. - Add the inventory's must_not_be_set names the hook lacked (OPENROUTER, QDRANT, MISTRAL, MASSIVE, PREFECT, MC, MSB, PAPERCLIP API keys; TWS_USERNAME, TWS_ACCOUNT, IBKR_ACCOUNT_ID). A new test asserts that every must_not_be_set name and every secret-class `variables` name of adoption/credential-inventory.json is in SECRET_NAMES, so the list cannot drift again; the hook still reads no file. - Refresh the hook's SHA256SUMS pin and the evidence manifest hashes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
CodeQL py/clear-text-logging-sensitive-data traced the boolean returned by user_secret_guard_hook() into the report print: its source was the function name, not a value. No credential value could reach output, but the code is now restructured so no value-bearing variable is reachable from print: settings env entries are reduced to key names plus truthiness in enabled_names(), parsed Claude/Codex settings stay local to helpers that return booleans only, and client_guards() fails closed unless every result is a fixed key with a boolean or null value. Tests use neutrally named, randomly generated fake sentinels and add a CLI check across plain, --json and --client-guards modes asserting the fake values never appear in stdout or stderr. The evidence manifest records the new bytes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review follow-ups on #182: claude_guard_booleans() also catches TypeError from well-formed settings with wrong field types (returns None instead of a traceback that reads like an unsafe result). The any-mode CLI test now asserts exit 0, no traceback, the exported name reported by name, and the text markers, so a crash can no longer pass it. New test for wrong-type settings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 24, 2026
Merged
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 24, 2026
…#94 cleanup, sync example agents - The decision record names v2026.09.24.1 and the settings file main added (#182). - The macOS step 4 re-pin note and next-host-stages' 'as the macOS page marks' go, since #94 is in the pinned tag. - The four example agents are byte-identical to agent-lab's again (source-scout maxTurns 100 and the foreground-acceptance sentence), as the recipe states. - The install-input docstring and comment name the secret-path guard's SHA256SUMS coverage. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 24, 2026
…staller agents, examples, committed project settings) (#196) * Run every Claude child at max effort under an Ultracode main loop Set effort: max in the five shipped agent definitions (adoption/agents/claude/*.md), keeping their task-matched models; the one-line change keeps them byte-identical to agent-lab's .claude/agents copies, which receive the same edit. Commit a project .claude/settings.json (byte-identical to the portable ultracode example and agent-lab's committed project settings) so single-repository cloud sessions on this repository start with Ultracode. It carries neither effortLevel (a max there is silently dropped) nor CLAUDE_CODE_EFFORT_LEVEL (any value but xhigh disables Ultracode, and max there overrides every child's effort). Update the Ultracode recipe: every child role at max with its task-matched model, the coordinator at Ultracode (xhigh plus orchestration), the measured constraints (max disables orchestration, cannot be persisted, and the environment variable overrides all child efforts), the per-stage rule (agent() calls pass effort: 'max'; a stage without it inherits xhigh) and GitHub/cloud-session guidance (claude-code-action has no effort input: use claude_args '--effort ultracode' or settings {"ultracode":true}, never '--effort max' beside Ultracode). Note the child-effort change in the native profile recipe. Record the decision with its probes, alternatives and overturn conditions in docs/decisions/2026-09-23-max-effort-default.md, linked from the Claude user profile decision, and rehash the two recipes in manifests/evidence.json. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Carry max effort through the examples, vendored lane and docs; fix the review findings Rebased onto the current main (v2026.09.23.1 pinned by #148). The five adoption/agents/claude definitions differ from that release, so adoption/bootstrap.md now says they "changed after `v2026.09.23.1`", and tests/test_adoption_docs_consistency.py treats what install_claude_profile.py copies (the agent and hook directories, the MCP template) as install inputs, matching a directory by its path, never its last component. The stale v2026.09.23 notes are dropped from the bootstrap, platform, README and next-host pages: the pinned release contains everything they described. Examples: the four example agents and every agent() stage of review-changes.js, readiness-audit.js and layer-verdict-lane.js bind effort max with models unchanged; the three scripts are byte-identical to agent-lab 915e73e. The vendored lane now names blind-lane-reviewer, vendored beside it and pinned under required_agents in vendored-lanes.json; SHA256SUMS, vendored-lanes.json and two appended lane-provenance.json entries rebind the new bytes. The contract suite pins max in ROUTING (which also clears the routing failure the old vendored lane already had on main) and now rejects a stage or agent at any other effort, a second effort line, MODEL drift, a routing table that restates an agent differently, CLAUDE_CODE_EFFORT_LEVEL or an effort cap in the settings, and instructions without the `effort: 'max'` literal; of 24 new mutations, 22 fail on their named assertion and two harmless edits do not. examples/claude-native/CLAUDE.md takes the host's current worker sentence, and AGENTS.md tells the coordinator to pass effort: 'max' on ad-hoc agent() calls and to stay at xhigh. Recipes and decision record: GitHub runs are non-interactive Agent SDK runs whose Workflow launches go through normal permission evaluation; autoContinueAtUsageLimit is no cost control and headless, SDK, GitHub and background runs fail at a usage limit; the claude-code-action pin is "read 2026-09-23" without an input count; the action's one CI smoke is cited and no Ultracode or effort run is claimed; /effort max is marked as not probed; Q1/Q2 replace the confounded P3/P4/P7/P8 as the evidence that a persisted max is dropped, and Q3 adds that a stage's effort wins over the frontmatter; CLAUDE_CODE_EFFORT_LEVEL is never set. catalogs/foundation/decisions.json's lean-workflow-child-routing activation names the max efforts and cites the decision. A new test requires one "effort: max" frontmatter line in each shipped agent. Checks (raw): validate.py, validate_catalogs.py, evidence_manifest.py --check, new_host_grand_list.py --check, build_ecosystem.py --check and release_due.py --strict-if-repinned origin/main (exit 0, not strict: no re-pin, nothing due) pass; the example suites pass as CI runs them (sha256sum --check) and as their README runs them (envelope 207/207, mutations 47/47, child-usage 16/16, usage-receipts 18/18, codex-envelope, check-syntax); the named unittest modules and the related ones pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Run the example contract suites in CI; fix the settings-example, adopt-step and concurrency wording Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Sync source-scout (maxTurns 100, foreground acceptance) and isolated-builder from agent-lab #46 (a7a67a3e) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Rebase onto #145 and v2026.09.24.1: merge the secret guard into project settings, blind-adjudicator routing row, current agent-lab pins - .claude/settings.json keeps main's secret-read deny rules and guard hook and adds the ultracode keys; the recipe and decision no longer call it byte-identical. - bootstrap.md marks adoption/agents/claude and adoption/hooks/claude as changed after v2026.09.24.1; the lane test and lane files are #145's. - child-usage mirror synced from agent-lab; examples routing table lists blind-adjudicator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Resolve the rebase review: current pin in the record, finish the macOS #94 cleanup, sync example agents - The decision record names v2026.09.24.1 and the settings file main added (#182). - The macOS step 4 re-pin note and next-host-stages' 'as the macOS page marks' go, since #94 is in the pinned tag. - The four example agents are byte-identical to agent-lab's again (source-scout maxTurns 100 and the foreground-acceptance sentence), as the recipe states. - The install-input docstring and comment name the secret-path guard's SHA256SUMS coverage. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 27, 2026
Preserve the historical RTK aggregate tables, synthetic exactness output and method, with private sample fragments removed and a dated M-R1 interpretation erratum. Publish 18 source-review currency records with live release metadata, explicit evidence classes and retrieval dates. Preserve scratch/current differences, including Headroom 0.39.1 and the context-hub PR #182 correction. Sources: rtk-ai/rtk v0.50.0 (1d87b8e719ce0a50c223cd93ca64dd16921f9aec), src/main.rs, src/hooks/decision.rs, src/discover/mod.rs and registry.rs; rtk-ai/rtk dev-0.51.0-rc.467 (a89a31494670fcec8ffa20d939dd94c64bd998fb). The 18 upstream repositories, pins and finding URLs are retained in evidence/artifacts/token-tool-currency-20260927/records/*.json. Release retrieval: https://docs.github.com/en/rest/releases/releases#get-the-latest-release and https://cli.github.com/manual/gh_api. Publication tests: https://docs.python.org/3/library/unittest.html. Leak scan: gitleaks/gitleaks v8.30.1 README.md. Validation: 5 offline unittest checks passed; native Gitleaks and identifier scans found no leaks. scripts/validate.py reports only the 25 new evidence files awaiting the coordinator's hash registration. No pins or runtime configuration changed. Recommended label: lane:foundation Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 29, 2026
…s (sanitized, revision-bound) (#409) * Publish sanitized RTK and token-tool currency evidence Preserve the historical RTK aggregate tables, synthetic exactness output and method, with private sample fragments removed and a dated M-R1 interpretation erratum. Publish 18 source-review currency records with live release metadata, explicit evidence classes and retrieval dates. Preserve scratch/current differences, including Headroom 0.39.1 and the context-hub PR #182 correction. Sources: rtk-ai/rtk v0.50.0 (1d87b8e719ce0a50c223cd93ca64dd16921f9aec), src/main.rs, src/hooks/decision.rs, src/discover/mod.rs and registry.rs; rtk-ai/rtk dev-0.51.0-rc.467 (a89a31494670fcec8ffa20d939dd94c64bd998fb). The 18 upstream repositories, pins and finding URLs are retained in evidence/artifacts/token-tool-currency-20260927/records/*.json. Release retrieval: https://docs.github.com/en/rest/releases/releases#get-the-latest-release and https://cli.github.com/manual/gh_api. Publication tests: https://docs.python.org/3/library/unittest.html. Leak scan: gitleaks/gitleaks v8.30.1 README.md. Validation: 5 offline unittest checks passed; native Gitleaks and identifier scans found no leaks. scripts/validate.py reports only the 25 new evidence files awaiting the coordinator's hash registration. No pins or runtime configuration changed. Recommended label: lane:foundation Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Correct evidence provenance and exercise publication controls Attribute RTK T7 to rc.467, qualify query timestamps and Serena's main distance, correct all metadata coordinates against repository commit 1c32ad2, and add discriminating unittest controls. Preserve historical output bytes and scratch records. Sources: rtk-ai/rtk v0.50.0 (1d87b8e719ce0a50c223cd93ca64dd16921f9aec), dev-0.51.0-rc.467 (a89a31494670fcec8ffa20d939dd94c64bd998fb), Cargo.toml:3; full-save/rtk/REPORT.md:37-41,193,205,270 and exactness.sh; Python unittest https://docs.python.org/3/library/unittest.html; Gitleaks v8.30.1 README; docs/acceptance-evidence-policy.md and docs/lanes.md. Refresh the PR evidence to distinguish reviewed a0ac1cde registration and passing harness logs from the unregistered b0fc8f5c repair checkout. Leave manifest registration, report generation and commits to the harness. Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Currency test reads the recorded metadata revision; METHOD states batch stamps and revision scope Coordinator corrections after the independent repair verification: test_pin_metadata_lines_contain_the_component_version now reads each cited file with git show <metadata_revision>:<path> (the records cite coordinates at 1c32ad2, which later changes may move); METHOD.md describes the two retrieval batches separately and no longer claims the revision's files match every later tree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Hot files: evidence registration and generated reports Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test_token_full_save_evidence: fail, never skip, a bad metadata coordinate (G1) The nested lines_at helper turned every failed git show into skipTest, so a nonexistent metadata path or revision bypassed provenance validation. The module-level metadata_lines helper follows tests/test_release_pin_contents.py (git cat-file -e <rev>^{commit}; fail in CI, skip locally): an absent revision skips only when GITHUB_ACTIONS is not "true" (validate.yml checks out full history, fetch-depth: 0), and a failed git show at an existing revision fails with the revision, path and git's first stderr line. Discriminating controls (red first against the old skip: FAILED (failures=3), 'skip' != 'fail'): a bogus path at HEAD fails with and without GITHUB_ACTIONS, and a missing revision skips locally but fails with GITHUB_ACTIONS=true. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * rtk-coverage-study METHOD: attribute the population to the unpublished report; record the bucket recount (G2, G3) G2: the transcript-file, subagent and window figures appear in no published aggregate; a lead-in now says the unpublished REPORT.md states every figure in that paragraph and that only the 59,041-call total is also retained in tables.txt and frame-summary.txt. The four original sentences are unchanged. G3: a scripted recount of the 22 bucket rows in frame-summary.txt gives 15,623 missed parts (matching TOTAL and defer+deny+NO_ROW) and 25,947 covered parts, 7 fewer than the TOTAL line and ask row (25,954). One dated sentence records both sums and that the published files do not attribute the difference. The data files are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * token-tool-currency METHOD: bound main distances without a retained main observation (G4) Of the nine batch-a records, context-mode, ccusage, qmd, repomix and toon state current-main, commits-ahead or unchanged main distances with no unreleased_main block and no returned head of main. context-mode (10 -> 12) and ccusage (183 -> 198) retain counts from compares ending at fixed commits, with nothing showing those commits were main's head; qmd, repomix and toon name a compare of the mutable main ref whose returned head and count were not retained, so their only retained distances are the 2026-09-26 scratch_record figures. One dated limitation paragraph names them and requires re-querying before use. headroom and rtk make no main-distance claim; markitdown and serena already bound theirs. No record changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence manifest: re-register the three files repaired for PR #409 review Hot-file protocol (docs/lanes.md): last commit only. register_file refreshes the hashes of tests/test_token_full_save_evidence.py and the two METHOD.md files; component_matrix.py --write and new_host_grand_list.py --write produced byte-identical reports. validate.py, evidence_manifest.py --check and component_matrix.py --check pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * token-tool-currency: README and METHOD state only the main-head observations each record retains The README said the records retain the observed heads, but only jcodemunch-mcp, codebase-memory-mcp and agentsview keep an unreleased_main head; the other five keep a 2026-09-26 scratch head and compare end commits. The METHOD limitation now names those per record. No number changed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * evidence manifest: re-register the two token-tool-currency files edited for the wording fix (manifests/evidence.json only) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds one secret-storage practice for every host. Secret values never reach git or GitHub, agents are blocked from reading them, and future sessions and new PCs can pick them up without extra steps.
Store.
${XDG_CONFIG_HOME:-$HOME/.config}/native-agent-stack/(mode 0700). For examplealpaca-paper.envandsec-contact.env. It is loaded by explicit path (--env-file, orset -a; . file; set +ainside the command)./proc/<pid>/environand child processes.runner.credentials) and extends it.Agent guardrails. Deny rules alone do not stop
Bash cat, so they are paired with a hook.scripts/hooks/secret_path_guard.pyis a deterministic PreToolUse hook. It blocks:*.envfiles, pointer variables or listed secret names;/proc/*/environ;git credential fillroutes..claude/settings.jsonaddspermissions.denyrules.install_claude_profile.py --only guardplusclaude.settings.template.json) installs the same hook and rules on every new PC. Both hooks are pinned by SHA-256 and checked before copying.[shell_environment_policy] inherit = "none", which is the setting the repo measured, and state plainly that this controls environment inheritance, not file reads.GitHub and git.
scripts/git-hooks/pre-commitruns gitleaks, in addition to the gitleaks job in CI..gitignoregains more patterns.Inventory and checker.
adoption/credential-inventory.jsonlists names, classes, lanes, statuses, store paths and loaders, and no values.scripts/credential_status.pyuses lstat and variable names only, and never opens a value into its output (tests use fake values). With--client-guardsit reports booleans only, including whether Claude telemetry logs content.docs/secret-storage.mdanddocs/decisions/2026-09-24-secret-storage.md. The decision record lists the alternatives rejected: systemd-creds, sops/age, pass, keyring, the 1Password and Bitwarden CLIs, and exported environment variables.Credential status (from the inventory):
Verification
validate.py,validate_catalogs.py,landscape.py,evidence_manifest.py --checkandverdict_review_gate("no verdict rows changed") all pass.credential_status.pyon this host: both required stores areok(0600 in a 0700 directory), no broker variable is set in the environment, and no sensitive file is tracked.test_effort_default_guard.test_shipped_guard_is_verbatim, comes from a stale host copy of that hook.Review. The design was adversarially refuted twice before the build: the deny rules alone did not stop
Bash cat, and Codex has no per-path read deny. Two independent review rounds followed, with every medium finding fixed in9c4a2d5and0936894.Host step after merge:
🤖 Generated with Claude Code