Skip to content

Telemetry writer identity (G1): one writer per token series, integrity alerts and a committed host recipe - #364

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
claude/monitor-metrics-integrity-20260926
Sep 27, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
claude/monitor-metrics-integrity-20260926

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Summary

Each Prometheus token series now has one writer. Before this change, every Claude process and every
directly started Codex process wrote without a service.instance.id, so concurrent cumulative streams
overwrote each other. On the workstation that meant 7,216 Claude counter resets in an hour,
Prometheus increase() at 73 to 193 times the Loki api_request sums, and about 3,518 of 8,160 delta
points rejected by delta_to_cumulative (read-only receipt host-before-20260926.json).

  • Claude: session ids are exported again (OTEL_METRICS_INCLUDE_SESSION_ID=true, the upstream
    default). The Collector turns them into service.instance.id (groupbyattrs/session, then
    transform/privacy) and sums the points of dropped attributes (aggregate_on_attributes).
  • Codex: observability/collector/codex-identity-launcher.sh.example gives each process its own id
    through OTEL_RESOURCE_ATTRIBUTES. This PR's host apply leaves it out; see Residuals.
  • Prometheus: created-timestamp-zero-ingestion and promql-extended-range-selectors. The
    collector-native job drops the Codex histogram buckets except token usage.
  • Alerts: the new native-telemetry-integrity group covers per-series resets, dropped delta
    points and unscoped writers. The dashboards exclude instance="unscoped".
  • Host recipe: committed with failing-first tests under
    evidence/artifacts/telemetry-writer-identity-20260926/host/ (apply.sh, rollback.sh,
    prove.sh/prove.py and three merge/setting helpers).

Scope

  • What this PR changes: the Collector metrics profile, both Claude settings templates, the Prometheus
    unit flags, scrape config and rules, the dashboards, the host recipe, the READMEs and new-host
    pages that describe them (marked "changed after v2026.09.26.2"), and their tests.
  • Base commit: 803bc351. git merge-tree against origin/main dde28cc2 is clean.
  • Lane: lane:foundation. manifests/evidence.json is only re-registered.
  • Owned paths touched: observability/collector/, observability/backends/,
    observability/native-data/, observability/grand-dashboard/render.py, adoption/,
    docs/decisions/, evidence/artifacts/telemetry-writer-identity-20260926/ and tests/.
  • Stacked on this PR: Tool invoke rates: tool, MCP server, skill and subagent names in Loki and Grafana (stacked on #364) #366, the tool invoke-rate change (branch claude/tool-invoke-rates-20260926). It retargets to main after this PR merges.

What was verified

Claim Evidence class Command / receipt
Repository checks pass at dfd5460f local integration scripts/validate.py (6,934 hashed files, 159 receipts), evidence_manifest.py --check, component_matrix.py --check, new_host_grand_list.py --check and build_ecosystem.py --check all exit 0
Unit and related modules synthetic 116 OK, 0 skipped: writer_identity, writer_identity_host, backends_alerts, grand_dashboard and native_dashboard_data. The otelcol-contrib 0.161.0 and promtool 3.15.0 cases ran.
Full suite synthetic unittest discover -s tests -q: 6,281 tests OK (skipped=654), first run
Queries parse on the pinned backends local integration the pinned Prometheus 3.15.0 parsed 20 of 20 prove.py queries on a scratch TSDB with both features on; the loopback Loki 3.7.8 accepted 16 of 16 LogQL queries
The new profile removes the overwrite synthetic (real otelcol-contrib 0.161.0 and Prometheus 3.15.0 on loopback) synthetic-ab-20260926.json: old profile Claude −51.3% (anchored), 4 resets, 24 ErrOutOfOrder points, Codex −6.7%; new profile 0.0% anchored for both, 0 resets, 0 dropped points
Host scripts are read-only until --apply host observation real-host apply.sh --dry-run exits 0 and creates no state directory; read-only prove.py reports FAIL before the apply, as expected (anchored modifier not enabled)
Static checks local integration shellcheck 0.11.0 is clean on apply.sh, rollback.sh and prove.sh, and bash -n passes

Reviews (bounded to two rounds; the notes are in the session scratchpad):

  • Round 1, adversarial review by a Claude Opus reviewer (p2rev/review-u1.md) found 12 defects
    and a nit, none blocking the design. All were fixed, plus 3 more found during the repair: the
    anchored modifier's position, the Codex active window and partial rollback. Failing-first:
    38 tests gave 22 failures and 3 errors before, and 38 of 38 passed after. The resolution is
    u1-p2/REPAIR-1.md, commit 6ad0d358.
  • Round 2, GPT-6 cross-family review (gpt-6-astra, effort max, read-only) found 1 high and
    6 medium issues. All seven held against the source, and all are fixed with tests that failed
    first (9 failures and 1 error before). The resolution is u1-p2/REPAIR-2.md, commit dfd5460f.
    • high: prove.py compares only processes that completed a turn;
    • concurrency needs overlapping activity spans;
    • the self-metrics check reads raw up samples with the scrape interval;
    • every --apply reads the running services back;
    • Prometheus is read at the rendered unit's address;
    • rollback guards the settings key;
    • the unscoped-writer alert keeps 15 minutes of samples.

Host-apply plan (not run yet)

The coordinator applies this in one window with the stacked invoke-rate change. $REPO is a
checkout holding both. The step order is in that PR's plan: install its Collector file without a
restart, then run this recipe once.

  1. evidence/artifacts/telemetry-writer-identity-20260926/host/apply.sh --dry-run --repo "$REPO" --only prometheus,dashboards,collector,claude-setting
    is read-only: it renders, validates (otelcol, promtool, systemd-analyze) and diffs.
  2. The same command with --apply backs up to
    ${XDG_STATE_HOME:-~/.local/state}/native-agent-stack/g1-writer-identity/backup-<UTC>/
    (0700, with manifest.json). It installs, restarts the Collector and Prometheus once, and reads
    both back. Any failed read-back stops it and names the rollback.
  3. prove.sh runs after a scenario window with Claude and concurrent Codex activity.
  4. Rollback: rollback.sh --apply is all or nothing and refuses files or the settings key if they
    changed after the apply, unless --force.

The codex-launcher step is deliberately left out of --only.

Residuals

  • Codex identity launcher: out of scope for this apply. An upstream check found no native Codex
    per-process OTel identity. At rust-v0.157.1, Codex builds its resource from the service name,
    version, env and OS, plus OTEL_RESOURCE_ATTRIBUTES, and exports no per-process attribute. The
    template stays in the tree as the documented local-integration path. Until it is installed, the
    host sees the following:
    • Codex token series stay instance="unscoped", except for callers that set their own id.
      Panels 1 and 26 ("Codex turn tokens") exclude them, so Codex tokens show only in Loki and the
      native-data panels.
    • EcosystemUnscopedTokenWriters fires while Codex exports counters.
    • prove.py fails setup.codex_launcher, so its PASS is unreachable as written. The coordinator
      decides whether to accept that failure or amend the proof.
  • Host apply and host acceptance have not run.
  • prove.py defaults to --prom :19090 and --loki :13100.
  • Activity spans coarsen past a 10,000 s window.
  • lastConfigTime has whole-second resolution.
  • The prove fake and the apply sandbox are synthetic test doubles.
  • The GPT-6 round was the last one; no third review ran.
  • Merge order: this branch and the token-savings branch both re-register manifests/evidence.json.
    Whichever merges second takes main's copy and re-registers its own files (hot-file protocol in
    docs/lanes.md).
  • Side findings, recorded by the stacked change:
    • otelcol-contrib 0.161.0 logs that the otlphttp alias is deprecated;
    • merge_collector.py replaces the whole metrics processor list and metric_statements, so a
      host-specific metrics processor would be dropped without a refusal.

SOTA sources

  • Claude Code monitoring docs, https://code.claude.com/docs/en/monitoring-usage (fetched
    2026-09-26T15:01Z): OTEL_METRICS_INCLUDE_SESSION_ID default true, session.id on metrics,
    cumulative temporality.
  • open-telemetry/opentelemetry-collector-contrib v0.161.0:
    • processor/deltatocumulativeprocessor/README.md (ErrOutOfOrder, ErrOlderStart);
    • processor/groupbyattrsprocessor/README.md;
    • processor/transformprocessor/README.md (aggregate_on_attributes);
    • exporter/prometheusexporter/README.md (metric_expiration 5m, start timestamps).
  • prometheus/prometheus v3.15.0:
    • docs/feature_flags.md (created-timestamp-zero-ingestion, promql-extended-range-selectors
      and anchored);
    • docs/configuration/configuration.md (metric_relabel_configs).
  • openai/codex rust-v0.157.1 (36650394c5b3): codex-rs/otel/src/metrics/client.rs and
    codex-rs/otel/src/provider.rs build the resource with no per-process attribute;
    OTEL_RESOURCE_ATTRIBUTES reaches it through EnvResourceDetector (opentelemetry_sdk 0.31.0).

Decision record

docs/decisions/2026-09-26-telemetry-writer-identity.md covers the alternatives, the overturn
conditions and the limits.

Window 2: host apply, proof, GPT-6 review and repair round 3 (coordinator, 2026-09-26/27)

  • Replay: onto main 34c56340 as 924ec302, under the hot-file protocol (main's manifests/evidence.json, 36 files re-registered). The native-data files from Native data (G2): token-tool savings into Loki and Grafana, one series per native scope #365 and Native-data snapshot: read token reports up to their documented embedded size #373 auto-merged and were checked. validate.py FAILS=0; verdict gate reports no rows changed; the full suite is OK (skipped=642).
  • Host apply, window 2: steps 0-9, applied and read back with no failed gate:
    • U6 filter, then u1 metrics side (one otelcol and one Prometheus restart);
    • OTEL_LOG_TOOL_DETAILS=1;
    • the Codex identity launcher, as its own step.
  • Proof, from fresh headless processes:
    • canary against production Loki: PASS;
    • u6 prove.sh: 33/33;
    • u1 prove.sh: exit 1 on the two permitted checks only (unscoped_token_writers and firing_integrity_alerts), both job=claude-code and attributed by process capture to the pre-change coordinator. Its 60 frozen unscoped series live until that process exits. Claude writers and Codex processes reconcile at 0.00%, and no Codex series is unscoped.
  • GPT-6 review of window 2 (gpt-6-astra, max effort, live loopback access): privacy clean across 16,632 live records (0 banned keys, 0 emails, 0 canary hits); every installed logs pipeline guarded; the launcher is authentic and argv intact. Resolved in f4f772cb, with failing-first tests:
    • (medium) the launcher overwrote inherited exported scratch names. They now use the __codex_identity_ prefix and are assigned in a subshell; the child environment and exact argv are tested.
    • (high, open since window 1) the collector merge could drop a statement. merge_collector.py now verifies statements, groups, settings and pipeline order, and refuses before writing a candidate.
    • (medium) stale-candidate race. apply.sh records render-time hashes and re-checks before and at each install.
      94 targeted tests OK; validate FAILS=0. The anti-pattern log records the three lessons.
  • Residuals:
    • Advisor (Fable) usage appears in the Claude token metric with no api_request event, so a writer that uses the advisor can fail the per-writer comparison. Recorded as a comparison and coverage limitation.
    • The scenario runner's rc masking and the proof checker's short-string filter live in scratch helpers; they move with the host-helper commit (PR-EV).
    • tests.test_landscape_sweep_harness fails intermittently on the workstation under concurrent load, on main as well; hosted CI decides.

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 26, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: dfd5460fb2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread evidence/artifacts/telemetry-writer-identity-20260926/host/merge_prometheus.py Outdated
Comment thread evidence/artifacts/telemetry-writer-identity-20260926/host/apply.sh
Scout and others added 2 commits September 26, 2026 20:46
…y alerts and a committed host recipe

One writer per Prometheus token series: Claude session ids become the
Collector's service.instance.id, a Codex identity launcher gives each process
its own id (an inherited id becomes the prefix), dropped attributes are summed,
and Prometheus runs with start-timestamp zero injection and anchored ranges.
The collector-native scrape drops the Codex histogram buckets other than token
usage, and the native-telemetry-integrity rules watch per-series resets,
dropped delta points and unscoped writers.

Review repair 1 (adversarial review of unit u1): the host recipe (apply.sh,
rollback.sh, prove.sh) is committed under
evidence/artifacts/telemetry-writer-identity-20260926/host with failing-first
tests. Validations and read-backs now stop apply.sh; backups live in a private
XDG state directory; rollback is all or nothing. The proof compares every
finished Codex process and each Claude writer away from the window edges,
requires the Collector self-metrics, allows one reset per series, writes the
anchored modifier where Prometheus 3.15 parses it and reports capacity
against the retention limit.

Repair 2 (GPT-6 cross-family review of repair 1; all seven findings held, 1
high and 6 medium, each fixed with a test that failed first):
- prove.py (high): a finished codex_exec process is compared only when it
  completed a turn (a non-prewarm response.completed in Loki, or turn tokens in
  Prometheus). Start-only processes are listed, not compared, so an unexercised
  window stays INCONCLUSIVE.
- prove.py: concurrency needs overlapping activity spans, not events in one
  minute. A span runs from the first to the last Loki event, at 1 s steps up to
  a 10,000 s window (Loki 3.7.8 refuses more than 11,000 points per series).
- prove.py: the collector self-metrics check reads the raw up samples of the
  window, with the scrape interval from /api/v1/targets. Every sample must be 1,
  and no gap may exceed 1.5 intervals, edges included.
- apply.sh: every --apply reads the running Prometheus and Collector back, also
  when their files are already in place. A service that started, or loaded its
  config, before the installed files is restarted, as is one that lacks a
  feature, rule group, bucket drop or pipeline. If it still differs, the script
  stops with the rollback command. A failed restart now names the rollback too.
- apply.sh and rollback.sh read Prometheus at the --web.listen-address of the
  rendered or restored unit, not a fixed 19090.
- rollback.sh refuses, without --force, when the Claude settings key no longer
  holds the value apply wrote. apply now records that value.
- EcosystemUnscopedTokenWriters keeps 15 minutes of unscoped samples
  (last_over_time) and fires after 5 minutes. A short Codex run, whose series
  the exporter drops after 5 minutes, now fires it.

Evidence (branch record): 116 unit and related tests passed, including the
promtool 3.15.0 and otelcol-contrib 0.161.0 cases. All 20 PromQL queries parse
on the pinned Prometheus, and the loopback Loki accepts all 16 LogQL queries.
Read-only host probes are recorded in the before-change receipt
(backend_api_observation). Host apply and host acceptance have still not run.

manifests/evidence.json: changed files re-registered on current main;
component_matrix and new_host_grand_list re-run with --write.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rge preservation, stale-candidate refusal

- F1 (medium): the Codex identity launcher assigned bare scratch variables,
  so a parent that exported id, key, value and the like handed Codex changed
  values. Scratch names now carry the __codex_identity_ prefix and are
  assigned inside a command substitution. Only OTEL_RESOURCE_ATTRIBUTES
  changes in the child, and argv is forwarded exactly (POSIX 2.5.2-2.5.3;
  opentelemetry-rust v0.31.0 resource/env.rs).
- F2 (high, still open from window 1): merge_collector.py verifies its
  result before returning it. Original statement order and multiplicity,
  group metadata, processor settings and every pipeline must be preserved,
  and unsupported host customisations are refused before a candidate is
  written (Collector v0.161.0 pipeline and transform ordering).
- F3 (medium): apply.sh records each target's sha256 at render time. It
  re-checks every selected target after validation and before any install,
  and again at each install, refusing on drift. Backups and rollback are
  unchanged.
Failing-first tests cover each case: the launcher environment and exact
argv, merge fixtures that drop statements or reorder processors, and target
changes during validation and after preflight. The anti-pattern log records
the three lessons.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/monitor-metrics-integrity-20260926 branch from 8172c17 to f4f772c Compare September 27, 2026 01:24
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 verification of repair round 3 (924ec302..f4f772cb): gpt-6-astra, max effort, read-only, live search, rc=0. Verdict: NO DEFECTS. All 58 targeted tests pass with no skips; the failing-first claims check out against the old code; full repository validation passes.

…3.2 runs it

validate-macos failed 5 launcher tests: bash 3.2, which is macOS bash and /bin/sh,
mis-parses a case statement and nested quoted expansions inside $(...)
("syntax error near unexpected token 'newline'"). The attribute parsing now
lives in __codex_identity_attributes, a top-level function that bash 3.2
parses normally, called as $(...), so its scratch assignments still stay
out of codex's environment. Reproduced and verified with bash 3.2.57 as
both bash and sh: 16 launcher tests OK; with bash 5, 58 targeted tests OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

validate-macos caught a real defect in repair round 3 (f4f772cb): macOS bash 3.2, which is also /bin/sh there, mis-parses the case statement and nested quoted expansions that the repair placed inside $(...), so 5 launcher tests failed. 8e661348 moves the parsing into a top-level function called as $(fn). bash 3.2 parses the function body normally, and the subshell still keeps its scratch assignments out of codex's environment. Reproduced and verified locally with bash 3.2.57 as both bash and sh: 16 launcher tests OK. With bash 5, 58 targeted tests OK; validate.py FAILS=0.

…reflight Claude settings

T2: merge_prometheus.py replaced the host's collector-native
metric_relabel_configs with the bucket rules, and the guard stripped the
whole key before comparing, so the loss passed. The merge now keeps the host
rules first in their order and appends only missing owned rules (a second
merge is a no-op); verify_merge checks the preserved prefix, the exact
additions and every other host setting. Prometheus applies relabel rules in
configured order (model/relabel/relabel.go at v3.15.0).

T3: with the claude-setting step selected, an absent, malformed, symlinked,
non-object or directory ~/.claude/settings.json failed only after earlier
steps had installed and restarted services. apply.sh now preflights it
through claude_setting.py's read path before rendering or installing.

T1 (extra host metrics processor or transform/privacy statement) was already
refused by round 3's verify step; its tests are named in the thread reply.

Nine failing-first tests (red logs retained); tests.test_observability_writer_identity_host
and tests.test_observability_writer_identity: 67 OK under bash 5 and under
bash 3.2.57 as both bash and sh.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins merged commit c71d66d into main Sep 27, 2026
24 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/monitor-metrics-integrity-20260926 branch September 27, 2026 02:45
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…emory, context-hub, agentsview and an RTK exactness fix

Replayed onto main c71d66d (after #364) through the hot-file protocol: the
five PR files are unchanged; manifests/evidence.json is re-registered and the
generated matrix and grand list are regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…emory, context-hub, agentsview and an RTK exactness fix (#375)

Replayed onto main c71d66d (after #364) through the hot-file protocol: the
five PR files are unchanged; manifests/evidence.json is re-registered and the
generated matrix and grand list are regenerated.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
Replayed onto main 881164e (after #364 and #375) with a 3-way squash: main had
changed adoption/bootstrap.md, which merged cleanly; manifests/evidence.json
was re-registered and the matrix and grand list regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…e and practice docs (fix plan steps 7-8)

Replayed onto main 881164e (after #364 and #375) with a 3-way squash. Main had
changed three of this PR's files:
- The settings template and the macOS page merged cleanly; the template keeps
  #364's OTEL_METRICS_INCLUDE_SESSION_ID and this PR's 11 Read(**/...) twins.
- adoption/bootstrap.md conflicted on one note: it keeps this PR's extended
  settings-template sentence, followed by #364's writer-identity paragraph.
manifests/evidence.json was re-registered, and the matrix and grand list were
regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…re-register U6 on main

Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that
nested quoted expansions inside $(...) parse correctly. What breaks is a case
pattern's closing parenthesis, which ends the substitution early. bash -n
passes; at run time the shell prints a syntax error and continues with a wrong
value.

U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json
re-registered and the matrix and grand list regenerated. validate FAILS=0.
438 observability and manifest tests OK; the writer-identity modules also pass
under bash 3.2.57 as bash and sh (67 OK).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
Replayed onto main 881164e (after #364 and #375) with a 3-way squash: main had
changed adoption/bootstrap.md, which merged cleanly; manifests/evidence.json
was re-registered and the matrix and grand list regenerated.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…e and practice docs (fix plan steps 7-8)

Replayed onto main 881164e (after #364 and #375) with a 3-way squash. Main had
changed three of this PR's files:
- The settings template and the macOS page merged cleanly; the template keeps
  #364's OTEL_METRICS_INCLUDE_SESSION_ID and this PR's 11 Read(**/...) twins.
- adoption/bootstrap.md conflicted on one note: it keeps this PR's extended
  settings-template sentence, followed by #364's writer-identity paragraph.
manifests/evidence.json was re-registered, and the matrix and grand list were
regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…re-register U6 on main

Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that
nested quoted expansions inside $(...) parse correctly. What breaks is a case
pattern's closing parenthesis, which ends the substitution early. bash -n
passes; at run time the shell prints a syntax error and continues with a wrong
value.

U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json
re-registered and the matrix and grand list regenerated. validate FAILS=0.
438 observability and manifest tests OK; the writer-identity modules also pass
under bash 3.2.57 as bash and sh (67 OK).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…e and practice docs (fix plan steps 7-8) (#374)

Replayed onto main 881164e (after #364 and #375) with a 3-way squash. Main had
changed three of this PR's files:
- The settings template and the macOS page merged cleanly; the template keeps
  #364's OTEL_METRICS_INCLUDE_SESSION_ID and this PR's 11 Read(**/...) twins.
- adoption/bootstrap.md conflicted on one note: it keeps this PR's extended
  settings-template sentence, followed by #364's writer-identity paragraph.
manifests/evidence.json was re-registered, and the matrix and grand list were
regenerated.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…re-register U6 on main

Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that
nested quoted expansions inside $(...) parse correctly. What breaks is a case
pattern's closing parenthesis, which ends the substitution early. bash -n
passes; at run time the shell prints a syntax error and continues with a wrong
value.

U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json
re-registered and the matrix and grand list regenerated. validate FAILS=0.
438 observability and manifest tests OK; the writer-identity modules also pass
under bash 3.2.57 as bash and sh (67 OK).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
… and Grafana (stacked on #364) (#366)

* Tool invoke rates: tool, MCP server, skill and subagent names in Loki and Grafana, with the dated tool-details exception

Stacked on the telemetry writer identity change (dfd5460).

- Collector: transform/tool_names runs in both logs pipelines before
  transform/privacy. From Claude Code's tool_parameters it copies only the
  MCP server and tool names and the Agent tool's subagent_type, plus one
  boolean, shell_rtk. Skill names come from skill_activated. Registry names
  that are not short identifiers become "other", and agent types outside the
  documented built-ins become "custom". tool_parameters, tool_input,
  arguments, output, content and error are deleted, and the log allowlist
  gains 19 name, id and derived keys.
- Dashboard row 28 (panels 29-38): calls by family, MCP server, skill,
  client and actor; subagent and workflow launches; MCP share; ctx and rtk
  adoption; and an integrity panel.
- OTEL_LOG_TOOL_DETAILS is "1" in both Claude settings examples.
  docs/secret-storage.md records it as a dated exception. On 2026-09-26, in
  the workstation coordinator session, the user answered "1 and full sota
  convergence practice we proceed". The exception states the control and the
  canary proof, and marks the template changed after `v2026.09.26.2`. The
  decision record's status and the new-host template notes in bootstrap.md
  and macos-arm64.md say the same.
- tests/test_observability_tool_names.py: its two fixture thread ids are now
  built at runtime with uuid.UUID(int=...), with the same values.
  scripts/validate.py rejected the UUID literals as possible local session
  ids.

Evidence (local integration and synthetic; the host is not applied):
otelcol-contrib 0.161.0 validate exits 0. The replay of real Claude Code and
Codex captures plus synthetic records passed 109 of 109 checks: 72 canary
strings, 0 found in the file exporter or Loki; 19 of 19 dashboard targets.
The unit test passes 11 of 11 and fails on dfd5460 alone. validate.py,
evidence_manifest --check and the related modules pass. Reviews: two GPT-6
cross-family rounds; every high and medium finding is fixed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* U6: commit the sanitized replay and live-proof receipt, fix the privacy checker, correct client grouping

Resolves the review threads on #366 and the window-2 GPT-6 findings 2 and 5.
- evidence/artifacts/tool-invoke-rates-20260926/:
  - the replay harness, its reference implementation and the sanitized
    109/109 result (otelcol-contrib 0.161.0, Loki 3.7.8);
  - the live host proof (2026-09-26T23:47:42Z-23:49:21Z, 33 passed), with
    counts and hashes only for real captures;
  - the fixed checker.
- prove_check.py and prove.sh: no length floor on forbidden strings, and every
  tagged record's body must equal the Collector placeholder exactly.
  - Synthetic A/B: the old checker passed a leaked "pwd" body; the fixed one
    fails it.
  - Offline re-check of the retained events-file sink of the 23:47Z proof:
    29 tagged lines, all "[content omitted]", 0 forbidden hits. Its Loki
    sink was not re-checked by the fixed checker, and the decision record
    says so.
- client is documented as a grouping by normalized originator name: separate
  app-server processes with the same originator share one label.
- Panel 34: one description sentence. LogQL rate over unwrap is sum/range
  (Loki v3.7.8 pkg/logql/range_vector.go rateLogs), so the query is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Launcher comment: only a case pattern inside $(...) breaks bash 3.2; re-register U6 on main

Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that
nested quoted expansions inside $(...) parse correctly. What breaks is a case
pattern's closing parenthesis, which ends the substitution early. bash -n
passes; at run time the shell prints a syntax error and continues with a wrong
value.

U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json
re-registered and the matrix and grand list regenerated. validate FAILS=0.
438 observability and manifest tests OK; the writer-identity modules also pass
under bash 3.2.57 as bash and sh (67 OK).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* U6 receipt repair (GPT-6 review of #366): check every OTLP record structurally; synthetic-only replay runs natively

Findings from the cross-family review of 5ddfd93, each fixed with failing-first controls:
- high: prove_check.py kept only the last body of each OTLP batch, and missing or
  non-string bodies passed. scan_otlp now keeps and validates every record's body.
- medium: resource, scope and nested (ArrayValue/KeyValueList) values were not
  scanned, and banned keys matched only minified JSON. The scan now traverses
  decoded keys and values at every OTLP level (opentelemetry-proto v1.9.0 logs.proto
  and common.proto). The live and offline paths share one events_privacy_checks().
  Structural controls: 2 passed, 10 failed before; 12 passed after.
- medium: the synthetic-only replay failed on max() of an empty capture and required
  the omitted private captures. It now has a default timestamp and runs the
  historical assertions only when the full capture set is present.
  - Stubbed control: 25/20 before, 25/0 after.
  - Native run by the coordinator on the host, on the pinned Collector 0.161.0 and
    Loki 3.7.8 with scratch ports 45700-45703: 67 passed, 0 failed.
- low: docs/secret-storage.md still said the live proof had not run.

The corrected offline re-check of the 23:47Z proof's retained events file ran at
03:55:26Z: 270 records in 29 batches, 3 passed. The exporter (10 MB x 3 backups)
rotated that file out at 04:05:11Z, and a re-run at 04:14Z found 0 tagged lines. The
receipt records that the result cannot be repeated. The three observability test
modules pass on the host (78 OK), where the sandbox had blocked 21 socket tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* U6 receipt, second repair (GPT-6 re-check of 7a15a2e): decode bytesValue; always check Loki bodies

- A ProtoJSON bytesValue (base64, standard or URL-safe, with or without padding)
  was scanned only as encoded text, so bytesValue "cHdk" (pwd) passed. scan_otlp
  now scans the decoded text too. An undecodable value adds a marker that the
  checks treat as a banned key, so they fail closed. Controls: 12 passed, 3 failed
  before (bytes attribute, nested bytes, undecodable); 15 passed, 0 failed after.
- Loki's fixed-body privacy assertion ran only with the historical capture set.
  replay.py loki_body_checks() now runs on every replay: every stored Loki line
  must equal the Collector placeholder. loki_body_control.py runs the real loki()
  against a stubbed Loki: on the prior replay.py no body check ran and 31 "pwd"
  lines went unnoticed (0/2); after the fix, clean records pass and the leak is
  caught (2/2).
- The native synthetic replay on the host (Collector 0.161.0, Loki 3.7.8, scratch
  ports) now reports 68 passed, 0 failed.
- Residual: Loki structured metadata keeps attribute values as strings, so a
  bytes-typed attribute stays base64 there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…asks,

metrics, gate, procedure)

Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the
baseline-receipt half already merged via #369) before any organic run:
77 frozen task definitions (16 reused Claude + 15 reused Codex information
needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3
metric/threshold/guardrail/outcome rule with its three "proposed" bounds now
stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the
AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process
qualification, the Workflow script that will run arms B/A/A0 (syntax-checked
only, never executed), and the measurement runbook.

Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6
cross-family review (2 findings, 1 high) plus one combined repair round (13/13
addressed with failing-first evidence): relocated from
blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/
per adoption-audit/PLAN.md line 402; moved the Workflow script out of
examples/claude-native/workflows/ to avoid that directory's shared
PACKET/ROUTING contract test, which this one-off frozen script correctly does
not match; corrected PR-A's status (its base branch merged as #369 -- baseline
only, tooling-fix scope still open); added the six AA §6 machine-readable
policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle
binding; blocked the one information need (an otel-tui receiver task) that no
current role may legally run, pending a named permitted role; froze the
coordinator's own main-dispatch task at xhigh instead of max; fixed a
denylist word-boundary gap that let underscore-qualified MCP names slip past;
and sealed the five normative files' SHA256 with an append-only amendment
rule. Also refreshes PR #376 and #364's status to merged (both landed while
this unit was in review), while recording that #364's field-preservation gap
(workflow.run_id/tool_use_id still missing from collector.yaml) persists after
its merge.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…asks, (#381)

metrics, gate, procedure)

Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the
baseline-receipt half already merged via #369) before any organic run:
77 frozen task definitions (16 reused Claude + 15 reused Codex information
needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3
metric/threshold/guardrail/outcome rule with its three "proposed" bounds now
stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the
AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process
qualification, the Workflow script that will run arms B/A/A0 (syntax-checked
only, never executed), and the measurement runbook.

Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6
cross-family review (2 findings, 1 high) plus one combined repair round (13/13
addressed with failing-first evidence): relocated from
blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/
per adoption-audit/PLAN.md line 402; moved the Workflow script out of
examples/claude-native/workflows/ to avoid that directory's shared
PACKET/ROUTING contract test, which this one-off frozen script correctly does
not match; corrected PR-A's status (its base branch merged as #369 -- baseline
only, tooling-fix scope still open); added the six AA §6 machine-readable
policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle
binding; blocked the one information need (an otel-tui receiver task) that no
current role may legally run, pending a named permitted role; froze the
coordinator's own main-dispatch task at xhigh instead of max; fixed a
denylist word-boundary gap that let underscore-qualified MCP names slip past;
and sealed the five normative files' SHA256 with an append-only amendment
rule. Also refreshes PR #376 and #364's status to merged (both landed while
this unit was in review), while recording that #364's field-preservation gap
(workflow.run_id/tool_use_id still missing from collector.yaml) persists after
its merge.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
Correct the Claude custom-attribute citation and resource/event/datapoint
description, whole-run prefix queries, deployment path and dated E2E gates.
Publish original and repair command/result excerpts directly in the PR body.

Explicitly delete Codex arguments with native OTTL and preserve tool namespace,
guarded MCP server, timing, success, correlation and supplied size metadata.
Test duplicate Claude run attributes and Codex records with and without server
metadata through installed Collector 0.161.0 and Loki 3.7.8.

Sources: open-telemetry/opentelemetry-collector-contrib v0.161.0,
processor/transformprocessor/README.md and pkg/ottl/ottlfuncs/README.md;
openai/codex rust-v0.157.1, codex-rs/otel/src/tool_result.rs:55-79,94,
codex-rs/otel/src/events/shared.rs:59-60 and
codex-rs/core/src/tools/call_trace.rs:38-54;
https://code.claude.com/docs/en/monitoring-usage#multi-team-organization-support;
https://grafana.com/docs/loki/latest/query/log_queries/ and send-data/otel/;
open-telemetry/opentelemetry-proto v1.9.0 docs/specification.md;
base ec27a30, #364 c71d66d and #366 a464d28 historical receipts.

Verification: failing-first C1 checks, then 83 covering unittests pass with no
skips; complete native Collector validate and git diff --check pass. Evidence
is local integration with synthetic fixtures, not live-provider/host acceptance.
The coordinator owns receipt hash registration and commit creation.

Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute
pool; committed by the coordinator harness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…dex tool arguments, whole-run query guidance (#408)

* Preserve Claude tool joins and estimated cost through Loki

Retain tool_use_id and cost_usd alongside the existing workflow.run_id,
session.id and ecosystem.task.id resource route. Keep the current structured
metadata policy and metric writer identity. Add red-before-green unittest
coverage through native Collector 0.161.0 and Loki 3.7.8, including per-arm
usage, resource-only runs, private tool joins and indexed-series checks.

Sources: Anthropic monitoring documentation, checked 2026-09-27:
https://code.claude.com/docs/en/monitoring-usage
anthropics/claude-code v2.1.283 release notes and installed client.
open-telemetry/opentelemetry-collector-contrib v0.161.0:
processor/transformprocessor/README.md and pkg/ottl/ottlfuncs/func_keep_keys.go.
grafana/loki v3.7.8 with https://grafana.com/docs/loki/latest/send-data/otel/
and https://grafana.com/docs/loki/latest/reference/loki-http-api/.
open-telemetry/opentelemetry-proto v1.9.0, docs/specification.md.

Validation: 81 covering unittest checks passed without skips; the complete
Collector profile passed native validation after creating documented scratch
directories. Local integration uses synthetic inputs, not a new model run.

Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute
pool; committed by the coordinator harness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Repair collector evidence and Codex tool metadata privacy

Correct the Claude custom-attribute citation and resource/event/datapoint
description, whole-run prefix queries, deployment path and dated E2E gates.
Publish original and repair command/result excerpts directly in the PR body.

Explicitly delete Codex arguments with native OTTL and preserve tool namespace,
guarded MCP server, timing, success, correlation and supplied size metadata.
Test duplicate Claude run attributes and Codex records with and without server
metadata through installed Collector 0.161.0 and Loki 3.7.8.

Sources: open-telemetry/opentelemetry-collector-contrib v0.161.0,
processor/transformprocessor/README.md and pkg/ottl/ottlfuncs/README.md;
openai/codex rust-v0.157.1, codex-rs/otel/src/tool_result.rs:55-79,94,
codex-rs/otel/src/events/shared.rs:59-60 and
codex-rs/core/src/tools/call_trace.rs:38-54;
https://code.claude.com/docs/en/monitoring-usage#multi-team-organization-support;
https://grafana.com/docs/loki/latest/query/log_queries/ and send-data/otel/;
open-telemetry/opentelemetry-proto v1.9.0 docs/specification.md;
base ec27a30, #364 c71d66d and #366 a464d28 historical receipts.

Verification: failing-first C1 checks, then 83 covering unittests pass with no
skips; complete native Collector validate and git diff --check pass. Evidence
is local integration with synthetic fixtures, not live-provider/host acceptance.
The coordinator owns receipt hash registration and commit creation.

Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute
pool; committed by the coordinator harness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Collector README: state that the sealed #381 receipts are unchanged; drop an unsupported #366 citation

Coordinator correction after the independent repair verification: the repair's receipt edits were withdrawn (RUNBOOK.md is hash-sealed; the E2E README changes only by amendment), so the README must not say they carry dated updates; the #366 receipt does not record a byte-identical host deployment.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Hot files: evidence registration and generated reports

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant