Repository navigation
Telemetry writer identity (G1): one writer per token series, integrity alerts and a committed host recipe - #364
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dfd5460fb2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
dfd5460 to
8172c17
Compare
…y alerts and a committed host recipe One writer per Prometheus token series: Claude session ids become the Collector's service.instance.id, a Codex identity launcher gives each process its own id (an inherited id becomes the prefix), dropped attributes are summed, and Prometheus runs with start-timestamp zero injection and anchored ranges. The collector-native scrape drops the Codex histogram buckets other than token usage, and the native-telemetry-integrity rules watch per-series resets, dropped delta points and unscoped writers. Review repair 1 (adversarial review of unit u1): the host recipe (apply.sh, rollback.sh, prove.sh) is committed under evidence/artifacts/telemetry-writer-identity-20260926/host with failing-first tests. Validations and read-backs now stop apply.sh; backups live in a private XDG state directory; rollback is all or nothing. The proof compares every finished Codex process and each Claude writer away from the window edges, requires the Collector self-metrics, allows one reset per series, writes the anchored modifier where Prometheus 3.15 parses it and reports capacity against the retention limit. Repair 2 (GPT-6 cross-family review of repair 1; all seven findings held, 1 high and 6 medium, each fixed with a test that failed first): - prove.py (high): a finished codex_exec process is compared only when it completed a turn (a non-prewarm response.completed in Loki, or turn tokens in Prometheus). Start-only processes are listed, not compared, so an unexercised window stays INCONCLUSIVE. - prove.py: concurrency needs overlapping activity spans, not events in one minute. A span runs from the first to the last Loki event, at 1 s steps up to a 10,000 s window (Loki 3.7.8 refuses more than 11,000 points per series). - prove.py: the collector self-metrics check reads the raw up samples of the window, with the scrape interval from /api/v1/targets. Every sample must be 1, and no gap may exceed 1.5 intervals, edges included. - apply.sh: every --apply reads the running Prometheus and Collector back, also when their files are already in place. A service that started, or loaded its config, before the installed files is restarted, as is one that lacks a feature, rule group, bucket drop or pipeline. If it still differs, the script stops with the rollback command. A failed restart now names the rollback too. - apply.sh and rollback.sh read Prometheus at the --web.listen-address of the rendered or restored unit, not a fixed 19090. - rollback.sh refuses, without --force, when the Claude settings key no longer holds the value apply wrote. apply now records that value. - EcosystemUnscopedTokenWriters keeps 15 minutes of unscoped samples (last_over_time) and fires after 5 minutes. A short Codex run, whose series the exporter drops after 5 minutes, now fires it. Evidence (branch record): 116 unit and related tests passed, including the promtool 3.15.0 and otelcol-contrib 0.161.0 cases. All 20 PromQL queries parse on the pinned Prometheus, and the loopback Loki accepts all 16 LogQL queries. Read-only host probes are recorded in the before-change receipt (backend_api_observation). Host apply and host acceptance have still not run. manifests/evidence.json: changed files re-registered on current main; component_matrix and new_host_grand_list re-run with --write. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rge preservation, stale-candidate refusal - F1 (medium): the Codex identity launcher assigned bare scratch variables, so a parent that exported id, key, value and the like handed Codex changed values. Scratch names now carry the __codex_identity_ prefix and are assigned inside a command substitution. Only OTEL_RESOURCE_ATTRIBUTES changes in the child, and argv is forwarded exactly (POSIX 2.5.2-2.5.3; opentelemetry-rust v0.31.0 resource/env.rs). - F2 (high, still open from window 1): merge_collector.py verifies its result before returning it. Original statement order and multiplicity, group metadata, processor settings and every pipeline must be preserved, and unsupported host customisations are refused before a candidate is written (Collector v0.161.0 pipeline and transform ordering). - F3 (medium): apply.sh records each target's sha256 at render time. It re-checks every selected target after validation and before any install, and again at each install, refusing on drift. Backups and rollback are unchanged. Failing-first tests cover each case: the launcher environment and exact argv, merge fixtures that drop statements or reorder processors, and target changes during validation and after preflight. The anti-pattern log records the three lessons. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
8172c17 to
f4f772c
Compare
|
GPT-6 verification of repair round 3 ( |
…3.2 runs it
validate-macos failed 5 launcher tests: bash 3.2, which is macOS bash and /bin/sh,
mis-parses a case statement and nested quoted expansions inside $(...)
("syntax error near unexpected token 'newline'"). The attribute parsing now
lives in __codex_identity_attributes, a top-level function that bash 3.2
parses normally, called as $(...), so its scratch assignments still stay
out of codex's environment. Reproduced and verified with bash 3.2.57 as
both bash and sh: 16 launcher tests OK; with bash 5, 58 targeted tests OK.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
validate-macos caught a real defect in repair round 3 ( |
…reflight Claude settings T2: merge_prometheus.py replaced the host's collector-native metric_relabel_configs with the bucket rules, and the guard stripped the whole key before comparing, so the loss passed. The merge now keeps the host rules first in their order and appends only missing owned rules (a second merge is a no-op); verify_merge checks the preserved prefix, the exact additions and every other host setting. Prometheus applies relabel rules in configured order (model/relabel/relabel.go at v3.15.0). T3: with the claude-setting step selected, an absent, malformed, symlinked, non-object or directory ~/.claude/settings.json failed only after earlier steps had installed and restarted services. apply.sh now preflights it through claude_setting.py's read path before rendering or installing. T1 (extra host metrics processor or transform/privacy statement) was already refused by round 3's verify step; its tests are named in the thread reply. Nine failing-first tests (red logs retained); tests.test_observability_writer_identity_host and tests.test_observability_writer_identity: 67 OK under bash 5 and under bash 3.2.57 as both bash and sh. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…emory, context-hub, agentsview and an RTK exactness fix Replayed onto main c71d66d (after #364) through the hot-file protocol: the five PR files are unchanged; manifests/evidence.json is re-registered and the generated matrix and grand list are regenerated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…emory, context-hub, agentsview and an RTK exactness fix (#375) Replayed onto main c71d66d (after #364) through the hot-file protocol: the five PR files are unchanged; manifests/evidence.json is re-registered and the generated matrix and grand list are regenerated. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e and practice docs (fix plan steps 7-8) Replayed onto main 881164e (after #364 and #375) with a 3-way squash. Main had changed three of this PR's files: - The settings template and the macOS page merged cleanly; the template keeps #364's OTEL_METRICS_INCLUDE_SESSION_ID and this PR's 11 Read(**/...) twins. - adoption/bootstrap.md conflicted on one note: it keeps this PR's extended settings-template sentence, followed by #364's writer-identity paragraph. manifests/evidence.json was re-registered, and the matrix and grand list were regenerated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re-register U6 on main Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that nested quoted expansions inside $(...) parse correctly. What breaks is a case pattern's closing parenthesis, which ends the substitution early. bash -n passes; at run time the shell prints a syntax error and continues with a wrong value. U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json re-registered and the matrix and grand list regenerated. validate FAILS=0. 438 observability and manifest tests OK; the writer-identity modules also pass under bash 3.2.57 as bash and sh (67 OK). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replayed onto main 881164e (after #364 and #375) with a 3-way squash: main had changed adoption/bootstrap.md, which merged cleanly; manifests/evidence.json was re-registered and the matrix and grand list regenerated. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e and practice docs (fix plan steps 7-8) Replayed onto main 881164e (after #364 and #375) with a 3-way squash. Main had changed three of this PR's files: - The settings template and the macOS page merged cleanly; the template keeps #364's OTEL_METRICS_INCLUDE_SESSION_ID and this PR's 11 Read(**/...) twins. - adoption/bootstrap.md conflicted on one note: it keeps this PR's extended settings-template sentence, followed by #364's writer-identity paragraph. manifests/evidence.json was re-registered, and the matrix and grand list were regenerated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re-register U6 on main Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that nested quoted expansions inside $(...) parse correctly. What breaks is a case pattern's closing parenthesis, which ends the substitution early. bash -n passes; at run time the shell prints a syntax error and continues with a wrong value. U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json re-registered and the matrix and grand list regenerated. validate FAILS=0. 438 observability and manifest tests OK; the writer-identity modules also pass under bash 3.2.57 as bash and sh (67 OK). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e and practice docs (fix plan steps 7-8) (#374) Replayed onto main 881164e (after #364 and #375) with a 3-way squash. Main had changed three of this PR's files: - The settings template and the macOS page merged cleanly; the template keeps #364's OTEL_METRICS_INCLUDE_SESSION_ID and this PR's 11 Read(**/...) twins. - adoption/bootstrap.md conflicted on one note: it keeps this PR's extended settings-template sentence, followed by #364's writer-identity paragraph. manifests/evidence.json was re-registered, and the matrix and grand list were regenerated. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re-register U6 on main Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that nested quoted expansions inside $(...) parse correctly. What breaks is a case pattern's closing parenthesis, which ends the substitution early. bash -n passes; at run time the shell prints a syntax error and continues with a wrong value. U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json re-registered and the matrix and grand list regenerated. validate FAILS=0. 438 observability and manifest tests OK; the writer-identity modules also pass under bash 3.2.57 as bash and sh (67 OK). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… and Grafana (stacked on #364) (#366) * Tool invoke rates: tool, MCP server, skill and subagent names in Loki and Grafana, with the dated tool-details exception Stacked on the telemetry writer identity change (dfd5460). - Collector: transform/tool_names runs in both logs pipelines before transform/privacy. From Claude Code's tool_parameters it copies only the MCP server and tool names and the Agent tool's subagent_type, plus one boolean, shell_rtk. Skill names come from skill_activated. Registry names that are not short identifiers become "other", and agent types outside the documented built-ins become "custom". tool_parameters, tool_input, arguments, output, content and error are deleted, and the log allowlist gains 19 name, id and derived keys. - Dashboard row 28 (panels 29-38): calls by family, MCP server, skill, client and actor; subagent and workflow launches; MCP share; ctx and rtk adoption; and an integrity panel. - OTEL_LOG_TOOL_DETAILS is "1" in both Claude settings examples. docs/secret-storage.md records it as a dated exception. On 2026-09-26, in the workstation coordinator session, the user answered "1 and full sota convergence practice we proceed". The exception states the control and the canary proof, and marks the template changed after `v2026.09.26.2`. The decision record's status and the new-host template notes in bootstrap.md and macos-arm64.md say the same. - tests/test_observability_tool_names.py: its two fixture thread ids are now built at runtime with uuid.UUID(int=...), with the same values. scripts/validate.py rejected the UUID literals as possible local session ids. Evidence (local integration and synthetic; the host is not applied): otelcol-contrib 0.161.0 validate exits 0. The replay of real Claude Code and Codex captures plus synthetic records passed 109 of 109 checks: 72 canary strings, 0 found in the file exporter or Loki; 19 of 19 dashboard targets. The unit test passes 11 of 11 and fails on dfd5460 alone. validate.py, evidence_manifest --check and the related modules pass. Reviews: two GPT-6 cross-family rounds; every high and medium finding is fixed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * U6: commit the sanitized replay and live-proof receipt, fix the privacy checker, correct client grouping Resolves the review threads on #366 and the window-2 GPT-6 findings 2 and 5. - evidence/artifacts/tool-invoke-rates-20260926/: - the replay harness, its reference implementation and the sanitized 109/109 result (otelcol-contrib 0.161.0, Loki 3.7.8); - the live host proof (2026-09-26T23:47:42Z-23:49:21Z, 33 passed), with counts and hashes only for real captures; - the fixed checker. - prove_check.py and prove.sh: no length floor on forbidden strings, and every tagged record's body must equal the Collector placeholder exactly. - Synthetic A/B: the old checker passed a leaked "pwd" body; the fixed one fails it. - Offline re-check of the retained events-file sink of the 23:47Z proof: 29 tagged lines, all "[content omitted]", 0 forbidden hits. Its Loki sink was not re-checked by the fixed checker, and the decision record says so. - client is documented as a grouping by normalized originator name: separate app-server processes with the same originator share one label. - Panel 34: one description sentence. LogQL rate over unwrap is sum/range (Loki v3.7.8 pkg/logql/range_vector.go rateLogs), so the query is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Launcher comment: only a case pattern inside $(...) breaks bash 3.2; re-register U6 on main Correcting the comment from #364 (8e66134). A bash 3.2.57 re-check shows that nested quoted expansions inside $(...) parse correctly. What breaks is a case pattern's closing parenthesis, which ends the substitution early. bash -n passes; at run time the shell prints a syntax error and continues with a wrong value. U6 is rebased onto main 881164e: its two commits, with manifests/evidence.json re-registered and the matrix and grand list regenerated. validate FAILS=0. 438 observability and manifest tests OK; the writer-identity modules also pass under bash 3.2.57 as bash and sh (67 OK). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * U6 receipt repair (GPT-6 review of #366): check every OTLP record structurally; synthetic-only replay runs natively Findings from the cross-family review of 5ddfd93, each fixed with failing-first controls: - high: prove_check.py kept only the last body of each OTLP batch, and missing or non-string bodies passed. scan_otlp now keeps and validates every record's body. - medium: resource, scope and nested (ArrayValue/KeyValueList) values were not scanned, and banned keys matched only minified JSON. The scan now traverses decoded keys and values at every OTLP level (opentelemetry-proto v1.9.0 logs.proto and common.proto). The live and offline paths share one events_privacy_checks(). Structural controls: 2 passed, 10 failed before; 12 passed after. - medium: the synthetic-only replay failed on max() of an empty capture and required the omitted private captures. It now has a default timestamp and runs the historical assertions only when the full capture set is present. - Stubbed control: 25/20 before, 25/0 after. - Native run by the coordinator on the host, on the pinned Collector 0.161.0 and Loki 3.7.8 with scratch ports 45700-45703: 67 passed, 0 failed. - low: docs/secret-storage.md still said the live proof had not run. The corrected offline re-check of the 23:47Z proof's retained events file ran at 03:55:26Z: 270 records in 29 batches, 3 passed. The exporter (10 MB x 3 backups) rotated that file out at 04:05:11Z, and a re-run at 04:14Z found 0 tagged lines. The receipt records that the result cannot be repeated. The three observability test modules pass on the host (78 OK), where the sandbox had blocked 21 socket tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * U6 receipt, second repair (GPT-6 re-check of 7a15a2e): decode bytesValue; always check Loki bodies - A ProtoJSON bytesValue (base64, standard or URL-safe, with or without padding) was scanned only as encoded text, so bytesValue "cHdk" (pwd) passed. scan_otlp now scans the decoded text too. An undecodable value adds a marker that the checks treat as a banned key, so they fail closed. Controls: 12 passed, 3 failed before (bytes attribute, nested bytes, undecodable); 15 passed, 0 failed after. - Loki's fixed-body privacy assertion ran only with the historical capture set. replay.py loki_body_checks() now runs on every replay: every stored Loki line must equal the Collector placeholder. loki_body_control.py runs the real loki() against a stubbed Loki: on the prior replay.py no body check ran and 31 "pwd" lines went unnoticed (0/2); after the fix, clean records pass and the leak is caught (2/2). - The native synthetic replay on the host (Collector 0.161.0, Loki 3.7.8, scratch ports) now reports 68 passed, 0 failed. - Residual: Loki structured metadata keeps attribute values as strings, so a bytes-typed attribute stays base64 there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…asks, metrics, gate, procedure) Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the baseline-receipt half already merged via #369) before any organic run: 77 frozen task definitions (16 reused Claude + 15 reused Codex information needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3 metric/threshold/guardrail/outcome rule with its three "proposed" bounds now stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process qualification, the Workflow script that will run arms B/A/A0 (syntax-checked only, never executed), and the measurement runbook. Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6 cross-family review (2 findings, 1 high) plus one combined repair round (13/13 addressed with failing-first evidence): relocated from blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/ per adoption-audit/PLAN.md line 402; moved the Workflow script out of examples/claude-native/workflows/ to avoid that directory's shared PACKET/ROUTING contract test, which this one-off frozen script correctly does not match; corrected PR-A's status (its base branch merged as #369 -- baseline only, tooling-fix scope still open); added the six AA §6 machine-readable policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle binding; blocked the one information need (an otel-tui receiver task) that no current role may legally run, pending a named permitted role; froze the coordinator's own main-dispatch task at xhigh instead of max; fixed a denylist word-boundary gap that let underscore-qualified MCP names slip past; and sealed the five normative files' SHA256 with an append-only amendment rule. Also refreshes PR #376 and #364's status to merged (both landed while this unit was in review), while recording that #364's field-preservation gap (workflow.run_id/tool_use_id still missing from collector.yaml) persists after its merge. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…asks, (#381) metrics, gate, procedure) Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the baseline-receipt half already merged via #369) before any organic run: 77 frozen task definitions (16 reused Claude + 15 reused Codex information needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3 metric/threshold/guardrail/outcome rule with its three "proposed" bounds now stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process qualification, the Workflow script that will run arms B/A/A0 (syntax-checked only, never executed), and the measurement runbook. Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6 cross-family review (2 findings, 1 high) plus one combined repair round (13/13 addressed with failing-first evidence): relocated from blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/ per adoption-audit/PLAN.md line 402; moved the Workflow script out of examples/claude-native/workflows/ to avoid that directory's shared PACKET/ROUTING contract test, which this one-off frozen script correctly does not match; corrected PR-A's status (its base branch merged as #369 -- baseline only, tooling-fix scope still open); added the six AA §6 machine-readable policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle binding; blocked the one information need (an otel-tui receiver task) that no current role may legally run, pending a named permitted role; froze the coordinator's own main-dispatch task at xhigh instead of max; fixed a denylist word-boundary gap that let underscore-qualified MCP names slip past; and sealed the five normative files' SHA256 with an append-only amendment rule. Also refreshes PR #376 and #364's status to merged (both landed while this unit was in review), while recording that #364's field-preservation gap (workflow.run_id/tool_use_id still missing from collector.yaml) persists after its merge. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Correct the Claude custom-attribute citation and resource/event/datapoint description, whole-run prefix queries, deployment path and dated E2E gates. Publish original and repair command/result excerpts directly in the PR body. Explicitly delete Codex arguments with native OTTL and preserve tool namespace, guarded MCP server, timing, success, correlation and supplied size metadata. Test duplicate Claude run attributes and Codex records with and without server metadata through installed Collector 0.161.0 and Loki 3.7.8. Sources: open-telemetry/opentelemetry-collector-contrib v0.161.0, processor/transformprocessor/README.md and pkg/ottl/ottlfuncs/README.md; openai/codex rust-v0.157.1, codex-rs/otel/src/tool_result.rs:55-79,94, codex-rs/otel/src/events/shared.rs:59-60 and codex-rs/core/src/tools/call_trace.rs:38-54; https://code.claude.com/docs/en/monitoring-usage#multi-team-organization-support; https://grafana.com/docs/loki/latest/query/log_queries/ and send-data/otel/; open-telemetry/opentelemetry-proto v1.9.0 docs/specification.md; base ec27a30, #364 c71d66d and #366 a464d28 historical receipts. Verification: failing-first C1 checks, then 83 covering unittests pass with no skips; complete native Collector validate and git diff --check pass. Evidence is local integration with synthetic fixtures, not live-provider/host acceptance. The coordinator owns receipt hash registration and commit creation. Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dex tool arguments, whole-run query guidance (#408) * Preserve Claude tool joins and estimated cost through Loki Retain tool_use_id and cost_usd alongside the existing workflow.run_id, session.id and ecosystem.task.id resource route. Keep the current structured metadata policy and metric writer identity. Add red-before-green unittest coverage through native Collector 0.161.0 and Loki 3.7.8, including per-arm usage, resource-only runs, private tool joins and indexed-series checks. Sources: Anthropic monitoring documentation, checked 2026-09-27: https://code.claude.com/docs/en/monitoring-usage anthropics/claude-code v2.1.283 release notes and installed client. open-telemetry/opentelemetry-collector-contrib v0.161.0: processor/transformprocessor/README.md and pkg/ottl/ottlfuncs/func_keep_keys.go. grafana/loki v3.7.8 with https://grafana.com/docs/loki/latest/send-data/otel/ and https://grafana.com/docs/loki/latest/reference/loki-http-api/. open-telemetry/opentelemetry-proto v1.9.0, docs/specification.md. Validation: 81 covering unittest checks passed without skips; the complete Collector profile passed native validation after creating documented scratch directories. Local integration uses synthetic inputs, not a new model run. Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Repair collector evidence and Codex tool metadata privacy Correct the Claude custom-attribute citation and resource/event/datapoint description, whole-run prefix queries, deployment path and dated E2E gates. Publish original and repair command/result excerpts directly in the PR body. Explicitly delete Codex arguments with native OTTL and preserve tool namespace, guarded MCP server, timing, success, correlation and supplied size metadata. Test duplicate Claude run attributes and Codex records with and without server metadata through installed Collector 0.161.0 and Loki 3.7.8. Sources: open-telemetry/opentelemetry-collector-contrib v0.161.0, processor/transformprocessor/README.md and pkg/ottl/ottlfuncs/README.md; openai/codex rust-v0.157.1, codex-rs/otel/src/tool_result.rs:55-79,94, codex-rs/otel/src/events/shared.rs:59-60 and codex-rs/core/src/tools/call_trace.rs:38-54; https://code.claude.com/docs/en/monitoring-usage#multi-team-organization-support; https://grafana.com/docs/loki/latest/query/log_queries/ and send-data/otel/; open-telemetry/opentelemetry-proto v1.9.0 docs/specification.md; base ec27a30, #364 c71d66d and #366 a464d28 historical receipts. Verification: failing-first C1 checks, then 83 covering unittests pass with no skips; complete native Collector validate and git diff --check pass. Evidence is local integration with synthetic fixtures, not live-provider/host acceptance. The coordinator owns receipt hash registration and commit creation. Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Collector README: state that the sealed #381 receipts are unchanged; drop an unsupported #366 citation Coordinator correction after the independent repair verification: the repair's receipt edits were withdrawn (RUNBOOK.md is hash-sealed; the E2E README changes only by amendment), so the README must not say they carry dated updates; the #366 receipt does not record a byte-identical host deployment. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Hot files: evidence registration and generated reports Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Summary
Each Prometheus token series now has one writer. Before this change, every Claude process and every
directly started Codex process wrote without a
service.instance.id, so concurrent cumulative streamsoverwrote each other. On the workstation that meant 7,216 Claude counter resets in an hour,
Prometheus
increase()at 73 to 193 times the Lokiapi_requestsums, and about 3,518 of 8,160 deltapoints rejected by
delta_to_cumulative(read-only receipthost-before-20260926.json).OTEL_METRICS_INCLUDE_SESSION_ID=true, the upstreamdefault). The Collector turns them into
service.instance.id(groupbyattrs/session, thentransform/privacy) and sums the points of dropped attributes (aggregate_on_attributes).observability/collector/codex-identity-launcher.sh.examplegives each process its own idthrough
OTEL_RESOURCE_ATTRIBUTES. This PR's host apply leaves it out; see Residuals.created-timestamp-zero-ingestionandpromql-extended-range-selectors. Thecollector-nativejob drops the Codex histogram buckets except token usage.native-telemetry-integritygroup covers per-series resets, dropped deltapoints and unscoped writers. The dashboards exclude
instance="unscoped".evidence/artifacts/telemetry-writer-identity-20260926/host/(apply.sh,rollback.sh,prove.sh/prove.pyand three merge/setting helpers).Scope
unit flags, scrape config and rules, the dashboards, the host recipe, the READMEs and new-host
pages that describe them (marked "changed after
v2026.09.26.2"), and their tests.803bc351.git merge-treeagainstorigin/maindde28cc2is clean.lane:foundation.manifests/evidence.jsonis only re-registered.observability/collector/,observability/backends/,observability/native-data/,observability/grand-dashboard/render.py,adoption/,docs/decisions/,evidence/artifacts/telemetry-writer-identity-20260926/andtests/.claude/tool-invoke-rates-20260926). It retargets tomainafter this PR merges.What was verified
dfd5460fscripts/validate.py(6,934 hashed files, 159 receipts),evidence_manifest.py --check,component_matrix.py --check,new_host_grand_list.py --checkandbuild_ecosystem.py --checkall exit 0unittest discover -s tests -q: 6,281 tests OK (skipped=654), first runprove.pyqueries on a scratch TSDB with both features on; the loopback Loki 3.7.8 accepted 16 of 16 LogQL queriessynthetic-ab-20260926.json: old profile Claude −51.3% (anchored), 4 resets, 24ErrOutOfOrderpoints, Codex −6.7%; new profile 0.0% anchored for both, 0 resets, 0 dropped points--applyapply.sh --dry-runexits 0 and creates no state directory; read-onlyprove.pyreports FAIL before the apply, as expected (anchored modifier not enabled)apply.sh,rollback.shandprove.sh, andbash -npassesReviews (bounded to two rounds; the notes are in the session scratchpad):
p2rev/review-u1.md) found 12 defectsand a nit, none blocking the design. All were fixed, plus 3 more found during the repair: the
anchored modifier's position, the Codex active window and partial rollback. Failing-first:
38 tests gave 22 failures and 3 errors before, and 38 of 38 passed after. The resolution is
u1-p2/REPAIR-1.md, commit6ad0d358.gpt-6-astra, effort max, read-only) found 1 high and6 medium issues. All seven held against the source, and all are fixed with tests that failed
first (9 failures and 1 error before). The resolution is
u1-p2/REPAIR-2.md, commitdfd5460f.prove.pycompares only processes that completed a turn;upsamples with the scrape interval;--applyreads the running services back;Host-apply plan (not run yet)
The coordinator applies this in one window with the stacked invoke-rate change.
$REPOis acheckout holding both. The step order is in that PR's plan: install its Collector file without a
restart, then run this recipe once.
evidence/artifacts/telemetry-writer-identity-20260926/host/apply.sh --dry-run --repo "$REPO" --only prometheus,dashboards,collector,claude-settingis read-only: it renders, validates (otelcol, promtool, systemd-analyze) and diffs.
--applybacks up to${XDG_STATE_HOME:-~/.local/state}/native-agent-stack/g1-writer-identity/backup-<UTC>/(0700, with
manifest.json). It installs, restarts the Collector and Prometheus once, and readsboth back. Any failed read-back stops it and names the rollback.
prove.shruns after a scenario window with Claude and concurrent Codex activity.rollback.sh --applyis all or nothing and refuses files or the settings key if theychanged after the apply, unless
--force.The
codex-launcherstep is deliberately left out of--only.Residuals
per-process OTel identity. At
rust-v0.157.1, Codex builds its resource from the service name,version, env and OS, plus
OTEL_RESOURCE_ATTRIBUTES, and exports no per-process attribute. Thetemplate stays in the tree as the documented local-integration path. Until it is installed, the
host sees the following:
instance="unscoped", except for callers that set their own id.Panels 1 and 26 ("Codex turn tokens") exclude them, so Codex tokens show only in Loki and the
native-data panels.
EcosystemUnscopedTokenWritersfires while Codex exports counters.prove.pyfailssetup.codex_launcher, so its PASS is unreachable as written. The coordinatordecides whether to accept that failure or amend the proof.
prove.pydefaults to--prom :19090and--loki :13100.lastConfigTimehas whole-second resolution.manifests/evidence.json.Whichever merges second takes
main's copy and re-registers its own files (hot-file protocol indocs/lanes.md).otlphttpalias is deprecated;merge_collector.pyreplaces the whole metrics processor list andmetric_statements, so ahost-specific metrics processor would be dropped without a refusal.
SOTA sources
2026-09-26T15:01Z):
OTEL_METRICS_INCLUDE_SESSION_IDdefaulttrue,session.idon metrics,cumulative temporality.
v0.161.0:processor/deltatocumulativeprocessor/README.md(ErrOutOfOrder,ErrOlderStart);processor/groupbyattrsprocessor/README.md;processor/transformprocessor/README.md(aggregate_on_attributes);exporter/prometheusexporter/README.md(metric_expiration5m, start timestamps).v3.15.0:docs/feature_flags.md(created-timestamp-zero-ingestion,promql-extended-range-selectorsand
anchored);docs/configuration/configuration.md(metric_relabel_configs).rust-v0.157.1(36650394c5b3):codex-rs/otel/src/metrics/client.rsandcodex-rs/otel/src/provider.rsbuild the resource with no per-process attribute;OTEL_RESOURCE_ATTRIBUTESreaches it throughEnvResourceDetector(opentelemetry_sdk 0.31.0).Decision record
docs/decisions/2026-09-26-telemetry-writer-identity.mdcovers the alternatives, the overturnconditions and the limits.
Window 2: host apply, proof, GPT-6 review and repair round 3 (coordinator, 2026-09-26/27)
34c56340as924ec302, under the hot-file protocol (main'smanifests/evidence.json, 36 files re-registered). The native-data files from Native data (G2): token-tool savings into Loki and Grafana, one series per native scope #365 and Native-data snapshot: read token reports up to their documented embedded size #373 auto-merged and were checked.validate.pyFAILS=0; verdict gate reports no rows changed; the full suite is OK (skipped=642).OTEL_LOG_TOOL_DETAILS=1;prove.sh: 33/33;prove.sh: exit 1 on the two permitted checks only (unscoped_token_writersandfiring_integrity_alerts), bothjob=claude-codeand attributed by process capture to the pre-change coordinator. Its 60 frozen unscoped series live until that process exits. Claude writers and Codex processes reconcile at 0.00%, and no Codex series is unscoped.f4f772cb, with failing-first tests:__codex_identity_prefix and are assigned in a subshell; the child environment and exact argv are tested.merge_collector.pynow verifies statements, groups, settings and pipeline order, and refuses before writing a candidate.apply.shrecords render-time hashes and re-checks before and at each install.94 targeted tests OK; validate FAILS=0. The anti-pattern log records the three lessons.
api_requestevent, so a writer that uses the advisor can fail the per-writer comparison. Recorded as a comparison and coverage limitation.tests.test_landscape_sweep_harnessfails intermittently on the workstation under concurrent load, on main as well; hosted CI decides.🤖 Generated with Claude Code