forked from NousResearch/hermes-agent
-
Notifications
You must be signed in to change notification settings - Fork 0
fix(compression): inject memory-provider on_pre_compress() text into the summary #5
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
Closed
Changes from all commits
Commits
Show all changes
74 commits
Select commit
Hold shift + click to select a range
8de2099
feat(discord): auto-detect choice prompts and offer clickable buttons
sam7894604 afdab2c
test(discord): cover auto-detected choice buttons
sam7894604 53f985e
feat(discord): prefix-gated + multi-select auto-choice buttons
sam7894604 8667125
docs(discord): add README and session choice hint
sam7894604 4e3dcd7
feat: split session row on mid-session model switch
e518b99
fix: gateway split fallback when no cached agent
767af03
fix: get_messages follows model_switch parent chain for context conti…
2cd0c6a
fix: gateway load_transcript follows model_switch ancestor chain
5d86f30
feat(tokens): add bit-packed codec for messages.token_count
sam7894604 9b514b7
feat(state): query API for bit-packed message token counts
sam7894604 62d4301
feat(gateway): track reasoning_tokens on SessionEntry
sam7894604 295e113
feat(tokens): write + display path for bit-packed message tokens
sam7894604 76d7c6b
fix(tokens): flatten packed token_count on read for display consumers
sam7894604 19dd6bc
refactor(tokens): make input-token prompt-tail attribution verifiable
sam7894604 e78d064
fix(tokens): decode token_count in JSON snapshots and exports
sam7894604 12acd8f
feat(api): expose decoded per-message token totals endpoint
sam7894604 5cbc946
feat(achievements): surface decoded per-session token usage
sam7894604 39d3eae
test: skip token-totals endpoint test when aiohttp is absent
sam7894604 61037a8
feat(gateway): /tokens toggle + per-message token footer on replies
sam7894604 930c337
feat(discord): native /tokens slash command
sam7894604 66c0731
test(api): lock chat transcript per-message tokens contract
sam7894604 39c5949
feat(analytics): provider quota reference + endpoint
sam7894604 7e8fc0b
feat(analytics): usage-rates endpoint (RPM/RPD/TPM/TPD vs limits)
sam7894604 ecbe5a3
feat(analytics): token-trends endpoint (avg/call, cache-hit, time ser…
sam7894604 5e3d755
feat(analytics): cost-estimate endpoint (per-tier pricing + projection)
sam7894604 411e748
test(analytics): gate only endpoint tests on aiohttp, not whole module
sam7894604 2c65fee
feat(gateway): /tokens on|off|always (per-session + global)
sam7894604 541f86f
feat(discord): /tokens always choice (global) in native slash command
sam7894604 c440a4d
feat(tui): /tokens per-message token display toggle
sam7894604 f3e81d4
feat(tokens): shared format_token_count + compact backtick footer
sam7894604 3122c5a
feat(tui): short K/M token footer matching the gateway format
sam7894604 b06d882
feat(dashboard): analytics UI for usage-rates/trends/cost/quotas
sam7894604 d479a38
fix(packaging): ship hermes_token_codec as a top-level py-module
sam7894604 37d08c5
test(tokens): standalone read-only verifier for bit-packed token_count
sam7894604 8a2e4f2
feat(tokens): opt-in live per-turn API-vs-packed verification log
sam7894604 aee2e5d
fix(usage): capture reasoning_tokens from completion_tokens_details
sam7894604 0878bc1
fix(tokens): restore first-writer-wins billing route after upstream r…
sam7894604 b18ff6a
fix(tokens): keep first accounted model on session row after upstream…
sam7894604 f4b5a41
feat(line): _LineClient name resolution API + bot-mention parse helper
sam7894604 2b2048c
feat(line): P1 WhitelistStore + reject notify/dedup
sam7894604 c3189d1
feat(line): P3 agent approval tool
sam7894604 e5cb847
feat(line): P2 dashboard whitelist plugin
sam7894604 ea0a74b
feat(line): wire dashboard discovery + line_whitelist toolset gating
sam7894604 de4a992
feat(line): P4 adapter integration — gate/mention/reject/observe/quote
sam7894604 6220854
fix(line): route unauthorized_notify to target platform (telegram:/di…
sam7894604 a2ef65b
fix(line): whitelist DELETE was 404 on success — remove() now returns…
sam7894604 7dd687b
feat(line): WhitelistStore pending-queue API
sam7894604 b139064
feat(line): dashboard pending-queue panel
sam7894604 53d7a40
feat(line): record unauthorized attempts into pending queue + name re…
sam7894604 1e41dcf
feat(line): telegram+discord interactive whitelist-decision cards
sam7894604 a339a58
feat(line): route unauthorized notify to interactive card on telegram…
sam7894604 255674d
feat(line): dashboard authorized-list with names + admin lock + env o…
sam7894604 0b7983e
fix(line): pending 'approve' was a no-op for entries lacking source_type
sam7894604 0bf2cac
fix(line): dashboard COMMUNICATION RECORDS always empty (session_key …
sam7894604 b8bd0e6
fix(line): dashboard add/remove USER failed with 400 'unknown scope: …
sam7894604 b9ab3e3
refactor(line): dashboard — drop redundant Allowlist entry list
sam7894604 35a29f5
fix(line): interactive-card admin check rejected the notify recipient
sam7894604 ee20886
feat(line): store — card_admins table + managed settings get/set
sam7894604 d73a63f
feat(line): dashboard Settings panel (card_admins + config settings +…
sam7894604 f02b48f
fix(line): fail-open @mention gate when bot userId is unknown
sam7894604 fe54237
fix(line): route inbound audio/video/file to correct cache (not image…
sam7894604 3fdf3ef
feat(line+gateway): preserve PDF filename + auto-extract PDF text (vi…
sam7894604 99fd277
feat(gateway): generalize auto-extraction beyond PDF (text/csv/docx/x…
sam7894604 60ba98d
feat(gateway): full Office coverage — LibreOffice bridge for pptx + l…
sam7894604 f1582bd
fix(gateway): legacy spreadsheets via LibreOffice convert to XLSX, no…
sam7894604 cb0c398
feat(line): quote-reply = implicit mention + pre-extract observed media
sam7894604 5430f0c
feat(line): on-demand media backfill (replaces observe pre-extraction)
sam7894604 a5338a1
feat(line): backfill extracts+caches media, injects as channel_context
sam7894604 50bfc63
fix(agent): reliable turbovault edit_note + verifier false-alarm (B+C)
sam7894604 93d3840
fix(line): convert markdown tables to bullets on outbound
sam7894604 d8423ff
fix(state): get_conversation_root walks all parent links after upstre…
sam7894604 16ddf28
fix(tokens): include token_count in _CONVERSATION_ROW_COLUMNS after u…
sam7894604 8792681
fix(compression): inject memory-provider on_pre_compress() text into …
sam7894604 6cf37ef
test: expect provider_context kwarg in force-bypass compress call (PR…
sam7894604 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,234 @@ | ||
| """Pure analytics computations over decoded token time-series. | ||
|
|
||
| DB-free so the math is unit-testable in isolation. Inputs come from | ||
| :meth:`hermes_state.SessionDB.get_message_token_timeseries` (1-minute | ||
| buckets) plus provider quota dicts from :mod:`agent.provider_quotas`. | ||
| """ | ||
| from __future__ import annotations | ||
|
|
||
| from typing import Any, Dict, List, Optional | ||
|
|
||
|
|
||
| def _pct(value: int, limit: Optional[int]) -> Optional[float]: | ||
| if not limit or limit <= 0: | ||
| return None | ||
| return round(100.0 * value / limit, 1) | ||
|
|
||
|
|
||
| def _stat(values: List[int]) -> Dict[str, float]: | ||
| """min/max/mean/median/p95 over a list (0s for empty).""" | ||
| if not values: | ||
| return {"min": 0, "max": 0, "mean": 0.0, "median": 0.0, "p95": 0} | ||
| s = sorted(values) | ||
| n = len(s) | ||
| mean = sum(s) / n | ||
| median = s[n // 2] if n % 2 else (s[n // 2 - 1] + s[n // 2]) / 2 | ||
| p95 = s[min(n - 1, int(round(0.95 * (n - 1))))] | ||
| return {"min": s[0], "max": s[-1], "mean": round(mean, 2), "median": round(median, 2), "p95": p95} | ||
|
|
||
|
|
||
| def compute_usage_rates( | ||
| minute_buckets: List[Dict[str, int]], | ||
| daily_totals: Dict[str, int], | ||
| provider_quotas: List[Dict[str, Any]], | ||
| ) -> Dict[str, Any]: | ||
| """RPM/TPM peaks + RPD/TPD totals, compared to provider limits. | ||
|
|
||
| ``minute_buckets``: 1-minute time-series rows (requests/input/output/…). | ||
| ``daily_totals``: aggregate over the last 24h ({requests,input,output,…}). | ||
| ``provider_quotas``: quota dicts for providers active in the window. | ||
| """ | ||
| req = [b.get("requests", 0) for b in minute_buckets] | ||
| tpm = [b.get("input", 0) + b.get("output", 0) for b in minute_buckets] | ||
| tpm_in = [b.get("input", 0) for b in minute_buckets] | ||
| tpm_out = [b.get("output", 0) for b in minute_buckets] | ||
|
|
||
| peak_rpm = max(req, default=0) | ||
| peak_tpm = max(tpm, default=0) | ||
| current_rpm = req[-1] if req else 0 | ||
| current_tpm = tpm[-1] if tpm else 0 | ||
|
|
||
| rpd = int(daily_totals.get("requests", 0)) | ||
| tpd = int(daily_totals.get("input", 0)) + int(daily_totals.get("output", 0)) | ||
|
|
||
| providers_view: List[Dict[str, Any]] = [] | ||
| for q in provider_quotas: | ||
| # Peak-vs-limit %, conservative (treats the global peak as if it all | ||
| # went to this provider — exact for the single-provider common case). | ||
| providers_view.append({ | ||
| "provider": q.get("provider"), | ||
| "display": q.get("display"), | ||
| "tier": q.get("tier"), | ||
| "limits": { | ||
| "rpm": q.get("rpm"), "rpd": q.get("rpd"), | ||
| "tpm_input": q.get("tpm_input"), "tpm_output": q.get("tpm_output"), | ||
| "tpd": q.get("tpd"), | ||
| }, | ||
| "pct_of_limit": { | ||
| "rpm": _pct(peak_rpm, q.get("rpm")), | ||
| "rpd": _pct(rpd, q.get("rpd")), | ||
| "tpm_input": _pct(max(tpm_in, default=0), q.get("tpm_input")), | ||
| "tpm_output": _pct(max(tpm_out, default=0), q.get("tpm_output")), | ||
| "tpd": _pct(tpd, q.get("tpd")), | ||
| }, | ||
| "source_url": q.get("source_url"), | ||
| "as_of": q.get("as_of"), | ||
| }) | ||
|
|
||
| return { | ||
| "rpm": {"current": current_rpm, "peak": peak_rpm}, | ||
| "tpm": { | ||
| "current": current_tpm, "peak": peak_tpm, | ||
| "peak_input": max(tpm_in, default=0), "peak_output": max(tpm_out, default=0), | ||
| }, | ||
| "rpd": rpd, | ||
| "tpd": tpd, | ||
| "window_totals": { | ||
| "requests": sum(req), | ||
| "input": sum(tpm_in), | ||
| "output": sum(tpm_out), | ||
| }, | ||
| "providers": providers_view, | ||
| } | ||
|
|
||
|
|
||
| def compute_token_trends(buckets: List[Dict[str, int]]) -> Dict[str, Any]: | ||
| """Per-bucket series + per-request averages + cache-hit rate. | ||
|
|
||
| ``buckets``: time-series rows at the caller's chosen granularity. | ||
| """ | ||
| series: List[Dict[str, Any]] = [] | ||
| per_call_input: List[int] = [] | ||
| per_call_output: List[int] = [] | ||
| total_input = total_output = total_cache = total_reasoning = total_req = 0 | ||
|
|
||
| for b in buckets: | ||
| reqs = int(b.get("requests", 0)) | ||
| inp = int(b.get("input", 0)) | ||
| out = int(b.get("output", 0)) | ||
| cache = int(b.get("cache_read", 0)) | ||
| reason = int(b.get("reasoning", 0)) | ||
| total_input += inp | ||
| total_output += out | ||
| total_cache += cache | ||
| total_reasoning += reason | ||
| total_req += reqs | ||
| cache_hit = round(100.0 * cache / inp, 1) if inp else None | ||
| series.append({ | ||
| "bucket_start": int(b.get("bucket_start", 0)), | ||
| "requests": reqs, | ||
| "input": inp, "output": out, "cache_read": cache, "reasoning": reason, | ||
| "cache_hit_rate": cache_hit, | ||
| "avg_input_per_request": round(inp / reqs, 1) if reqs else 0, | ||
| "avg_output_per_request": round(out / reqs, 1) if reqs else 0, | ||
| }) | ||
| if reqs: | ||
| per_call_input.append(round(inp / reqs)) | ||
| per_call_output.append(round(out / reqs)) | ||
|
|
||
| overall_cache_hit = round(100.0 * total_cache / total_input, 1) if total_input else None | ||
| return { | ||
| "series": series, | ||
| "totals": { | ||
| "requests": total_req, "input": total_input, "output": total_output, | ||
| "cache_read": total_cache, "reasoning": total_reasoning, | ||
| }, | ||
| "averages_per_request": { | ||
| "input": round(total_input / total_req, 1) if total_req else 0, | ||
| "output": round(total_output / total_req, 1) if total_req else 0, | ||
| "reasoning": round(total_reasoning / total_req, 1) if total_req else 0, | ||
| "cache_read": round(total_cache / total_req, 1) if total_req else 0, | ||
| "input_distribution": _stat(per_call_input), | ||
| "output_distribution": _stat(per_call_output), | ||
| }, | ||
| "cache_hit_rate": overall_cache_hit, | ||
| } | ||
|
|
||
|
|
||
| def _price(per_million: Any, tokens: int) -> Optional[float]: | ||
| """Cost in USD for ``tokens`` at ``per_million`` USD/1M, or None if unpriced.""" | ||
| if per_million is None: | ||
| return None | ||
| return round(float(per_million) * tokens / 1_000_000.0, 6) | ||
|
|
||
|
|
||
| def compute_cost_estimate( | ||
| groups: List[Dict[str, Any]], | ||
| window_seconds: int, | ||
| price_lookup, | ||
| ) -> Dict[str, Any]: | ||
| """Per-model cost broken down by price tier, with daily/monthly projection. | ||
|
|
||
| ``groups``: rows from | ||
| :meth:`hermes_state.SessionDB.get_session_cost_aggregates`. | ||
| ``price_lookup(model, provider, base_url)``: returns a pricing entry with | ||
| ``input_cost_per_million`` / ``output_cost_per_million`` / | ||
| ``cache_read_cost_per_million`` / ``cache_write_cost_per_million`` (or | ||
| None). When a model has no known pricing the stored ``estimated_cost_usd`` | ||
| is used as a fallback and flagged. | ||
| """ | ||
| models: List[Dict[str, Any]] = [] | ||
| total = 0.0 | ||
| total_input_cost = total_output_cost = total_cache_cost = 0.0 | ||
| any_unpriced = False | ||
|
|
||
| for g in groups: | ||
| entry = None | ||
| try: | ||
| entry = price_lookup(g.get("model"), g.get("billing_provider"), g.get("billing_base_url")) | ||
| except Exception: | ||
| entry = None | ||
|
|
||
| in_cost = out_cost = cr_cost = cw_cost = None | ||
| if entry is not None: | ||
| in_cost = _price(getattr(entry, "input_cost_per_million", None), g["input_tokens"]) | ||
| out_cost = _price(getattr(entry, "output_cost_per_million", None), g["output_tokens"]) | ||
| cr_cost = _price(getattr(entry, "cache_read_cost_per_million", None), g["cache_read_tokens"]) | ||
| cw_cost = _price(getattr(entry, "cache_write_cost_per_million", None), g["cache_write_tokens"]) | ||
|
|
||
| priced = any(c is not None for c in (in_cost, out_cost, cr_cost, cw_cost)) | ||
| if priced: | ||
| grp_total = round(sum(c or 0.0 for c in (in_cost, out_cost, cr_cost, cw_cost)), 6) | ||
| source = "pricing" | ||
| else: | ||
| grp_total = round(float(g.get("estimated_cost_usd") or 0.0), 6) | ||
| source = "stored_estimate" | ||
| any_unpriced = True | ||
|
|
||
| total += grp_total | ||
| total_input_cost += in_cost or 0.0 | ||
| total_output_cost += out_cost or 0.0 | ||
| total_cache_cost += (cr_cost or 0.0) + (cw_cost or 0.0) | ||
|
|
||
| models.append({ | ||
| "model": g.get("model"), | ||
| "provider": g.get("billing_provider"), | ||
| "sessions": g.get("sessions", 0), | ||
| "tokens": { | ||
| "input": g["input_tokens"], "output": g["output_tokens"], | ||
| "cache_read": g["cache_read_tokens"], "cache_write": g["cache_write_tokens"], | ||
| "reasoning": g["reasoning_tokens"], | ||
| }, | ||
| "cost_breakdown": { | ||
| "input": in_cost, "output": out_cost, | ||
| "cache_read": cr_cost, "cache_write": cw_cost, | ||
| }, | ||
| "cost_usd": grp_total, | ||
| "cost_source": source, | ||
| }) | ||
|
|
||
| models.sort(key=lambda m: m["cost_usd"], reverse=True) | ||
| total = round(total, 6) | ||
| days = max(window_seconds / 86400.0, 1e-9) | ||
| daily = round(total / days, 6) | ||
| return { | ||
| "total_cost_usd": total, | ||
| "cost_by_tier": { | ||
| "input": round(total_input_cost, 6), | ||
| "output": round(total_output_cost, 6), | ||
| "cache": round(total_cache_cost, 6), | ||
| }, | ||
| "projection": {"daily_usd": daily, "monthly_usd": round(daily * 30, 6)}, | ||
| "has_unpriced_models": any_unpriced, | ||
| "models": models, | ||
| } |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
Append a redacted provider-context section locally instead of asking the LLM to echo raw text.
This block sends
_pending_provider_contextto the auxiliary model before redaction and only requests verbatim preservation. If the provider text contains credentials, it bypasses the existing_serialize_for_summary()redaction path; if the summarizer omits/alters it or summary generation falls back to_build_static_fallback_summary(), the provider context still does not reliably survive compaction. Prefer formatting a redacted## Memory Provider Contextsection in Python and appending it to both successful LLM summaries and deterministic fallback summaries.Suggested direction
Then append
provider_sectionafterredact_sensitive_text(content.strip()), and also after_build_static_fallback_summary(...)whensummaryis missing.🤖 Prompt for AI Agents