Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
74 commits
Select commit Hold shift + click to select a range
8de2099
feat(discord): auto-detect choice prompts and offer clickable buttons
sam7894604 Jun 26, 2026
afdab2c
test(discord): cover auto-detected choice buttons
sam7894604 Jun 26, 2026
53f985e
feat(discord): prefix-gated + multi-select auto-choice buttons
sam7894604 Jun 26, 2026
8667125
docs(discord): add README and session choice hint
sam7894604 Jun 27, 2026
4e3dcd7
feat: split session row on mid-session model switch
Jun 23, 2026
e518b99
fix: gateway split fallback when no cached agent
Jun 23, 2026
767af03
fix: get_messages follows model_switch parent chain for context conti…
Jun 23, 2026
2cd0c6a
fix: gateway load_transcript follows model_switch ancestor chain
Jun 24, 2026
5d86f30
feat(tokens): add bit-packed codec for messages.token_count
sam7894604 Jun 25, 2026
9b514b7
feat(state): query API for bit-packed message token counts
sam7894604 Jun 25, 2026
62d4301
feat(gateway): track reasoning_tokens on SessionEntry
sam7894604 Jun 25, 2026
295e113
feat(tokens): write + display path for bit-packed message tokens
sam7894604 Jun 25, 2026
76d7c6b
fix(tokens): flatten packed token_count on read for display consumers
sam7894604 Jun 25, 2026
19dd6bc
refactor(tokens): make input-token prompt-tail attribution verifiable
sam7894604 Jun 25, 2026
e78d064
fix(tokens): decode token_count in JSON snapshots and exports
sam7894604 Jun 25, 2026
12acd8f
feat(api): expose decoded per-message token totals endpoint
sam7894604 Jun 25, 2026
5cbc946
feat(achievements): surface decoded per-session token usage
sam7894604 Jun 25, 2026
39d3eae
test: skip token-totals endpoint test when aiohttp is absent
sam7894604 Jun 25, 2026
61037a8
feat(gateway): /tokens toggle + per-message token footer on replies
sam7894604 Jun 25, 2026
930c337
feat(discord): native /tokens slash command
sam7894604 Jun 25, 2026
66c0731
test(api): lock chat transcript per-message tokens contract
sam7894604 Jun 25, 2026
39c5949
feat(analytics): provider quota reference + endpoint
sam7894604 Jun 25, 2026
7e8fc0b
feat(analytics): usage-rates endpoint (RPM/RPD/TPM/TPD vs limits)
sam7894604 Jun 25, 2026
ecbe5a3
feat(analytics): token-trends endpoint (avg/call, cache-hit, time ser…
sam7894604 Jun 25, 2026
5e3d755
feat(analytics): cost-estimate endpoint (per-tier pricing + projection)
sam7894604 Jun 25, 2026
411e748
test(analytics): gate only endpoint tests on aiohttp, not whole module
sam7894604 Jun 25, 2026
2c65fee
feat(gateway): /tokens on|off|always (per-session + global)
sam7894604 Jun 25, 2026
541f86f
feat(discord): /tokens always choice (global) in native slash command
sam7894604 Jun 25, 2026
c440a4d
feat(tui): /tokens per-message token display toggle
sam7894604 Jun 25, 2026
f3e81d4
feat(tokens): shared format_token_count + compact backtick footer
sam7894604 Jun 25, 2026
3122c5a
feat(tui): short K/M token footer matching the gateway format
sam7894604 Jun 25, 2026
b06d882
feat(dashboard): analytics UI for usage-rates/trends/cost/quotas
sam7894604 Jun 25, 2026
d479a38
fix(packaging): ship hermes_token_codec as a top-level py-module
sam7894604 Jun 30, 2026
37d08c5
test(tokens): standalone read-only verifier for bit-packed token_count
sam7894604 Jun 30, 2026
8a2e4f2
feat(tokens): opt-in live per-turn API-vs-packed verification log
sam7894604 Jun 30, 2026
aee2e5d
fix(usage): capture reasoning_tokens from completion_tokens_details
sam7894604 Jun 30, 2026
0878bc1
fix(tokens): restore first-writer-wins billing route after upstream r…
sam7894604 Jul 9, 2026
b18ff6a
fix(tokens): keep first accounted model on session row after upstream…
sam7894604 Jul 11, 2026
f4b5a41
feat(line): _LineClient name resolution API + bot-mention parse helper
sam7894604 Jul 6, 2026
2b2048c
feat(line): P1 WhitelistStore + reject notify/dedup
sam7894604 Jul 6, 2026
c3189d1
feat(line): P3 agent approval tool
sam7894604 Jul 6, 2026
e5cb847
feat(line): P2 dashboard whitelist plugin
sam7894604 Jul 6, 2026
ea0a74b
feat(line): wire dashboard discovery + line_whitelist toolset gating
sam7894604 Jul 6, 2026
de4a992
feat(line): P4 adapter integration — gate/mention/reject/observe/quote
sam7894604 Jul 6, 2026
6220854
fix(line): route unauthorized_notify to target platform (telegram:/di…
sam7894604 Jul 6, 2026
a2ef65b
fix(line): whitelist DELETE was 404 on success — remove() now returns…
sam7894604 Jul 6, 2026
7dd687b
feat(line): WhitelistStore pending-queue API
sam7894604 Jul 6, 2026
b139064
feat(line): dashboard pending-queue panel
sam7894604 Jul 6, 2026
53d7a40
feat(line): record unauthorized attempts into pending queue + name re…
sam7894604 Jul 6, 2026
1e41dcf
feat(line): telegram+discord interactive whitelist-decision cards
sam7894604 Jul 6, 2026
a339a58
feat(line): route unauthorized notify to interactive card on telegram…
sam7894604 Jul 6, 2026
255674d
feat(line): dashboard authorized-list with names + admin lock + env o…
sam7894604 Jul 6, 2026
0b7983e
fix(line): pending 'approve' was a no-op for entries lacking source_type
sam7894604 Jul 6, 2026
0bf2cac
fix(line): dashboard COMMUNICATION RECORDS always empty (session_key …
sam7894604 Jul 6, 2026
b8bd0e6
fix(line): dashboard add/remove USER failed with 400 'unknown scope: …
sam7894604 Jul 6, 2026
b9ab3e3
refactor(line): dashboard — drop redundant Allowlist entry list
sam7894604 Jul 6, 2026
35a29f5
fix(line): interactive-card admin check rejected the notify recipient
sam7894604 Jul 6, 2026
ee20886
feat(line): store — card_admins table + managed settings get/set
sam7894604 Jul 6, 2026
d73a63f
feat(line): dashboard Settings panel (card_admins + config settings +…
sam7894604 Jul 6, 2026
f02b48f
fix(line): fail-open @mention gate when bot userId is unknown
sam7894604 Jul 9, 2026
fe54237
fix(line): route inbound audio/video/file to correct cache (not image…
sam7894604 Jul 10, 2026
3fdf3ef
feat(line+gateway): preserve PDF filename + auto-extract PDF text (vi…
sam7894604 Jul 11, 2026
99fd277
feat(gateway): generalize auto-extraction beyond PDF (text/csv/docx/x…
sam7894604 Jul 11, 2026
60ba98d
feat(gateway): full Office coverage — LibreOffice bridge for pptx + l…
sam7894604 Jul 12, 2026
f1582bd
fix(gateway): legacy spreadsheets via LibreOffice convert to XLSX, no…
sam7894604 Jul 12, 2026
cb0c398
feat(line): quote-reply = implicit mention + pre-extract observed media
sam7894604 Jul 13, 2026
5430f0c
feat(line): on-demand media backfill (replaces observe pre-extraction)
sam7894604 Jul 13, 2026
a5338a1
feat(line): backfill extracts+caches media, injects as channel_context
sam7894604 Jul 13, 2026
50bfc63
fix(agent): reliable turbovault edit_note + verifier false-alarm (B+C)
sam7894604 Jul 13, 2026
93d3840
fix(line): convert markdown tables to bullets on outbound
sam7894604 Jul 16, 2026
d8423ff
fix(state): get_conversation_root walks all parent links after upstre…
sam7894604 Jul 17, 2026
16ddf28
fix(tokens): include token_count in _CONVERSATION_ROW_COLUMNS after u…
sam7894604 Jul 19, 2026
8792681
fix(compression): inject memory-provider on_pre_compress() text into …
sam7894604 Jul 4, 2026
6cf37ef
test: expect provider_context kwarg in force-bypass compress call (PR…
sam7894604 Jul 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions agent/agent_runtime_helpers.py
Original file line number Diff line number Diff line change
Expand Up @@ -2155,6 +2155,38 @@ def switch_model(agent, new_model, new_provider, api_key='', base_url='', api_mo
reason="switch_model",
shared=True,
)

# ── Split session when model changes ──
_session_db = getattr(agent, '_session_db', None)
_old_sess_id = getattr(agent, 'session_id', None)
if _session_db is not None and _old_sess_id:
import uuid
_new_sess_id = f"{_old_sess_id.split('_')[0]}_{uuid.uuid4().hex[:8]}"
try:
_session_db.split_session(
_old_sess_id,
_new_sess_id,
model=new_model,
billing_provider=new_provider,
billing_base_url=agent.base_url,
billing_mode=getattr(agent, 'api_mode', None),
source=getattr(agent, 'platform', None),
user_id=getattr(agent, 'user_id', None),
cwd=getattr(agent, 'cwd', None),
)
agent.session_id = _new_sess_id
agent._transition_context_engine_session(
old_session_id=_old_sess_id,
new_session_id=_new_sess_id,
carry_over_context=True,
)
except Exception as _split_exc:
logger.warning(
"Session split on model switch failed (non-fatal): %s",
_split_exc,
)
# ──────────────────────────────────────────────

except Exception:
# Rollback every mutated field to the pre-swap snapshot so the agent
# is left consistent (old model + old provider + old client) and the
Expand Down
234 changes: 234 additions & 0 deletions agent/analytics.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,234 @@
"""Pure analytics computations over decoded token time-series.

DB-free so the math is unit-testable in isolation. Inputs come from
:meth:`hermes_state.SessionDB.get_message_token_timeseries` (1-minute
buckets) plus provider quota dicts from :mod:`agent.provider_quotas`.
"""
from __future__ import annotations

from typing import Any, Dict, List, Optional


def _pct(value: int, limit: Optional[int]) -> Optional[float]:
if not limit or limit <= 0:
return None
return round(100.0 * value / limit, 1)


def _stat(values: List[int]) -> Dict[str, float]:
"""min/max/mean/median/p95 over a list (0s for empty)."""
if not values:
return {"min": 0, "max": 0, "mean": 0.0, "median": 0.0, "p95": 0}
s = sorted(values)
n = len(s)
mean = sum(s) / n
median = s[n // 2] if n % 2 else (s[n // 2 - 1] + s[n // 2]) / 2
p95 = s[min(n - 1, int(round(0.95 * (n - 1))))]
return {"min": s[0], "max": s[-1], "mean": round(mean, 2), "median": round(median, 2), "p95": p95}


def compute_usage_rates(
minute_buckets: List[Dict[str, int]],
daily_totals: Dict[str, int],
provider_quotas: List[Dict[str, Any]],
) -> Dict[str, Any]:
"""RPM/TPM peaks + RPD/TPD totals, compared to provider limits.

``minute_buckets``: 1-minute time-series rows (requests/input/output/…).
``daily_totals``: aggregate over the last 24h ({requests,input,output,…}).
``provider_quotas``: quota dicts for providers active in the window.
"""
req = [b.get("requests", 0) for b in minute_buckets]
tpm = [b.get("input", 0) + b.get("output", 0) for b in minute_buckets]
tpm_in = [b.get("input", 0) for b in minute_buckets]
tpm_out = [b.get("output", 0) for b in minute_buckets]

peak_rpm = max(req, default=0)
peak_tpm = max(tpm, default=0)
current_rpm = req[-1] if req else 0
current_tpm = tpm[-1] if tpm else 0

rpd = int(daily_totals.get("requests", 0))
tpd = int(daily_totals.get("input", 0)) + int(daily_totals.get("output", 0))

providers_view: List[Dict[str, Any]] = []
for q in provider_quotas:
# Peak-vs-limit %, conservative (treats the global peak as if it all
# went to this provider — exact for the single-provider common case).
providers_view.append({
"provider": q.get("provider"),
"display": q.get("display"),
"tier": q.get("tier"),
"limits": {
"rpm": q.get("rpm"), "rpd": q.get("rpd"),
"tpm_input": q.get("tpm_input"), "tpm_output": q.get("tpm_output"),
"tpd": q.get("tpd"),
},
"pct_of_limit": {
"rpm": _pct(peak_rpm, q.get("rpm")),
"rpd": _pct(rpd, q.get("rpd")),
"tpm_input": _pct(max(tpm_in, default=0), q.get("tpm_input")),
"tpm_output": _pct(max(tpm_out, default=0), q.get("tpm_output")),
"tpd": _pct(tpd, q.get("tpd")),
},
"source_url": q.get("source_url"),
"as_of": q.get("as_of"),
})

return {
"rpm": {"current": current_rpm, "peak": peak_rpm},
"tpm": {
"current": current_tpm, "peak": peak_tpm,
"peak_input": max(tpm_in, default=0), "peak_output": max(tpm_out, default=0),
},
"rpd": rpd,
"tpd": tpd,
"window_totals": {
"requests": sum(req),
"input": sum(tpm_in),
"output": sum(tpm_out),
},
"providers": providers_view,
}


def compute_token_trends(buckets: List[Dict[str, int]]) -> Dict[str, Any]:
"""Per-bucket series + per-request averages + cache-hit rate.

``buckets``: time-series rows at the caller's chosen granularity.
"""
series: List[Dict[str, Any]] = []
per_call_input: List[int] = []
per_call_output: List[int] = []
total_input = total_output = total_cache = total_reasoning = total_req = 0

for b in buckets:
reqs = int(b.get("requests", 0))
inp = int(b.get("input", 0))
out = int(b.get("output", 0))
cache = int(b.get("cache_read", 0))
reason = int(b.get("reasoning", 0))
total_input += inp
total_output += out
total_cache += cache
total_reasoning += reason
total_req += reqs
cache_hit = round(100.0 * cache / inp, 1) if inp else None
series.append({
"bucket_start": int(b.get("bucket_start", 0)),
"requests": reqs,
"input": inp, "output": out, "cache_read": cache, "reasoning": reason,
"cache_hit_rate": cache_hit,
"avg_input_per_request": round(inp / reqs, 1) if reqs else 0,
"avg_output_per_request": round(out / reqs, 1) if reqs else 0,
})
if reqs:
per_call_input.append(round(inp / reqs))
per_call_output.append(round(out / reqs))

overall_cache_hit = round(100.0 * total_cache / total_input, 1) if total_input else None
return {
"series": series,
"totals": {
"requests": total_req, "input": total_input, "output": total_output,
"cache_read": total_cache, "reasoning": total_reasoning,
},
"averages_per_request": {
"input": round(total_input / total_req, 1) if total_req else 0,
"output": round(total_output / total_req, 1) if total_req else 0,
"reasoning": round(total_reasoning / total_req, 1) if total_req else 0,
"cache_read": round(total_cache / total_req, 1) if total_req else 0,
"input_distribution": _stat(per_call_input),
"output_distribution": _stat(per_call_output),
},
"cache_hit_rate": overall_cache_hit,
}


def _price(per_million: Any, tokens: int) -> Optional[float]:
"""Cost in USD for ``tokens`` at ``per_million`` USD/1M, or None if unpriced."""
if per_million is None:
return None
return round(float(per_million) * tokens / 1_000_000.0, 6)


def compute_cost_estimate(
groups: List[Dict[str, Any]],
window_seconds: int,
price_lookup,
) -> Dict[str, Any]:
"""Per-model cost broken down by price tier, with daily/monthly projection.

``groups``: rows from
:meth:`hermes_state.SessionDB.get_session_cost_aggregates`.
``price_lookup(model, provider, base_url)``: returns a pricing entry with
``input_cost_per_million`` / ``output_cost_per_million`` /
``cache_read_cost_per_million`` / ``cache_write_cost_per_million`` (or
None). When a model has no known pricing the stored ``estimated_cost_usd``
is used as a fallback and flagged.
"""
models: List[Dict[str, Any]] = []
total = 0.0
total_input_cost = total_output_cost = total_cache_cost = 0.0
any_unpriced = False

for g in groups:
entry = None
try:
entry = price_lookup(g.get("model"), g.get("billing_provider"), g.get("billing_base_url"))
except Exception:
entry = None

in_cost = out_cost = cr_cost = cw_cost = None
if entry is not None:
in_cost = _price(getattr(entry, "input_cost_per_million", None), g["input_tokens"])
out_cost = _price(getattr(entry, "output_cost_per_million", None), g["output_tokens"])
cr_cost = _price(getattr(entry, "cache_read_cost_per_million", None), g["cache_read_tokens"])
cw_cost = _price(getattr(entry, "cache_write_cost_per_million", None), g["cache_write_tokens"])

priced = any(c is not None for c in (in_cost, out_cost, cr_cost, cw_cost))
if priced:
grp_total = round(sum(c or 0.0 for c in (in_cost, out_cost, cr_cost, cw_cost)), 6)
source = "pricing"
else:
grp_total = round(float(g.get("estimated_cost_usd") or 0.0), 6)
source = "stored_estimate"
any_unpriced = True

total += grp_total
total_input_cost += in_cost or 0.0
total_output_cost += out_cost or 0.0
total_cache_cost += (cr_cost or 0.0) + (cw_cost or 0.0)

models.append({
"model": g.get("model"),
"provider": g.get("billing_provider"),
"sessions": g.get("sessions", 0),
"tokens": {
"input": g["input_tokens"], "output": g["output_tokens"],
"cache_read": g["cache_read_tokens"], "cache_write": g["cache_write_tokens"],
"reasoning": g["reasoning_tokens"],
},
"cost_breakdown": {
"input": in_cost, "output": out_cost,
"cache_read": cr_cost, "cache_write": cw_cost,
},
"cost_usd": grp_total,
"cost_source": source,
})

models.sort(key=lambda m: m["cost_usd"], reverse=True)
total = round(total, 6)
days = max(window_seconds / 86400.0, 1e-9)
daily = round(total / days, 6)
return {
"total_cost_usd": total,
"cost_by_tier": {
"input": round(total_input_cost, 6),
"output": round(total_output_cost, 6),
"cache": round(total_cache_cost, 6),
},
"projection": {"daily_usd": daily, "monthly_usd": round(daily * 30, 6)},
"has_unpriced_models": any_unpriced,
"models": models,
}
55 changes: 55 additions & 0 deletions agent/chat_completion_helpers.py
Original file line number Diff line number Diff line change
Expand Up @@ -1469,6 +1469,24 @@ def build_assistant_message(agent, assistant_message, finish_reason: str) -> dic
tool_calls.append(tc_dict)
msg["tool_calls"] = tool_calls

# Bit-pack (output, reasoning) token counts onto the assistant row.
# conversation_loop stashes the call's CanonicalUsage on agent._last_usage
# right after the response arrives (before this message is built/appended),
# so the value here is this turn's. Stored NEGATIVE per the codec's
# F=sign-bit convention; stripped from provider payloads in build_api_kwargs
# / the api_messages loop. Skipped when usage is unavailable (e.g. an
# interim message built before usage on a truncated stream).
_usage = getattr(agent, "_last_usage", None)
if _usage is not None:
from hermes_token_codec import pack_assistant_tokens, log_assistant_pack_verification
msg["token_count"] = pack_assistant_tokens(
getattr(_usage, "output_tokens", 0) or 0,
getattr(_usage, "reasoning_tokens", 0) or 0,
)
log_assistant_pack_verification(
getattr(agent, "session_id", None), msg["token_count"], _usage
)

return msg


Expand Down Expand Up @@ -1888,6 +1906,43 @@ def try_activate_fallback(agent, reason: "FailoverReason | None" = None) -> bool
# short-circuit the freshly activated fallback before it gets a
# single stream attempt.
_reset_stale_streak(agent)

# ── Split session when fallback changes model ──
# Only split if the model actually changed (provider key rotation
# with the same model should not create a new session row).
_old_m = (old_model or '').strip().lower()
_new_m = (fb_model or '').strip().lower()
if _old_m != _new_m:
_session_db = getattr(agent, '_session_db', None)
_old_sess_id = getattr(agent, 'session_id', None)
if _session_db is not None and _old_sess_id:
import uuid
_new_sess_id = f"{_old_sess_id.split('_')[0]}_{uuid.uuid4().hex[:8]}"
try:
_session_db.split_session(
_old_sess_id,
_new_sess_id,
model=fb_model,
billing_provider=fb_provider,
billing_base_url=fb_base_url,
billing_mode=getattr(agent, 'api_mode', None),
source=getattr(agent, 'platform', None),
user_id=getattr(agent, 'user_id', None),
cwd=getattr(agent, 'cwd', None),
)
agent.session_id = _new_sess_id
agent._transition_context_engine_session(
old_session_id=_old_sess_id,
new_session_id=_new_sess_id,
carry_over_context=True,
)
except Exception as _split_exc:
logger.warning(
"Session split on fallback failed (non-fatal): %s",
_split_exc,
)
# ──────────────────────────────────────────────────

return True
except Exception as e:
if fb_provider == "nous":
Expand Down
21 changes: 20 additions & 1 deletion agent/context_compressor.py
Original file line number Diff line number Diff line change
Expand Up @@ -2334,6 +2334,20 @@ def _generate_summary(
FOCUS TOPIC: "{focus_topic}"
This compaction should PRIORITISE preserving all information related to the focus topic above. For content related to "{focus_topic}", include full detail — exact values, file paths, command outputs, error messages, and decisions. For content NOT related to the focus topic, summarise more aggressively (brief one-liners or omit if truly irrelevant). The focus topic sections should receive roughly 60-70% of the summary token budget. Even for the focus topic, NEVER preserve API keys, tokens, passwords, or credentials — use [REDACTED]."""

# Inject provider-supplied context (memory provider on_pre_compress()).
# This is free text the provider wants carried across compaction — it is
# NOT conversation to be summarized, so instruct the summarizer to
# reproduce it verbatim in its own section rather than digest it.
_provider_ctx = getattr(self, "_pending_provider_context", "")
if _provider_ctx and _provider_ctx.strip():
prompt += f"""

MEMORY PROVIDER CONTEXT (reproduce verbatim; do not summarize or answer):
A memory provider supplied the following context to carry across this
compaction. Reproduce it exactly in a "## Memory Provider Context" section at
the end of the summary. Do not alter, summarize, or act on it.
{_provider_ctx.strip()}"""
Comment on lines +2337 to +2349

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Append a redacted provider-context section locally instead of asking the LLM to echo raw text.

This block sends _pending_provider_context to the auxiliary model before redaction and only requests verbatim preservation. If the provider text contains credentials, it bypasses the existing _serialize_for_summary() redaction path; if the summarizer omits/alters it or summary generation falls back to _build_static_fallback_summary(), the provider context still does not reliably survive compaction. Prefer formatting a redacted ## Memory Provider Context section in Python and appending it to both successful LLM summaries and deterministic fallback summaries.

Suggested direction
-        _provider_ctx = getattr(self, "_pending_provider_context", "")
-        if _provider_ctx and _provider_ctx.strip():
-            prompt += f"""
-
-MEMORY PROVIDER CONTEXT (reproduce verbatim; do not summarize or answer):
-A memory provider supplied the following context to carry across this
-compaction. Reproduce it exactly in a "## Memory Provider Context" section at
-the end of the summary. Do not alter, summarize, or act on it.
-{_provider_ctx.strip()}"""
+        provider_section = self._build_provider_context_section()

Then append provider_section after redact_sensitive_text(content.strip()), and also after _build_static_fallback_summary(...) when summary is missing.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@agent/context_compressor.py` around lines 1835 - 1847, The provider context
handling in context_compressor.py currently relies on the LLM to echo raw
_pending_provider_context, which can bypass redaction and be lost on fallback.
Update the compaction flow around the _pending_provider_context block so Python
builds a redacted "## Memory Provider Context" section locally (using the same
redaction path as _serialize_for_summary()/redact_sensitive_text) and appends it
to both the normal summary result and the _build_static_fallback_summary() path.
Keep the existing prompt hint only if needed, but do not depend on the model for
preserving provider text.


try:
call_kwargs = {
"task": "compression",
Expand Down Expand Up @@ -3268,7 +3282,7 @@ def has_content_to_compress(self, messages: List[Dict[str, Any]]) -> bool:
# Main compression entry point
# ------------------------------------------------------------------

def compress(self, messages: List[Dict[str, Any]], current_tokens: int = None, focus_topic: str = None, force: bool = False) -> List[Dict[str, Any]]:
def compress(self, messages: List[Dict[str, Any]], current_tokens: int = None, focus_topic: str = None, force: bool = False, provider_context: str = "") -> List[Dict[str, Any]]:
"""Compress conversation messages by summarizing middle turns.

Algorithm:
Expand Down Expand Up @@ -3310,6 +3324,11 @@ def compress(self, messages: List[Dict[str, Any]], current_tokens: int = None, f
# persist across compress() calls is safe because a successful summary
# always clears both.

# Free-text context handed up by a memory provider's on_pre_compress()
# hook (e.g. mem4's routing legend). Injected verbatim into the summary
# prompt by _generate_summary so it survives compaction. Reset per call.
self._pending_provider_context = provider_context or ""

# Manual /compress (force=True) bypasses the failure cooldown so the
# user can retry immediately after an auto-compress abort. Without
# this, /compress would silently no-op for 30-60s after a failure.
Expand Down
Loading