Skip to content

fix(agent): refresh DeepSeek pricing to the 2026-08-16 official rates - #94243

Open
Parker-Fawcett wants to merge 3 commits into
NousResearch:mainfrom
Parker-Fawcett:fix/deepseek-pricing-refresh-94221
Open

Parker-Fawcett wants to merge 3 commits into
NousResearch:mainfrom
Parker-Fawcett:fix/deepseek-pricing-refresh-94221

Conversation

@Parker-Fawcett

Copy link
Copy Markdown

Partial for #94221 (point 5 — the Python-side slice; the desktop Node/ekko layer lives outside this repo).

What & why

DeepSeek changed API pricing effective 2026-08-16 and added peak-hour
billing. The pricing table still carried the retired 2026-07 snapshot, so
every DeepSeek session's cost display under-reported real spend:

model old (table) new official
v4-pro 0.435 / 0.87 in-out 0.66 / 1.98 off-peak
v4-flash (+ deprecated chat/reasoner aliases) 0.14 / 0.28 0.22 / 0.66 off-peak
cache-hit input 0.003625 / 0.0028 0.022 / 0.007

Verified today against DeepSeek's own pricing page
(api-docs.deepseek.com/quick_start/pricing), not third-party trackers.

Peak-hour caveat, documented in-table: DeepSeek now bills
01:00–04:00 and 06:00–10:00 UTC Mon–Fri at exactly 2× the rates above.
PricingEntry has no time-of-day axis and estimate_usage_cost sees no
timestamp, so rows carry the OFF-PEAK rates — estimates are an off-peak
floor, never a peak quote. Modeling the split would be a schema change;
flagging it as possible follow-up rather than smuggling it into a data
refresh.

Also adds the now-listed deepseek-v4-flash-vision-exp row at flash
parity (image tokens bill as input tokens per the docs), preventing the
next "unknown cost" report for that model.

How to test

from agent.usage_pricing import get_pricing_entry
e = get_pricing_entry("deepseek-v4-pro", provider="deepseek")
float(e.input_cost_per_million)   # 0.66 (was 0.435)

Updated test_deepseek_v4_pro_pricing_entry_exists to the new rates;
added test_deepseek_vision_exp_prices_as_v4_flash parity invariant; the
existing alias invariant (chat/reasoner ≡ flash) passes untouched.

Suite: scripts/run_tests.sh tests/agent/test_usage_pricing.py tests/agent/test_billing_usage.py tests/agent/test_account_usage.py → 53 passed, 0 failed at head. macOS 26, Python 3.11.

Parker-Fawcett and others added 3 commits August 23, 2026 09:58
…eads

Under gateway multiplexing, get_secret fails closed on unscoped reads
(rather than risk returning another profile's credential). hindsight's
writer, daemon-start and prefetch threads are spawned raw — no contextvars
propagation — so local_embedded could never boot its daemon: _get_client's
get_secret('HINDSIGHT_LLM_API_KEY') raised UnscopedSecretError on every
start and retain.

Capture the profile HERMES_HOME at construction (always scoped) and wrap
each background body in a _profile_scope context manager that re-installs
set_secret_scope(build_profile_secret_scope(home)) plus the home override —
the same contract gateway/run.py applies to its own worker threads.
Per-job wrapping in the writer loop keeps .env edits visible to later
retains; sentinel exit is unaffected.

Closes NousResearch#92608.
Review findings on NousResearch#93028:

1. Cache the built profile scope at construction instead of re-parsing the
   profile .env on every writer job / prefetch recall; docstring documents
   the snapshot lifecycle trade-off.
2. Debug-log a misconstruction signal: multiplex active with no
   HERMES_HOME override at construction means the captured home is likely
   the process default.
3. Extract _spawn_embedded_daemon/_daemon_start_body and
   _prefetch_background from their closures so both wrapper paths are
   directly testable; add per-path regressions observing
   current_secret_scope()/get_hermes_home() inside each body.
4. Reuse _bare_provider across all scope tests; switch the multiplex
   activation to the public set_multiplex_active hook.
DeepSeek changed API pricing on 2026-08-16 and introduced peak-hour
billing (01:00-04:00 and 06:00-10:00 UTC Mon-Fri at exactly 2x). The
table carried the retired 2026-07 snapshot (pro 0.435/0.87, flash
0.14/0.28), so session cost displays under-reported real spend.

Refresh all deepseek rows to the current OFF-PEAK rates verified against
api-docs.deepseek.com/quick_start/pricing (pro 0.66/1.98 cache 0.022,
flash 0.22/0.66 cache 0.007) and document that the estimator has no
time-of-day axis: entries are an off-peak floor, not a peak quote.
Add the now-listed deepseek-v4-flash-vision-exp row at flash parity
(image tokens bill as input tokens per the docs).

Partial for NousResearch#94221 (its remaining scope lives in the Studio Node layer,
outside this repo).

Closes NousResearch#94221 point 5.
@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/plugins Plugin system and bundled plugins tool/memory Memory tool and memory providers provider/deepseek DeepSeek API area/billing Account usage, credit usage, billing (cross-cutting) area/usage-cost Token accounting, usage reporting, billing, cost tracking P2 Medium — degraded but workaround exists labels Aug 24, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

Overall: two independent fixes bundled (DeepSeek pricing refresh + hindsight background-thread secret scoping) — both look correct individually; consider splitting them in the PR description since they touch unrelated subsystems and reviewers/CI-bisect will treat them separately. Per-change notes:

  1. Off-peak-floor semantics deserve user-visible surfacing — agent/usage_pricing.py:509-513: documenting "estimates are an off-peak floor" in the table comment is right, but anyone reading hermes insights cost numbers during peak hours (01:00–04:00 / 06:00–10:00 UTC Mon–Fri) will see up-to-2× understated costs with no hint. If the CostResult path can carry a caveat string cheaply, surfacing "peak hours may bill ~2×" next to deepseek rows would prevent trust erosion in the estimator; otherwise at least mirror the caveat where estimates render.

  2. Misconstruction under multiplexing should warn, not debug-log — plugins/memory/hindsight/__init__.py:768-777: a provider constructed with no HERMES_HOME override while multiplexing doesn't just risk wrong credentials — self._profile_home also scopes the memory bank identity, so retains/recalls silently land in the process-default profile's bank (cross-profile memory contamination, harder to notice than a secret failure). That's a correctness hazard worth logger.warning, matching the fail-closed posture the rest of this change enforces.

  3. Minor: _daemon_start_body (hindsight/__init__.py:1860) leaves open(log_path, "a") handles owned by the Rich Console unclosed on every daemon start — pre-existing, but this PR touches exactly these lines; a with open(...) as f: around the console assignment (or an explicit close after daemon start settles) is a one-line cleanup.

  4. Nit: _prefetch_background re-reads self._prefetch_result fields after recall — unchanged from before, but note the thread-reference overwrite in queue_prefetch means two rapid queries race on self._prefetch_*; pre-existing, out of scope.

  5. Positive: the TestBackgroundSecretScope sync-thread fixture design (making threading.Thread synchronous and capturable) tests the actual scope-install behavior rather than mocking it away — nice.

@Parker-Fawcett

Copy link
Copy Markdown
Author

Thanks for the review — clarifying the grouping: this PR is pricing-only (DeepSeek rows + tests). The hindsight background-thread fixes you mention live on #93028 (and the debug-log vs warning question there was per Enough1122's original guidance for a debug-level diagnostic — warning would be noisy when multiplexing is off, which is the common case).

On point 1 (off-peak floor surfacing) — agreed the table comment alone won't reach users staring at hermes insights during peak. The rows here intentionally carry the off-peak floor with pricing_version deepseek-pricing-2026-08-offpeak so estimates never overstate the bill (peak is exactly 2×, per the docs). If you want a UI caveat (peak hours may bill ~2×) plumbed through CostResult, happy to add it in a follow-up — kept this PR data-only to stay focused and keep the diff reviewable. The estimator has no timestamp today, so any peak-aware math would be a schema/behavior change beyond a snapshot refresh.

@Enough1122

Copy link
Copy Markdown

Thanks for clarifying the scope split — understood that this PR is intentionally pricing-only and the hindsight background-thread fixes live on #93028. The off-peak floor rationale (estimates never overstate the bill, peak is exactly 2×, no timestamp in estimator today so peak-aware math would be a behavior change) makes sense to keep this data-only; a UI caveat via CostResult as a follow-up works. And the debug-log vs warning point being per the original guidance for a debug-level diagnostic when multiplexing is off — noted, warning would be noisy in the common case.

No further items from me on this PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/billing Account usage, credit usage, billing (cross-cutting) area/usage-cost Token accounting, usage reporting, billing, cost tracking comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/plugins Plugin system and bundled plugins P2 Medium — degraded but workaround exists provider/deepseek DeepSeek API tool/memory Memory tool and memory providers type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants