fix(agent): refresh DeepSeek pricing to the 2026-08-16 official rates - #94243
Parker-Fawcett wants to merge 3 commits into
Conversation
…eads
Under gateway multiplexing, get_secret fails closed on unscoped reads
(rather than risk returning another profile's credential). hindsight's
writer, daemon-start and prefetch threads are spawned raw — no contextvars
propagation — so local_embedded could never boot its daemon: _get_client's
get_secret('HINDSIGHT_LLM_API_KEY') raised UnscopedSecretError on every
start and retain.
Capture the profile HERMES_HOME at construction (always scoped) and wrap
each background body in a _profile_scope context manager that re-installs
set_secret_scope(build_profile_secret_scope(home)) plus the home override —
the same contract gateway/run.py applies to its own worker threads.
Per-job wrapping in the writer loop keeps .env edits visible to later
retains; sentinel exit is unaffected.
Closes NousResearch#92608.
Review findings on NousResearch#93028: 1. Cache the built profile scope at construction instead of re-parsing the profile .env on every writer job / prefetch recall; docstring documents the snapshot lifecycle trade-off. 2. Debug-log a misconstruction signal: multiplex active with no HERMES_HOME override at construction means the captured home is likely the process default. 3. Extract _spawn_embedded_daemon/_daemon_start_body and _prefetch_background from their closures so both wrapper paths are directly testable; add per-path regressions observing current_secret_scope()/get_hermes_home() inside each body. 4. Reuse _bare_provider across all scope tests; switch the multiplex activation to the public set_multiplex_active hook.
DeepSeek changed API pricing on 2026-08-16 and introduced peak-hour billing (01:00-04:00 and 06:00-10:00 UTC Mon-Fri at exactly 2x). The table carried the retired 2026-07 snapshot (pro 0.435/0.87, flash 0.14/0.28), so session cost displays under-reported real spend. Refresh all deepseek rows to the current OFF-PEAK rates verified against api-docs.deepseek.com/quick_start/pricing (pro 0.66/1.98 cache 0.022, flash 0.22/0.66 cache 0.007) and document that the estimator has no time-of-day axis: entries are an off-peak floor, not a peak quote. Add the now-listed deepseek-v4-flash-vision-exp row at flash parity (image tokens bill as input tokens per the docs). Partial for NousResearch#94221 (its remaining scope lives in the Studio Node layer, outside this repo). Closes NousResearch#94221 point 5.
Overall: two independent fixes bundled (DeepSeek pricing refresh + hindsight background-thread secret scoping) — both look correct individually; consider splitting them in the PR description since they touch unrelated subsystems and reviewers/CI-bisect will treat them separately. Per-change notes:
|
|
Thanks for the review — clarifying the grouping: this PR is pricing-only (DeepSeek rows + tests). The hindsight background-thread fixes you mention live on #93028 (and the debug-log vs warning question there was per Enough1122's original guidance for a debug-level diagnostic — warning would be noisy when multiplexing is off, which is the common case). On point 1 (off-peak floor surfacing) — agreed the table comment alone won't reach users staring at |
|
Thanks for clarifying the scope split — understood that this PR is intentionally pricing-only and the hindsight background-thread fixes live on #93028. The off-peak floor rationale (estimates never overstate the bill, peak is exactly 2×, no timestamp in estimator today so peak-aware math would be a behavior change) makes sense to keep this data-only; a UI caveat via No further items from me on this PR. |
Partial for #94221 (point 5 — the Python-side slice; the desktop Node/ekko layer lives outside this repo).
What & why
DeepSeek changed API pricing effective 2026-08-16 and added peak-hour
billing. The pricing table still carried the retired 2026-07 snapshot, so
every DeepSeek session's cost display under-reported real spend:
Verified today against DeepSeek's own pricing page
(
api-docs.deepseek.com/quick_start/pricing), not third-party trackers.Peak-hour caveat, documented in-table: DeepSeek now bills
01:00–04:00 and 06:00–10:00 UTC Mon–Fri at exactly 2× the rates above.
PricingEntryhas no time-of-day axis andestimate_usage_costsees notimestamp, so rows carry the OFF-PEAK rates — estimates are an off-peak
floor, never a peak quote. Modeling the split would be a schema change;
flagging it as possible follow-up rather than smuggling it into a data
refresh.
Also adds the now-listed
deepseek-v4-flash-vision-exprow at flashparity (image tokens bill as input tokens per the docs), preventing the
next "unknown cost" report for that model.
How to test
Updated
test_deepseek_v4_pro_pricing_entry_existsto the new rates;added
test_deepseek_vision_exp_prices_as_v4_flashparity invariant; theexisting alias invariant (
chat/reasoner≡ flash) passes untouched.Suite:
scripts/run_tests.sh tests/agent/test_usage_pricing.py tests/agent/test_billing_usage.py tests/agent/test_account_usage.py→ 53 passed, 0 failed at head. macOS 26, Python 3.11.