Skip to content

fix(insights): count mixed tool usage across sessions - #9896

Closed
MonkeyLeeT wants to merge 1 commit into
NousResearch:mainfrom
MonkeyLeeT:codex/issue-9814-insights-undercount
Closed

MonkeyLeeT wants to merge 1 commit into
NousResearch:mainfrom
MonkeyLeeT:codex/issue-9814-insights-undercount

Conversation

@MonkeyLeeT

Copy link
Copy Markdown

What does this PR do?

Fixes insights tool-usage undercounting when one session records tools via tool_name rows and another records them only via assistant tool_calls. The merge now happens per (session_id, tool_name) so mixed datasets add across sessions without double-counting duplicate representations inside a single session.

Related Issue

Fixes #9814

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • agent/insights.py: aggregate tool usage by (session_id, tool_name) before collapsing to repo-wide totals.
  • tests/agent/test_insights.py: add regression coverage for one tool_name-only session plus one tool_calls-only session for the same tool.

How to Test

  1. Run uv run --extra dev python -m pytest tests/agent/test_insights.py -q.
  2. Create one session with a tool row containing tool_name="search_files".
  3. Create a second session with only assistant tool_calls for search_files and confirm insights reports a count of 2.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: macOS

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — or N/A
  • I've updated cli-config.yaml.example if I added/changed config keys — or N/A
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — or N/A
  • I've updated tool descriptions/schemas if I changed tool behavior — or N/A

Screenshots / Logs

  • uv run --extra dev python -m pytest tests/agent/test_insights.py -q → 55 passed
  • uv run --extra dev python -m pytest tests/ -q is not green in this environment because of unrelated existing failures and missing optional dependencies such as acp, fastapi, and faster_whisper

@MonkeyLeeT
MonkeyLeeT force-pushed the codex/issue-9814-insights-undercount branch 13 times, most recently from d1a86bc to f73b5fa Compare April 21, 2026 21:25
@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint labels Apr 21, 2026
@alt-glitch

Copy link
Copy Markdown

Related to #9990 — both fix the same insights tool-usage counting bug (#9814) via different approaches.

@MonkeyLeeT
MonkeyLeeT force-pushed the codex/issue-9814-insights-undercount branch 14 times, most recently from 3d6345b to 1e7b089 Compare April 26, 2026 06:16
@MonkeyLeeT
MonkeyLeeT force-pushed the codex/issue-9814-insights-undercount branch 6 times, most recently from 717252f to 430558c Compare April 27, 2026 22:36
@MonkeyLeeT
MonkeyLeeT force-pushed the codex/issue-9814-insights-undercount branch from 430558c to a3cb4e3 Compare April 28, 2026 16:44
@MonkeyLeeT MonkeyLeeT closed this Apr 28, 2026
teknium1 pushed a commit that referenced this pull request Sep 12, 2026
…urces

`_get_tool_usage()` merged `tool_name` rows and assistant `tool_calls` JSON with
a GLOBAL per-tool max. That is right inside one session (both columns describe
the same call) but wrong across sessions: a gateway session recording
`tool_name` only plus a CLI session recording `tool_calls` only for the same
tool reported 1 use instead of 2.

Group both queries by (session_id, tool_name), reconcile with max per session,
then sum across sessions.

Port of PR #9896 by @MonkeyLeeT onto the `_scoped` query layout; one invariant
test covering disjoint sessions AND a paired session.

Fixes #9814
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Insights undercount tool usage when tool_name and tool_calls sources coexist

2 participants