Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions DEVJOURNAL.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,38 @@
# Development Journal

## 2026-03-12: Payload Visualization Feature

### Goal
Visualize what goes into each API payload: system prompt, tool definitions, user messages, assistant messages, tool results. Show actual cached tokens. Display as stacked bar (per-turn detail) and area chart (session growth over turns).

### Changes

**`run_agent.py` — instrumentation:**
- `_compute_payload_breakdown(api_kwargs)` — uses tiktoken (`gpt-4o` encoding) to count tokens per component. Handles both Codex Responses and Chat Completions API modes. ~3ms overhead.
- Breakdown computed after `_build_api_kwargs()` + preflight at main loop only (not memory flush or compression side tasks).
- `breakdown` dict merged into JSONL entry: `{"system": N, "tool_defs": N, "user": N, "assistant": N, "tool_results": N}`

**`dashboard/data.py` — endpoint:**
- `get_payload_breakdown(session_id)` — reads JSONL, filters by session, returns turns with breakdown. Old entries without breakdown gracefully skipped.

**`dashboard/server.py` — route:**
- `GET /api/payload-breakdown?session_id=<id>`

**`dashboard/static/index.html` — visualization:**
- New "Payload" tab with session selector
- Canvas area chart: stacked token layers (system blue, tool_defs purple, user green, assistant orange, tool_results red) with dashed cached-tokens overlay line
- Horizontal stacked bars: per-turn proportions, click any turn for detailed breakdown
- Per-turn drill-down: component bars with token counts, percentages, actual vs estimated vs cached
- Version display (v0.2.0) added to topbar

### Design decisions
- tiktoken for estimated counts based on JSON serialization (more accurate than the rough estimator in `model_metadata.py`, but not exact chat-ml counts)
- Extend existing JSONL rather than new file or DB table (old entries without breakdown handled gracefully)
- Instrument only main agent loop, not side tasks (memory flush, compression)
- `message_extras` extension table planned for reasoning content (avoids upstream schema conflicts on merge)

---

## 2026-03-09: OpenAI Codex Caching Improvements

### Problem
Expand Down
39 changes: 39 additions & 0 deletions dashboard/data.py
Original file line number Diff line number Diff line change
Expand Up @@ -222,6 +222,45 @@ def get_usage(days: int = 30, source: Optional[str] = None) -> Dict[str, Any]:
return result


def get_payload_breakdown(session_id: str) -> List[Dict[str, Any]]:
"""Get per-turn payload breakdown for a session from the JSONL token log.

Returns entries that have a 'breakdown' field, ordered by api_call number.
Entries without breakdown (older data) are skipped.
"""
if not JSONL_PATH.exists():
return []
results = []
try:
with open(JSONL_PATH, "r") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
entry = json.loads(line)
except json.JSONDecodeError:
continue
if entry.get("session_id") != session_id:
continue
if "breakdown" not in entry:
continue
results.append({
"api_call": entry.get("api_call", 0),
"prompt_tokens": entry.get("prompt_tokens", 0),
"completion_tokens": entry.get("completion_tokens", 0),
"cached_tokens": entry.get("cached_tokens", 0),
"cache_write_tokens": entry.get("cache_write_tokens", 0),
"breakdown": entry["breakdown"],
"tools": entry.get("tools", []),
"model": entry.get("model", ""),
})
except Exception:
return []
results.sort(key=lambda x: x["api_call"])
return results


def get_insights(days: int = 30, source: Optional[str] = None) -> Dict[str, Any]:
"""Generate insights report using InsightsEngine."""
if not DB_PATH.exists():
Expand Down
8 changes: 8 additions & 0 deletions dashboard/server.py
Original file line number Diff line number Diff line change
Expand Up @@ -112,6 +112,13 @@ async def handle_memory(request):
return json_response(data.get_memory())


async def handle_payload_breakdown(request):
session_id = request.query.get("session_id", "")
if not session_id:
return json_response({"error": "session_id required"}, status=400)
return json_response(data.get_payload_breakdown(session_id))


async def handle_insights(request):
days = int(request.query.get("days", 30))
source = request.query.get("source")
Expand Down Expand Up @@ -143,6 +150,7 @@ def create_app() -> web.Application:
app.router.add_get("/api/activity", handle_activity)
app.router.add_get("/api/memory", handle_memory)
app.router.add_get("/api/insights", handle_insights)
app.router.add_get("/api/payload-breakdown", handle_payload_breakdown)

# Static files
app.router.add_static("/static/", STATIC_DIR, name="static")
Expand Down
Loading