Skip to content

docs: how we improve tallyman from session reviews, the error corpus and the eval suite - #290

Open
paddymul wants to merge 1 commit into
mainfrom
docs/improving-from-sessions
Open

paddymul wants to merge 1 commit into
mainfrom
docs/improving-from-sessions

Conversation

@paddymul

Copy link
Copy Markdown
Contributor

Problem

Three ways of finding what to fix in tallyman have produced most of the recent PRs, and none of them is written down:

An agent asked to "look at session X" currently has to work all of this out again: where transcripts and errors.jsonl live, that tallyman errors are in-band (is_error is false), the rules Paddy gave along the way, and the traps.

What this adds

docs/improving-from-sessions.md, covering:

  • where the evidence is: transcript paths and format, and the per-project errors.jsonl, events.jsonl and telemetry.jsonl. It includes a short script that counts tool calls and prints tallyman errors; run on 4a93fb4e and b9d2057f it gives the same Bash counts the reviews reported (48 and 42).
  • how to run a driver session and ask for a review, and the steps a useful review takes.
  • a table mapping each kind of finding to a change: new tool, reply field, docstring, hint, lint, issue, upstream workaround, or nothing.
  • the error-corpus check against the hint code, and rules for writing a hint (anchored regexes, neutral column names, a negative test).
  • the eval suite: its files, how to run it, how to add a scenario from a transcript, and what to do with its findings.
  • the prompt-pack repo, a checklist for an agent, and a dated table of what is in flight.

It also adds a line to docs/architecture.md's doc index, and a short CLAUDE.md section telling agents to read the doc before reviewing sessions, touching hints or running evals.

Things the doc points out that need a decision

Checks

Docs only. The doc's anchors and relative links resolve against the headings in architecture.md and in the doc itself, and the embedded script was taken out of the markdown and run on two transcripts. CI was not watched.

🤖 Generated with Claude Code

…and the eval suite

docs/improving-from-sessions.md describes the three loops that have driven
recent fixes: a second Claude session reviewing a driver session's transcript
(#287, #288, #289 came from 4a93fb4e), running the hint code over every
recorded error, and replaying recorded sessions in the eval suite (#258,
whose findings became #259-#267). It says where transcripts and the
per-project error logs live, what each kind of finding turns into, and how to
add a hint or a scenario. architecture.md's doc index and CLAUDE.md point
to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant