Repository navigation
Conversation
…and the eval suite docs/improving-from-sessions.md describes the three loops that have driven recent fixes: a second Claude session reviewing a driver session's transcript (#287, #288, #289 came from 4a93fb4e), running the hint code over every recorded error, and replaying recorded sessions in the eval suite (#258, whose findings became #259-#267). It says where transcripts and the per-project error logs live, what each kind of finding turns into, and how to add a hint or a scenario. architecture.md's doc index and CLAUDE.md point to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Three ways of finding what to fix in tallyman have produced most of the recent PRs, and none of them is written down:
4a93fb4eled to feat(mcp): catalog_peek and catalog_query read rows without persisting anything #287 (catalog_query/catalog_peek), feat(mcp): a tool response carries errors recorded since the last one #288 (new_errors) and docs(mcp): a display klass formatter reference, and a hint when a build error has a known fix #289 (display-klass docs and hints).An agent asked to "look at session X" currently has to work all of this out again: where transcripts and
errors.jsonllive, that tallyman errors are in-band (is_erroris false), the rules Paddy gave along the way, and the traps.What this adds
docs/improving-from-sessions.md, covering:errors.jsonl,events.jsonlandtelemetry.jsonl. It includes a short script that counts tool calls and prints tallyman errors; run on4a93fb4eandb9d2057fit gives the same Bash counts the reviews reported (48 and 42).It also adds a line to
docs/architecture.md's doc index, and a shortCLAUDE.mdsection telling agents to read the doc before reviewing sessions, touching hints or running evals.Things the doc points out that need a decision
~/code/tallyman_nfl_demo/.mcp.jsonruns the MCP server from/Users/paddy/tallyman, which is at22f6a3e, 136 commits behindmain. Both NFL reviews partly described tools thatmainhad already replaced. Any NFL demo run started now tests that old checkout.docs/adr-007-009-implemented, which has since reachedmainthrough feat(cache): implement ADR-007, ADR-008 and ADR-009 (tallyman-owned materialization, row order, digest stability) #189, so it needs retargeting.~/code/tallyman2still has the untracked first version of the eval files (513c32a), which test: eval suite — multi-step sessions checked after every step, ported to ADR-011, with results on #216 #258 supersedes.Checks
Docs only. The doc's anchors and relative links resolve against the headings in
architecture.mdand in the doc itself, and the embedded script was taken out of the markdown and run on two transcripts. CI was not watched.🤖 Generated with Claude Code