fix(batch_runner): write a discard tombstone so resume skips no-reasoning prompts - #93542
Closed
chelsealong wants to merge 1 commit into
Closed
chelsealong wants to merge 1 commit into
chelsealong wants to merge 1 commit into
Conversation
…ning prompts
The no-reasoning discard path marked prompts completed in the checkpoint
but never wrote a row to batch_N.jsonl. --resume's content-based scan
only reads those files, so any discarded prompt not yet flushed to the
checkpoint (e.g. the batch was interrupted before it fully returned) is
invisible to it and gets reprocessed at full cost, then discarded again.
Write a tombstone row ({"discarded": "no_reasoning", ...}) for each
discarded prompt so the scan sees it as done. The trajectory combine
step excludes these tombstones from trajectories.jsonl, and
statistics.json now reports discarded_no_reasoning.
Closed
14 tasks done
teknium1
pushed a commit
that referenced
this pull request
Aug 24, 2026
…ning prompts The no-reasoning discard branch in _process_batch_worker continued before writing any JSONL row, so run(resume=True) — which filters solely via _scan_completed_prompts_by_content over batch_*.jsonl — never saw discarded prompts and re-ran them at full cost on every resume. Write a tombstone row on discard, exclude tombstones from the trajectories.jsonl merge, and report discarded_no_reasoning in final statistics. Salvaged from #93542. Fixes #93527
teknium1
pushed a commit
that referenced
this pull request
Aug 24, 2026
…bstones Complete the #93527 fix: the tombstone now carries the human prompt text via _entry_prompt_text (handling flat prompt, ShareGPT, and chat-style shapes), _scan_completed_prompts_by_content counts discarded rows as completed instead of only reading ShareGPT conversations, and the merge step reports excluded tombstones in the combined-count summary. Adds a dedicated regression suite covering the tombstone round-trip, the all-discarded-batch resume path, and merge exclusion. Salvaged from #93579 (issue reporter's PR), building on #93542. Fixes #93527
Collaborator
and7777
pushed a commit
to and7777/hermes-agent
that referenced
this pull request
Aug 27, 2026
…ning prompts The no-reasoning discard branch in _process_batch_worker continued before writing any JSONL row, so run(resume=True) — which filters solely via _scan_completed_prompts_by_content over batch_*.jsonl — never saw discarded prompts and re-ran them at full cost on every resume. Write a tombstone row on discard, exclude tombstones from the trajectories.jsonl merge, and report discarded_no_reasoning in final statistics. Salvaged from NousResearch#93542. Fixes NousResearch#93527
and7777
pushed a commit
to and7777/hermes-agent
that referenced
this pull request
Aug 27, 2026
…bstones Complete the NousResearch#93527 fix: the tombstone now carries the human prompt text via _entry_prompt_text (handling flat prompt, ShareGPT, and chat-style shapes), _scan_completed_prompts_by_content counts discarded rows as completed instead of only reading ShareGPT conversations, and the merge step reports excluded tombstones in the combined-count summary. Adds a dedicated regression suite covering the tombstone round-trip, the all-discarded-batch resume path, and merge exclusion. Salvaged from NousResearch#93579 (issue reporter's PR), building on NousResearch#93542. Fixes NousResearch#93527
melon-xf
added a commit
to melon-xf/hermes-agent
that referenced
this pull request
Sep 3, 2026
…ning prompts The no-reasoning discard branch in _process_batch_worker continued before writing any JSONL row, so run(resume=True) — which filters solely via _scan_completed_prompts_by_content over batch_*.jsonl — never saw discarded prompts and re-ran them at full cost on every resume. Write a tombstone row on discard, exclude tombstones from the trajectories.jsonl merge, and report discarded_no_reasoning in final statistics. Salvaged from NousResearch#93542. Fixes NousResearch#93527
melon-xf
added a commit
to melon-xf/hermes-agent
that referenced
this pull request
Sep 3, 2026
…bstones Complete the NousResearch#93527 fix: the tombstone now carries the human prompt text via _entry_prompt_text (handling flat prompt, ShareGPT, and chat-style shapes), _scan_completed_prompts_by_content counts discarded rows as completed instead of only reading ShareGPT conversations, and the merge step reports excluded tombstones in the combined-count summary. Adds a dedicated regression suite covering the tombstone round-trip, the all-discarded-batch resume path, and merge exclusion. Salvaged from NousResearch#93579 (issue reporter's PR), building on NousResearch#93542. Fixes NousResearch#93527
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes #93527. The "no reasoning in any turn" discard path in
_process_batch_worker()marked the prompt as completed in thein-memory/checkpoint tracking but never wrote a row to
batch_N.jsonl.--resume's only dataset filter,_scan_completed_prompts_by_content(),scans exclusively those
batch_*.jsonlfiles for prompt text — so anydiscarded prompt whose batch hadn't yet returned from the worker pool
(e.g. the run was interrupted mid-batch, a very normal way for a long
batch job to end) is invisible to it. On
--resumeit gets reprocessedat full agent/API cost and discarded again, repeating indefinitely.
The fix: write a lightweight tombstone row for each discarded prompt
(
{"prompt_index": N, "conversations": ..., "discarded": "no_reasoning"})into the same
batch_N.jsonlfile used for real trajectories, so theexisting content-based scan naturally picks it up as already processed
(it only special-cases
failedrows for retry; anything else with ahuman-turn goes into the completed-prompts set). Two follow-on gaps
called out in the issue are fixed too:
trajectories.jsonlnowexplicitly skips these tombstones, so they never leak into the
training data.
statistics.jsonnow reportsdiscarded_no_reasoning(previouslyonly printed to stdout, so totals couldn't be reconciled after the
fact).
Related Issue
Fixes #93527
Type of Change
Changes Made
batch_runner.py:_process_batch_worker()writes a discard tombstonerow (with
fsync, matching the durability of the success path) insteadof writing nothing.
batch_runner.py: the batch-file combine step inBatchRunner.run()skips rows with a
discardedkey before they can reachtrajectories.jsonl.batch_runner.py:discarded_no_reasoningis aggregated across batchesand included in
statistics.json.tests/test_batch_runner_checkpoint.py: updated the existingtest_discarded_no_reasoning_prompts_are_marked_completed(whichasserted the old, buggy behavior — no row written) to assert the
tombstone is written instead, and added
test_resume_after_all_discarded_batch_reruns_zero_prompts, a directregression test for the issue: it drives
_process_batch_worker()todiscard a prompt, then drives
_scan_completed_prompts_by_content()and_filter_dataset_by_completed()exactly asrun()does, and assertsthe discarded prompt is (a) visible to the content scan and (b) not
rescheduled on resume.
How to Test
Verified the new/updated tests fail without the fix and pass with it, by
diffing out the
batch_runner.pychange and rerunning:Also ran, both clean:
Checklist
Code
fix(scope):,feat(scope):, etc.)pytest tests/ -q(the specific test files above) and all tests passDocumentation & Housekeeping
docs/, docstrings) — N/A, no public API/config surface changedcli-config.yaml.exampleif I added/changed config keys — N/ACONTRIBUTING.mdorAGENTS.mdif I changed architecture or workflows — N/Ascripts/check-windows-footguns.pyAI assistance disclosure
This PR was prepared with AI assistance (Claude), including the root-cause
analysis, the fix, and the tests. All changes were verified locally
(tests, ruff, windows-footguns check) before submission.