Skip to content

perf(agent): avoid quadratic work in tool history pairing - #105660

Closed
ymybxx wants to merge 1 commit into
NousResearch:mainfrom
ymybxx:perf/tool-history-pairing
Closed

ymybxx wants to merge 1 commit into
NousResearch:mainfrom
ymybxx:perf/tool-history-pairing

Conversation

@ymybxx

@ymybxx ymybxx commented Sep 8, 2026

Copy link
Copy Markdown

What does this PR do?

Long tool-heavy histories incur two avoidable quadratic costs during message pairing: positional repair copies the entire remaining history for each assistant call, and compression orphan classification scans result-ID sets again for each declared call. This removes both costs while preserving the existing pairing decisions.

Related Issue

Found by profiling current main (fef0e16fe19b79ded929209f87c7434270b03825); no linked issue. This preserves the positional pairing policy from #97203. I searched open and closed PRs by the helper names and quadratic/tool-pairing terms; the nearby #104454 changes persistence invalidation, not traversal cost.

Type of Change

  • Performance refactor (no behavior change)

Changes Made

  • agent/agent_runtime_helpers.py: walk the immediately following tool-result run by index, avoiding eager suffix copies. Stop at the same first non-tool message.
  • Reuse the existing union of result-ID variants to find missing calls. An orphan result is disjoint from all declared call IDs, so including it in that union cannot create a match. Alias handling, output ordering, object identity, and the separate positional/global pairing policies stay the same.
  • tests/agent/test_tool_pairing_scaling.py: two parametrized operation-count contracts (64 and 256 pairs), covering adjacency, composite aliases, missing calls, orphan results, and retained object identity. No wall-clock assertions.

How to Test

scripts/run_tests.sh tests/agent/test_tool_pairing_scaling.py tests/run_agent/test_message_sequence_repair.py tests/agent/test_cursor_optimizations_parity.py tests/agent/test_context_compressor.py tests/agent/test_send_path_history_isolation.py --file-retries 0 -j 4 -q --tb=short

218 passed. Copying just the new test file onto the base produces 4 failures; this branch passes all 4. At 256 pairs, the base performs 66,821 history reference reads/copies and 33,152 call/result intersections, exceeding the linear budgets.

A separate 1,000-case seeded differential check through the real repair and compression-cleanup callers also preserved repair counts, persistence cursors, serialized transcripts, and retained message identities. Ruff, git diff --check, and scripts/check_compat_pointers.py pass.

Performance

macOS 26.6.2 arm64, Python 3.11.15; synthetic histories, 5-run medians after warming existing argument-validation cursors. The small reproduction below was run on the base and this branch with the same venv.

Messages Local phase Base This PR
501 Per-iteration preparation 1.66 ms 1.53 ms
501 Compression pair cleanup 2.51 ms 0.75 ms
6,001 Per-iteration preparation 38.78 ms 17.64 ms
6,001 Compression pair cleanup 250.55 ms 9.02 ms
12,001 Per-iteration preparation 124.37 ms 35.47 ms
12,001 Compression pair cleanup 995.41 ms 19.09 ms

These measure local phases, excluding model inference, network latency, and summary generation. Short conversations see small absolute savings. Large compression-cleanup inputs are reachable when current-task retention keeps a long tool chain in the tail: with the default lean-tail settings and a 1M-token window, the real boundary/tail helpers retained 12,002 tail messages in a synthetic input estimated at 519,989 tokens. This PR does not change compression thresholds or retention.

Benchmark reproduction (run from each checkout root with its dev Python)

Run the following with python - or save it in the checkout root. The outer agent is a minimal stub; the iteration-preparation and compressor methods are real imports. No model request is made. JSON equality checks are outside the timed region.

import gc, json, os, statistics, tempfile, time
from types import SimpleNamespace

with tempfile.TemporaryDirectory(prefix="hermes-pairing-bench-") as bench_home:
    os.environ["HERMES_HOME"] = bench_home
    from agent.agent_runtime_helpers import sanitize_tool_call_arguments
    from agent.context_compressor import ContextCompressor
    from agent.turn_iteration_prep import prepare_iteration

    compressor = object.__new__(ContextCompressor)
    compressor.quiet_mode = True
    for count in (250, 3000, 6000):
        messages = [{"role": "user", "content": "Inspect the repository."}]
        for i in range(count):
            call = {
                "id": f"call_{i}", "call_id": f"logical_{i}",
                "response_item_id": f"fc_{i}", "type": "function",
                "function": {"name": "read_file", "arguments": "{}"},
            }
            messages.extend([
                {"role": "assistant", "content": "", "tool_calls": [call]},
                {"role": "tool", "tool_call_id": f"logical_{i}", "content": "ok"},
            ])
        before = json.dumps(messages)
        agent = SimpleNamespace(
            step_callback=None, _skill_nudge_interval=0, valid_tool_names=set(),
            _adopt_nous_key_before_expiry=lambda: None,
            _drain_pending_steer=lambda: None, session_id="pairing-benchmark",
            _sanitize_tool_call_arguments=sanitize_tool_call_arguments,
            budget_warning_ratio=None, _last_flushed_db_idx=0,
        )
        phases = {
            "prepare_iteration": lambda: prepare_iteration(
                agent, messages=messages, api_call_count=count,
            ).messages,
            "compression_pair_cleanup": lambda: compressor._sanitize_tool_pairs(messages),
        }
        for name, phase in phases.items():
            phase()
            samples = []
            for _ in range(5):
                gc.collect()
                start = time.perf_counter()
                result = phase()
                samples.append(1000 * (time.perf_counter() - start))
                assert json.dumps(result) == before
            print(len(messages), name, round(statistics.median(samples), 3), "ms")

Checklist

  • Read CONTRIBUTING.md and the root/agent development guides.
  • Conventional commit; one focused commit with production code and regression tests.
  • Searched existing open/closed PRs for duplicates.
  • Canonical isolated runner: 218 related tests pass, file retries disabled.
  • Full repository test suite (not run; focused coverage listed above).
  • Tested on macOS arm64; changes use host-independent Python list/set operations.
  • Documentation/config/tool schemas: N/A; no new settings, dependencies, or user-facing behavior.

@alt-glitch alt-glitch added type/perf Performance improvement or optimization P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint labels Sep 8, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

PR #105660 — avoid quadratic work in tool history pairing

Both optimizations check out. The slice→index change preserves the break-on-first-non-tool semantics while removing the O(n) copy per row, and the union reuse is provably equivalent (an orphan result is by definition disjoint from every declared call's variants, so it can never rescue a "missing" call — the comment states exactly this). The scaling tests assert bounds structurally (items read / intersections) rather than with wall-clock timing, so they won't flake on slow CI.

  • Non-blocking — no issues found. test_global_pairing_bounds_id_comparisons also re-verifies the returned sets against uninstrumented helpers, guarding against the optimization changing results, not just speed. Nice.

kshitijk4poor pushed a commit that referenced this pull request Sep 23, 2026
_classify_tool_call_orphans rebuilt the surviving-result variant list and
scanned all of it for every declared tool call (O(calls x results)).
Orphan result variants are disjoint from every declared call, so matching
against the already-computed union of all result variants is equivalent
and O(1) per call. 1000 pairs: 50.9 ms -> 2.4 ms per classification.

Partial salvage of #105660: kept the _classify_tool_call_orphans union
reuse; dropped the messages[i+1:] slice -> range() hunk (sub-ms) and
tests/agent/test_tool_pairing_scaling.py (read-count assertions are a
change-detector). Hunk re-derived by hand: PR base was ~7.9k commits behind.

(cherry picked from commit 6879a69)
@kshitijk4poor

Copy link
Copy Markdown

Thanks for this. It is included in #120406 as part of the perf sweep. The orphan-classification union reuse in _classify_tool_call_orphans is kept, re-applied by hand on current main with the original commit's author (illidan) kept. The messages[i+1:] slice hunk is left out.

Closing this in favour of #120406, which is armed to rebase-merge. If that PR doesn't land, I'll reopen this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have type/perf Performance improvement or optimization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants