Skip to content

test: native provider E2E suites: Bedrock Converse, Gemini, Vertex, Codex app-server and Copilot ACP wire contracts proven against validating fakes - #121345

Merged
teknium1 merged 12 commits into
mainfrom
tests/e2e3-providers-native
Sep 24, 2026
Merged

teknium1 merged 12 commits into
mainfrom
tests/e2e3-providers-native

Conversation

@teknium1

@teknium1 teknium1 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Hermes' five native (non-chat-completions) provider dialects — Bedrock Converse, Gemini native, Vertex, the Codex app-server protocol and Copilot ACP — now have end-to-end wire-conformance suites that drive the real hermes chat -q CLI against fakes that validate every request the way the vendor would, with 18 real bugs pinned as strict xfails.

What is real vs faked

  • Real: the hermes CLI subprocess (hermetic fake HOME/HERMES_HOME, --resume in a new process), runtime provider resolution, the adapters, the agent loop, real tools (read_file, terminal), auto compaction, state.db.
  • Faked (vendor boundary only), in tests/fakes/providers/:
    • bedrock_converse.py: real boto3 reaches it via AWS_ENDPOINT_URL_BEDROCK_RUNTIME. It recomputes SigV4, validates bodies with the botocore bedrock-runtime Converse/ConverseStream ParamValidator plus the semantic rules (toolUse/toolResult pairing, reasoning signatures), and streams real vnd.amazon.eventstream frames with CRCs.
    • gemini_native.py: a loopback HTTPS CONNECT proxy that terminates TLS for generativelanguage.googleapis.com with a throwaway CA (HTTPS_PROXY + SSL_CERT_FILE). The native adapter only engages for the real host, and no Hermes code is patched. It validates against the published v1beta/v1 shapes and returns Google's 400 INVALID_ARGUMENT body.
    • vertex.py: a Google OAuth /token endpoint. Real google-auth signs an RS256 JWT from a generated service-account key and the fake verifies it. It also runs the same TLS proxy for *-aiplatform.googleapis.com, which validates the project/location path, the minted bearer, tool pairing and Gemini 3 thought signatures.
    • codex_app_server.py: a fake codex app-server executable speaking newline-delimited JSON-RPC. Its schema subset and serde-style errors come from codex app-server generate-json-schema and an offline round trip with the real binary. It validates Hermes' replies to approval requests and persists thread rollouts, so resume works across processes.
    • copilot_acp.py: a fake ACP agent executable. It answers --help with --acp, validates every request against the acp.schema pydantic models and records a timestamped JSONL transcript with PIDs.
  • Every fake also has a "fake rejects what the vendor rejects" test, so a pass-through fake can't make the suite vacuous.

Scenarios (test | what it proves | sabotage → observed red)

Every passing test was proven red by at least one temporary production edit, then reverted. The table has 56 rows, and several rows record more than one sabotage. Two were re-run by the lane parent on the merged branch: Gemini _tool_call_extra_signature returns None, and Bedrock toolResult toolUseId + "-x". Both went red as recorded.

Dialect Test Proves Sabotage → observed red
Bedrock Converse turns::test_every_call_is_sigv4_signed_for_bedrock_with_the_profile_credentials The real boto3 client is reached through AWS_ENDPOINT_URL_BEDROCK_RUNTIME. The fake recomputes and checks each SigV4 signature (AWS4-HMAC-SHA256, Credential=<.env key>//us-east-1/bedrock/aws4_request). bedrock_adapter bind_bedrock_runtime: _bedrock_region = 'us-west-2' → AssertionError: credential scope region 'us-west-2'
Bedrock Converse turns::test_parallel_tool_uses_round_trip_as_tool_results_paired_by_id Parallel toolUse blocks come back as toolResult blocks paired by toolUseId. The signed reasoningContent is replayed before the toolUse in-process. Every request passes the botocore Converse/ConverseStream ParamValidator and the fake's semantic checks. (a) toolUseId + '-x' in convert_messages_to_converse; (b) reasoning signature reversed on replay → (a) ValidationException: Expected toolResult blocks at messages.2.content for the following Ids; (b) messages.1.content.0: reasoning block signature is missing or invalid (6 tests red for each)
Bedrock Converse turns::test_tool_turn_prints_final_answer_and_persists_paired_rows The final eventstream text is printed exactly once, and the state.db assistant, tool and tool_call rows are paired. stream_converse_with_callbacks: text = delta['text'].upper() → exit=0, but the final answer was missing from stdout
Bedrock Converse turns::test_resume_in_new_process_replays_tool_history_valid_for_converse --resume in a new process rebuilds toolUse/toolResult history that Converse accepts. hermes_state_messages conversation replay: tool_calls never deserialized → The resumed request lacks the toolUse block
Bedrock Converse turns::test_compaction_in_reasoning_session_keeps_every_request_converse_valid Auto-compaction in a signed-reasoning tool session: the aux summary goes over Converse, and every post-compaction request carries the summary and passes validation (no orphan pairs, no unsigned final thinking). auxiliary_client Bedrock adapter sends messages[:0] → ValidationException: A conversation must start with a user message
Bedrock Converse turns::test_compaction_persists_summary_and_archives_compacted_rows The summary is persisted, compacted rows are archived, and no live tool row is orphaned. All compressor pair guards disabled (_sanitize_tool_pairs, _align_boundary_forward/backward, repair_message_sequence) → orphaned live tool rows: {'tooluse_001704075'}
Bedrock Converse turns::test_fake_rejects_requests_bedrock_rejects[7 cases] The fake is not a pass-through. Using real boto3 against it: orphan toolResult, tampered signature, unsigned final thinking, a union with two members and blank text give ValidationException; a wrong secret gives InvalidSignatureException; a valid request gives OK. fake validate_conversation made a no-op (sabotage of the fake, not production) → 4 of the 5 semantic cases red. The union case is still caught by ParamValidator and the bad-secret case by SigV4.
Bedrock Converse faults::test_retryable_fault_is_retried_once_into_the_answer[throttle_http,unavailable_http,throttle_in_stream] 429 ThrottlingException, 503 ServiceUnavailableException and a mid-text :message-type exception throttlingException frame are each retried by Hermes (botocore retries off via AWS_MAX_ATTEMPTS=1). The fake sees exactly 2 requests with identical history. The answer is printed once, with one assistant row. conversation_loop max_retries forced to 1 → exit=1: 'AWS Bedrock rate-limited every one of 1 attempts'
Bedrock Converse faults::test_mid_stream_drop_retries_without_duplicated_persisted_content A TCP drop 3 text deltas into the answer is retried. The partial text is not persisted or duplicated. max_retries=1, and separately _V_UNKNOWN non-retryable → exit=1 with 1 request
Bedrock Converse faults::test_streaming_iam_denial_falls_back_to_converse An AccessDenied on ConverseStream for bedrock:InvokeModelWithResponseStream falls back to unary Converse with the same history. Reasoning and text are persisted. is_streaming_access_denied_error returns False → exit=1
Gemini native test_native_gemini_tools.py::test_function_call_round_trip_pairs_response_and_persists A real read_file functionCall comes back as a functionResponse with the same name and id, carrying the seeded canary. The functionCall part carries the signature the fake issued, byte for byte. The CLI prints the answer, and state.db has the tool_call, a matching tool row and the answer once. agent/gemini_native_adapter.py::_translate_tool_result_to_gemini: function_response id = tool_call_id + '-x' → The fake returned 400: 'functionResponse ids in contents[2] do not match the functionCall ids of the function call turn.'
Gemini native test_native_gemini_tools.py::test_resume_replays_thought_signatures_verbatim After --resume in a new process, turn 1's functionCall is sent back with its exact issued thoughtSignature (not the skip-validator dummy) and stays paired with its result. The resumed turn's own call is also signed. hermes_state_messages.py::get_messages_as_conversation: pop extra_content from loaded tool_calls → assert 'skip_thought_signature_validator' == '' for fc-0001
Gemini native test_native_gemini_tools.py::test_every_request_authenticates_and_targets_the_configured_model Every generate call sends x-goog-api-key equal to the .env key, to /v1beta/models/gemini-3-flash-preview, and no other Google path is requested. agent/gemini_native_adapter.py::GeminiNativeClient._headers: rename x-goog-api-key header → AssertionError: generate call without the configured API key: {... 'x-goog-api-kee': ...}
Gemini native test_native_gemini_schema.py::test_v1beta_json_schema_declaration_is_sanitized A tool from a real stdio MCP server with a hostile schema ($schema, $ref/$defs, oneOf, const, list type, integer enum, required naming no property) is declared with parametersJsonSchema. The declaration has no $ref, $schema or $defs; the $ref-typed parameter is inlined to enum fast/slow; required lists only real properties. The MCP call round-trips. agent/gemini_schema.py::prepare_gemini_tool_parameters: skip _inline_refs → AssertionError: reference/meta keywords reached Google: ['$ref']
Gemini native test_native_gemini_schema.py::test_v1_proto_schema_declaration_is_accepted With model.base_url pinned to /v1, the same hostile MCP schema goes out as legacy parameters that pass Google's field-by-field proto Schema checks, and the MCP call round-trips. agent/gemini_schema.py::sanitize_gemini_schema: disable the _GEMINI_SCHEMA_ALLOWED_KEYS filter → The fake returned 400: 'Invalid JSON payload received. Unknown name "$schema" at tools[0].function_declarations[20].parameters: Cannot find field.'
Gemini native test_native_gemini_errors.py::test_retryable_error_is_retried_then_succeeds[rate_limited_429] Two 429 RESOURCE_EXHAUSTED responses (with RetryInfo) are retried and the third call succeeds: the fake served statuses 429, 429, 200, and the answer is printed and persisted once. agent/error_classifier.py::_STATUS_HANDLERS[429] -> non-retryable format_error → Statuses were only [429]; the CLI exited 1 with a 'rejected this request as malformed' message.
Gemini native test_native_gemini_errors.py::test_retryable_error_is_retried_then_succeeds[unavailable_503] A 503 UNAVAILABLE is retried and the next call succeeds (statuses 503, 200); the answer is persisted once. agent/error_classifier.py::_STATUS_HANDLERS[503] -> non-retryable format_error → Statuses were only [503]; the CLI exited 1.
Gemini native test_native_gemini_errors.py::test_non_retryable_surfaced_once[invalid_argument_400] A 400 INVALID_ARGUMENT gets exactly one request. It is shown once with Google's message, the exit code is non-zero, and no scripted follow-up text is shown or persisted. agent/error_classifier.py::_STATUS_HANDLERS[400] -> _V_OVERLOADED → 'terminal Google response was re-sent 2x'; the forbidden scripted text was printed.
Gemini native test_native_gemini_errors.py::test_non_retryable_surfaced_once[safety] A candidate with finishReason SAFETY gets one request and is shown once as a safety block; nothing false is persisted as the answer. agent/gemini_native_adapter.py::_FINISH_REASON_MAP SAFETY -> 'stop' → 'terminal Google response was re-sent 2x'; the CLI printed the second scripted reply (exit 0).
Gemini native test_native_gemini_errors.py::test_non_retryable_surfaced_once[recitation] A candidate with finishReason RECITATION gets one request and is shown once; it is not retried and no fake success is persisted. agent/gemini_native_adapter.py::_FINISH_REASON_MAP RECITATION -> 'stop' → 'terminal Google response was re-sent 2x'; the CLI printed the second scripted reply (exit 0).
Gemini native test_native_gemini_errors.py::test_stream_drop_recovers_without_duplicate_rows When the TLS connection is reset mid-SSE after a partial chunk, Hermes retries once and the full answer is printed once. The partial text is never persisted and there are no duplicate assistant rows. agent/gemini_native_adapter.py::GeminiNativeClient._stream_completion: swallow httpx.HTTPError (return) → Assistant rows were ['GEMINI-PARTIAL-DROPPED', 'GEMINI-FULL-AFTER-DROP'] instead of only the full answer.
Gemini native test_native_gemini_compaction.py::test_compaction_happens_on_the_wire A large promptTokenCount triggers compaction mid-turn in a resumed reasoning session. A summary request with no tools reaches Google, and the next request is shorter than it would be without compaction and contains the summary. agent/context_compressor.py::should_compress_info: always return (False, None) → AssertionError: no compaction summary request reached Google
Gemini native test_native_gemini_compaction.py::test_post_compaction_requests_are_valid_and_signed No request is rejected before or after compaction (role alternation, call/response pairing by count, name and id, current-turn signatures). Every functionCall left after compaction carries its own issued signature verbatim, and there is no orphan functionResponse. agent/context_compressor.py::_finalize_compressed: strip extra_content from compacted tool_calls → 'signature for fc-0001 lost/altered': got 'skip_thought_signature_validator'
Gemini native test_native_gemini_compaction.py::test_persisted_transcript_keeps_pairs_and_signatures After compaction, every active tool row in state.db has a matching assistant tool_call, and the final answer is persisted once. agent/context_compressor.py::_finalize_compressed: pop tool_calls from the first assistant messages (makes orphan tool results) → AssertionError: orphan tool row fc-0001-...
Vertex test_native_vertex_tools.py::test_tool_result_goes_back_paired_to_the_signed_call A real read_file result goes back as a tool message paired to the Vertex call id. The thought signature (extra_content.google.thought_signature) is replayed byte-for-byte. The CLI prints the final answer. state.db has user/assistant(tool_calls)/tool/assistant rows. agent/transports/chat_completions.py::_model_consumes_thought_signature -> return False; separately agent/tool_dispatch_helpers.py tool message tool_call_id -> 'call_sabotaged' → sig strip: Vertex 400 'Function call is missing a thought_signature ... default_api:read_file'. tool_call_id: fake 400 function-response/call count mismatch; test red.
Vertex test_native_vertex_tools.py::test_wire_scheme_bearer_and_single_token_mint Every request goes to us-central1-aiplatform.googleapis.com/v1beta1/projects/<config project, not the SA-embedded one>/locations/us-central1/endpoints/openapi/chat/completions. The bearer was minted by the fake OAuth endpoint from a real google-auth RS256 JWT (signature, iss, aud, scope, iat/exp verified). One mint per process, reused across API calls. Model is sent in google/ form. agent/vertex_adapter.py::get_vertex_credentials ignores the project override; separately _needs_refresh -> return True → project: fake 404 on the SA-embedded project path, every turn fails. always-refresh: token-exchange count >1 per process, assertion red.
Vertex test_native_vertex_tools.py::test_reasoning_effort_reaches_vertex_as_thinking_config --reasoning high reaches the wire as the documented literal extra_body.google.thinking_config (thinking_level=high) and never together with reasoning_effort. plugins/model-providers/vertex/init.py::VertexProfile.build_extra_body -> return {} → thinking_config missing, assertion red
Vertex test_native_vertex_tools.py::test_thought_signature_replayed_after_resume_in_new_process After --resume in a NEW process, the first request replays turn 1's signed call (same id, same signature bytes) loaded from state.db. All issued signatures are on the wire in order, and Vertex accepts the resumed turn. hermes_state_messages.py::get_messages_as_conversation strips extra_content from deserialized tool_calls → replayed extra_content != issued signature, assertion red
Vertex test_native_vertex_tools.py::test_expired_token_is_reminted_and_request_retried The vendor expires the bearer mid-turn. The 401 UNAUTHENTICATED triggers a real JWT re-exchange, and the same request is retried once with the new token (statuses 200,401,200). No 401 text reaches the user and no rows are duplicated. agent/turn_recovery.py vertex 401 branch condition -> False → refresh turn exits 1 with the UNAUTHENTICATED error
Vertex test_native_vertex_errors.py::test_transient_error_is_retried_then_answers[rate_limited|unavailable] 429 RESOURCE_EXHAUSTED x2 (list-wrapped google.rpc body) and 503 UNAVAILABLE are retried with an identical body and the same token, then the answer is printed. agent/error_classifier.py::_STATUS_HANDLERS 429/503 -> _V_FORMAT_ERROR → both params red: single request, exit 1
Vertex test_native_vertex_errors.py::test_terminal_error_surfaced_once_without_retry[invalid_argument|permission_denied] 400 INVALID_ARGUMENT and 403 PERMISSION_DENIED: exactly one request, non-zero exit, and the vendor message shown exactly once. agent/error_classifier.py::_STATUS_HANDLERS 400/403 -> _V_SERVER_ERROR → both params red: retried (>1 request)
Vertex test_native_vertex_errors.py::test_rejected_bearer_refreshes_once_then_surfaces When every bearer 401s, there is at most one refresh+retry, then the error is shown once, with no loop. agent/turn_recovery.py remove '_retry.vertex_auth_retry_attempted = True' → refresh loop, the CLI never exits, run_chat TimeoutExpired -> module fixture ERROR
Vertex test_native_vertex_errors.py::test_oauth_invalid_grant_sends_nothing_and_names_the_credential invalid_grant at the Google token endpoint (with a valid JWT): zero Vertex requests, non-zero exit, and the output names VERTEX_CREDENTIALS_PATH / GOOGLE_APPLICATION_CREDENTIALS. hermes_cli/runtime_provider.py::_resolve_vertex_runtime substitutes a placeholder token instead of raising → 'requests reached Vertex without a token: [Bearer unminted]'
Vertex test_native_vertex_recovery.py::test_compacted_signed_session_stays_valid_for_vertex Auto-compaction (tiny threshold_tokens, large reported prompt tokens) in a --reasoning high signed session across --resume: the summary call goes to Vertex with the minted bearer and later requests carry the summary. Every post-compaction request is accepted (tool pairs intact, current-turn calls signed) and each signature stays on its own call id. agent/context_compressor.py::_finalize_compressed rebuilds tool_calls as {id,type,function}, dropping sidecars → turn 2 fails with Vertex 400 missing thought_signature (position 2)
Vertex test_native_vertex_recovery.py::test_compacted_session_rows_keep_tool_pairs state.db active rows after compaction: every tool_call id has its tool row, and the final answer is persisted once. same compressor tool_calls rebuild → turn 2 failed, test red
Vertex test_native_vertex_recovery.py::test_stream_drop_retried_without_duplicate_content After a signed tool step, the SSE body is cut mid-chunk (incomplete chunked TLS response). Hermes resends the identical valid request (signature intact), prints the recovered answer, and does not persist or print the partial. The answer row is stored once. agent/chat_completion_helpers.py::_handle_stream_error undelivered branch no longer resets delivery tracking → 'the retry changed the conversation' (a continuation user message was injected)
Codex app-server test_native_codex_app_server.py::test_every_request_is_schema_valid_and_handshake_ordered In every app-server process, initialize comes first, then the initialized notification. Every request and response Hermes sends matches the codex-cli 0.147 schema, with no fields codex would silently drop. agent/transports/codex_app_server.py::CodexAppServerClient.initialize: removed clientInfo.version → protocol violations: ... 'violation': 'missing field version'
Codex app-server test_native_codex_app_server.py::test_command_approval_round_trip_and_tool_rows The commandExecution approval reply reuses the server request id and says decision=accept. Hermes persists the exec_command tool call and a tool row paired by id with the output, then one final-answer row. The CLI prints the answer once. agent/transports/codex_app_server_session.py::_SERVER_REQUEST handler table: exec approval now returns {'decision':'approved'} → AssertionError: {... 'result': {'decision': 'approved'}} (reply != accept)
Codex app-server test_native_codex_app_server.py::test_reasoning_projected_once_and_never_as_content The reasoning summary lands on exactly one assistant row (the tool-call row), including after --resume. It never appears in content and is not printed in quiet mode. agent/transports/codex_event_projector.py::_assistant_message: stopped clearing _pending_reasoning → reasoning must attach to exactly one row (after resume too)
Codex app-server test_native_codex_app_server.py::test_resume_in_new_process_resumes_the_same_thread --resume in a new process sends thread/resume with turn 1's thread id and no thread/start. It does not re-send the history seed. turn/start goes to the same thread with the new input, and turn-2 rows are appended once. agent/codex_runtime.py::_stored_codex_thread_id: return None → resume must not start a fresh thread
Codex app-server test_native_codex_app_server.py::test_resume_fallback_seeds_history_then_rebinds_new_thread When the rollout is missing (thread/resume returns -32600), Hermes falls back to thread/start with a history seed. The seed holds the prior user turn, tool result and answer once each, but no reasoning and not the new prompt. Run 3 resumes the replacement thread. agent/codex_runtime_history_seed.py::render_history_seed: rows = [] → fallback thread/start must carry the prior conversation
Codex app-server test_native_codex_app_server.py::test_native_compaction_keeps_thread_and_transcript A contextCompaction item from codex mid-turn does not retire the thread: the next process resumes the same id. Hermes sends no thread/compact/start, the session is not split, and transcript rows are not rewritten or duplicated. agent/codex_runtime.py::_record_codex_app_server_compaction: added _store_codex_thread_id(agent, None) → codex-native compaction must not retire the thread
Codex app-server test_native_codex_app_server_faults.py::test_crash_mid_item_surfaced_once_without_partial_content_or_orphan When the app-server exits mid-item, 'exited unexpectedly' plus its stderr tail are shown once. Partial deltas are neither printed nor persisted, there is no respawn or replay (one turn/start), and the app-server PID is dead after the CLI exits. agent/transports/codex_app_server_session.py::_format_error_with_stderr: joined = '' → the app-server's stderr tail must reach the user once
Codex app-server test_native_codex_app_server_faults.py::test_terminal_turn_error_surfaced_once_not_retried[failed] turn/completed status=failed with an error is shown once. It is not retried, not persisted as assistant text, and the exit code is non-zero. codex_app_server_session.py::_run_started_turn.on_note: failed-status check replaced with if False: → exit=0 ... (marker count != 1)
Codex app-server test_native_codex_app_server_faults.py::test_terminal_turn_error_surfaced_once_not_retried[start_error] A JSON-RPC -32600 error on turn/start is shown once with the server message, not retried, and the exit code is non-zero. codex_app_server_session.py::_request_for: CodexAppServerError now sets a bare label with no message → exit=1 ... START-ERR-99 count != 1
Codex app-server test_native_codex_app_server_faults.py::test_will_retry_error_notification_is_not_terminal An error notification with willRetry=true, followed by success, does not end the turn: exit 0, final answer printed and persisted once, the notice is not shown, one turn/start. codex_app_server_session.py::on_note: treat method=='error' as terminal → exit=1 wall=6.0s (turn aborted on retry notice)
Copilot ACP test_native_copilot_acp.py::test_hermes_tool_call_round_trips_through_acp_and_persists Across three ACP agent processes, the agent emits a text tool call, Hermes runs its real read_file on a seeded canary file, and the result comes back in the next session/prompt after the user turn. The CLI prints the final answer. In state.db, the assistant row's tool_call id matches the tool row's tool_call_id (the id the fake chose). agent_thought_chunk text is kept as reasoning. No assistant text is duplicated. agent/acp_openai_bridge.py::extract_tool_calls_from_text returns [] ; agent/copilot_acp_client.py sinks drops agent_thought_chunk → 'the read_file result must follow the user turn in the next call's prompt' ; 'agent_thought_chunk text was not kept as reasoning'
Copilot ACP test_native_copilot_acp.py::test_every_request_is_schema_valid_and_selects_the_configured_model Every client request passes validation against the acp.schema pydantic models, with unknown root fields rejected and cwd required to be absolute. Each process sends initialize, then session/new (cwd = project), then session/set_config_option (the configured model), then session/prompt, using the issued sessionId. _model_selection_request returns None ; session/new sends relative cwd + extra 'model' field → method sequence ['initialize','session/new','session/prompt'] ; the fake rejects session/new with -32602 so the CLI exits 1
Copilot ACP test_native_copilot_acp.py::test_agent_permission_is_never_granted_and_fs_reads_stay_in_cwd Hermes answers the agent's session/request_permission with outcome cancelled (fails closed) and never grants it. fs/read_text_file works for a file inside the session cwd and returns an error for a secret file outside it, without leaking the content. permission response set to selected/allow-once ; _ensure_path_within_cwd relative_to check removed → 'Hermes granted an agent-side permission' ; 'fs read escaped the session cwd: ... OUTSIDE-CWD-SECRET-5512'
Copilot ACP test_native_copilot_acp.py::test_resume_reaches_the_agent_with_persisted_history_in_order --resume runs in a new agent process whose prompt carries the prior user, tool and assistant turns in order, then the new question. The answer is printed and persisted once in the same session. How the ACP session is opened (seeded session/new or session/load) is deliberately not asserted. _format_messages_as_prompt drops assistant messages → 'resumed prompt lost or reordered history: [39663, 39739, -1, 39858]'
Copilot ACP test_native_copilot_acp.py::test_no_agent_process_outlives_the_cli Every agent PID the fake recorded is gone after the CLI exits. This includes an agent wedged so it ignores both SIGTERM and stdin EOF, so it proves the SIGKILL fallback runs. _terminate_process kill() fallback removed → 'timed out after 10.0s waiting for agent processes [...] to exit' (the fixture teardown reaped the orphans)
Copilot ACP test_native_copilot_acp.py::test_compaction_in_an_acp_session_keeps_the_next_prompt_valid_and_grounded Eight reads push the turn over compression.threshold_tokens and auto compaction fires mid-turn. The summarizer call goes to the ACP agent itself and receives the history. The next main prompt passes schema validation and contains the summary, the user's request and the latest tool result, while the summarized tool output is gone. The final answer is persisted once. agent/context_compressor.py should_compress_info always returns False → 'compaction never called the summarizer through the ACP provider'
Copilot ACP test_native_copilot_acp_errors.py::test_acp_failure_is_retried_per_semantics_and_surfaced_once[internal_error_retried|crash_mid_stream_retried|auth_required_not_retried|invalid_params_bounded|crash_every_time_bounded] Transient errors (-32603, or the agent crashing mid-stream) are retried and recover: the partial chunk is never persisted and the final answer is persisted once. -32000 Authentication required gets exactly 1 attempt. Persistent -32602 or repeated crashes stop at api_max_retries and the error is shown once. No agent PID is left behind in any case. agent_init _api_retries=1 ; _request returns {} when the process died (crash becomes a partial success) ; error_classifier 'authentication' pattern removed → 4 rows: 'N: 1 model calls, expected 2' ; crash rows: '1 model calls, expected 2' ; auth row: '2 model calls, expected 1'
Copilot ACP test_native_copilot_acp_errors.py::test_auth_failure_remedy_is_an_actionable_command (strict xfail #121290) Any hermes ... fix command suggested in the auth-failure message must actually run, not report 'not implemented'. fix-simulation: turn_failure_copy oauth template changed to suggest copilot login → XPASS(strict) under the fix simulation; xfails for the KnownSymptom on main
Copilot ACP test_native_copilot_acp.py::test_message_chunk_after_prompt_result_reaches_the_user (strict xfail #65788) An agent_message_chunk sent after the session/prompt result should still reach the user. fix-simulation: after the prompt result, _request keeps reading and handling messages for 0.6s → XPASS(strict) under the fix simulation; on main the chunk is dropped, the turn looks empty and Hermes retries
Copilot ACP test_native_copilot_acp_streaming.py::test_long_reasoning_turn_streams_progress_before_it_completes[tui_stream|nested_acp] (strict xfail #120550, #101507) During a turn of about 4.5s of thought chunks, progress must reach the surface before the prompt result, timed against the fake's own clock. Checked on two surfaces: tui_gateway reasoning.delta/message.delta events, and session/update from hermes acp running on top of copilot-acp. fix-simulation: drop the acp:// no-stream guard in agent/turn_api_call.py and stream chunks live from CopilotACPClient → both XPASS(strict) under the fix simulation; on main no chunk arrives until the result (turn_api_call disables streaming for acp:// and the client buffers the turn)

KNOWN bugs (strict xfail with raises=KnownSymptom, per-file KNOWN dict; KnownSymptom is raised only at the bug's symptom)

Dialect Issue Symptom
Bedrock Converse #121293 (new) After --resume, Bedrock assistant turns are replayed without their signed reasoningContent. The bedrock_content_blocks sidecar is never persisted, so the wire request differs from the in-process one. xfail test: test_resume_replays_signed_reasoning_verbatim.
Bedrock Converse #121294 (new) A 400 ValidationException (and a non-streaming 403 AccessDenied) is retried and reported as 'temporarily unavailable'. botocore ClientError has no status_code, so the classifier returns unknown/retryable. Confirmed with classify_api_error. xfail test: test_validation_exception_is_surfaced_once_without_retry. It flips to XPASS→red when unknown is made non-retryable.
Bedrock Converse #98468 (existing) Streamed Bedrock reasoning is persisted with '\n\n' between every delta. xfail test: test_persisted_reasoning_equals_the_streamed_reasoning_text.
Bedrock Converse #109988 (existing) A ConverseStream that cleanly ends before messageStop is accepted: the truncated 'Recovered answer ZEBR' is persisted as the final answer, exit 0, no retry. xfail test: test_stream_ending_before_message_stop_is_not_accepted.
Gemini native #121317 (new) When Google blocks the prompt itself (promptFeedback.blockReason, no candidates), the adapter treats it as an empty stream. It is sent 9 times (3 stream retries × 3 attempts) and the user is told Google AI Studio is 'temporarily unavailable'; the block reason is never shown. Covered by strict xfail test_non_retryable_surfaced_once[prompt_blocked].
Gemini native #99438 (existing) On the legacy v1 parameters path, $ref/$defs are dropped instead of inlined, so the $ref-typed MCP parameter is sent as an empty schema {}. v1beta parametersJsonSchema inlines it correctly. Covered by strict xfail test_v1_ref_parameter_keeps_its_shape.
Gemini native #71804 (existing) On v1, an array parameter without items is sent as-is and Google returns 400 'parameters.properties[tags].value.items: missing field.' Covered by strict xfail test_v1_array_without_items_is_accepted.
Vertex #109115 (existing) terminal.notify anyOf[boolean, array] is rejected by Vertex's FunctionDeclaration translator ('For schema with items, schema type should be ARRAY'), so every default-toolset turn 400s. strict xfail test_default_toolset_schemas_accepted_by_vertex.
Vertex #121295 (new) Vertex 401 UNAUTHENTICATED / 403 PERMISSION_DENIED is reported as 'rejected your API key ... Update it in Settings → Providers'. Vertex has no API key (auth_kind falls through to api_key). Also, the 401 refresh resends the identical cached bearer. strict xfail test_auth_failure_guidance_is_vertex_specific[permission_denied|unauthenticated].
Codex app-server #121296 (new) In hermes chat -q, a codex commandExecution approval waits on the interactive prompt for the full approvals.timeout (300 s by default) and then declines. It ignores approvals.single_query_mode, and the turn loop is blocked the whole time. xfail: reply arrived 8.0 s later with timeout=8.
Codex app-server #121297 (new) The reply to item/permissions/requestApproval is {'decision':'decline'}. The schema requires permissions and has no decision field, so the fake flags 'missing field permissions'.
Codex app-server #121298 (new) The chat -q exit path never calls agent.close()/_close_codex_session. The app-server dies only when its stdin hits EOF, so close()'s descendant reap never runs and children in their own session (like MCP servers) are orphaned. Related to closed #66671.
Codex app-server #121299 (new) If a turn fails after an agentMessage completed, the CLI prints the message as the answer, exits 1 and never shows the failure reason. A failure-only turn does show it.
Codex app-server #121301 (new) A codex-native contextCompaction item is persisted as an assistant row '[codex contextCompaction] {json}' through the projector's opaque fallback.
Copilot ACP #121290 (new) When the agent returns ACP -32000 'Authentication required', Hermes suggests hermes auth add copilot-acp --type oauth, which exits 'not implemented for auth type oauth yet'. The user is left with a fix command that doesn't work.
Copilot ACP #65788 (existing) An agent_message_chunk sent after the session/prompt result is dropped. The turn looks empty, so Hermes retries it with a new agent process.
Copilot ACP #120550 (existing) copilot-acp turns are not streamed: turn_api_call disables streaming for acp:// and the client buffers the whole turn, so the TUI/Desktop gets no reasoning or message delta while the turn is running.
Copilot ACP #101507 (existing) With hermes acp running on top of copilot-acp, the outer client gets the inner chunks only after the inner turn has ended.

Local runtime

scripts/run_tests.sh --include-integration tests/e2e/core/providers/: 14 files, 67 passed + 19 strict xfail, 0 failed. (The earlier "20" counted a codex teardown error that the unguarded xfail swallowed; see Review fixes.)

Run Runner wall Slowest file
1 28.0 s 28.0 s
2 31.7 s —
3 27.0 s —
4 27.2 s —
5 27.8 s 27.6 s
loaded (2 copies in parallel) 39.7 s / 39.2 s, both green —

Every file finishes in 30 s or less, well under the CI 180 s per-file budget. ruff, check_no_tmp_literals, check-windows-footguns --all and git diff --check are all clean. No production code changed.

Review fixes

Rebased on current origin/main (no XPASS: all 19 KNOWN cells still reproduce).

Finding Change Before → after
MINOR 1: strict xfails without raises= let harness errors pass as the known bug Every KNOWN xfail in all 11 xfail-bearing files (Bedrock, Gemini, Codex, Vertex, Copilot) now uses raises=KnownSymptom, a shared AssertionError subclass in _native_helpers.py. It is raised only at the bug's symptom. Preconditions (exit codes, request counts, the approval exchange, the descendant pid) are plain asserts, next(...)/tuple unpacks became length-checked lists, and wait_until(..., error=KnownSymptom) covers the orphan timeout. --runxfail: all 19 cells fail with KnownSymptom. Sabotage: q_approval scenario issues no command (no approval request). On PR head it XFAILED (StopIteration swallowed); now it FAILS with approval request/reply not exchanged once: [] []. The codex faults teardown error that PR head counted as a 5th xfail now shows as 1e (seen mid-fix, before the MINOR 2 fix).
MINOR 2: orphaned grandchild (#121298) killed with os.kill from module teardown, no guard bypass; sleep(600) leaked on dev boxes The fake's own-session grandchild now waits on a release file; CodexRun.cleanup() drops that file and waits for it to exit, with a SIGKILL fallback. Both codex modules also carry @pytest.mark.live_system_guard_bypass (same convention as tests/agent/test_shell_hooks_tree_kill.py). The dev-box guard re-arms before module-fixture finalizers run, so the marker alone did not help: the cooperative release is what fixes the leak. PR head under run_tests.sh: 4✓ 5xf plus one leaked sleep(600) (ppid 1) per run. Now 4✓ 4xf, 0 leaked across 5 full-suite runs.
MINOR 3: copilot resume test asserted session/load is never sent, which is not documented anywhere Dropped that assertion and renamed the test. The new-agent-process, in-order history and no-duplicate assertions stay, and the module docstring was corrected. The existing _format_messages_as_prompt sabotage still turns it red (history-order assert).
NIT: body said 20 strict xfails Now 19 (the 20th was the swallowed teardown error). —
(extra) Copilot teardowns os.kill after the liveness check now tolerates ProcessLookupError (race), so teardown can't error spuriously. —

Determinism after the fixes: scripts/run_tests.sh --include-integration tests/e2e/core/providers/ 3× serial (46.6 s / 34.4 s / 37.4 s), plus 2 copies in parallel (72.0 s / 66.8 s). Every run: 67 passed, 19 xfailed, 0 failed, 0 errors, 0 leaked processes. The CI test_agent_turn_liveness[provider_hang] red is a pre-existing main flake owned by another lane and is not touched here.

NOT COVERED (honest gaps)

  • Bedrock Converse: Non-streaming Converse as the primary path (only covered through the IAM streaming-denial fallback)
  • Bedrock Converse: The AnthropicBedrock SDK route (bedrock anthropic.claude-* via the anthropic_messages transport); only the Converse dialect is covered
  • Bedrock Converse: Guardrail / safety-block stopReason (guardrail_intervened, content_filtered) and ModelErrorException/ModelTimeoutException
  • Bedrock Converse: ServiceQuotaExceededException and the credential-pool rotation/fallback chain on throttle
  • Bedrock Converse: Image/document content blocks and cachePoint blocks
  • Gemini native: Unary generateContent (non-streaming): no CLI path used it (main turns and compaction summaries both streamed). The fake still serves and validates it.
  • Gemini native: Text-part thoughtSignature replay: Google documents text-part signatures as not validated, and the adapter does not replay them. Not asserted.
  • Gemini native: Parallel function calls (signature only on the first part) in a single step: the fake supports them, but no scenario scripts one.
  • Gemini native: Checks on parametersJsonSchema contents beyond an object root: Google does not publish which JSON-Schema keywords it rejects there, so the fake does not enforce any.
  • Gemini native: MAX_TOKENS / OTHER finish reasons and 500 INTERNAL are not scripted.
  • Vertex: Vertex thought-summary text output (include_thoughts): no published OpenAI-compat wire shape to emit without guessing; related open issues [Bug]: Gemini/Vertex thought summaries leak into user-visible content with display.show_reasoning: false (effort and visibility coupled via include_thoughts) #64035 and [Bug]: Vertex Gemini exposes interim assistant messages during tool calls #76997
  • Vertex: The partial-visible-delivery stream-drop path (continuation stub): -Q oneshot has no display consumer, so drops always take the undelivered same-prefix retry path
  • Vertex: ADC / gcloud user credentials and the metadata server: only the service-account JWT flow is faked (NO_GCE_CHECK is set)
  • Vertex: global and us/eu multi-region hosts (Vertex us/eu multi-region locations build invalid API hosts #72063): only regional us-central1 is exercised
  • Vertex: Aux-client-only failure modes ([Bug]: Every auxiliary task silently fails on provider: vertex #61852, Auxiliary client cannot use Vertex AI — the auth_type="vertex" ADC branch is unreachable (vertex missing from PROVIDER_REGISTRY) #66674): compaction's summary call to Vertex works here, but other aux tasks are disabled
  • Vertex: Separate sabotage of the compressor's tool-pair guards: disabling _align_boundary_backward + _sanitize_tool_pairs together stayed green in this scenario (no split group arises, or a later wire-repair layer covers it). The compaction tests are proven via signature loss instead.
  • Codex app-server: Codex app-server approvals bypass pre/post approval hooks #108368 approval hooks: needs an interactive approval prompt (PTY); in -q the approval path is broken by codex_app_server: approval requests in 'hermes chat -q' wait the full approvals.timeout instead of resolving from single_query_mode #121296 and --yolo never prompts
  • Codex app-server: [Bug]: codex-runtime progress does not refresh the 600s liveness watchdog #118410 liveness: agent.turn_liveness.timeout_s=2 never fired in chat -q, even on a turn silent for 6 s, so this path cannot be made to fail deterministically
  • Codex app-server: codex_app_server turn_timeout is a hardcoded 600s with no config surface #118486 turn timeout: hardcoded 600 s with no config key; a test would have to invent the key
  • Codex app-server: Hermes-driven thread/compact/start (compression.codex_app_server_auto: hermes): preflight compaction runs before the codex session exists in a one-turn -q process, so it cannot be reached through chat -q (the fake supports the method)
  • Codex app-server: fileChange approval and mcpServer elicitation round trips: the fake validates their response schemas but no scenario drives them
  • Codex app-server: Retryable-then-succeed from Hermes' side: codex retries internally and Hermes does not retry codex turns by design; covered only as 'willRetry notice is not terminal'
  • Copilot ACP: Real Copilot CLI binary: only the fake agent is exercised
  • Copilot ACP: session/cancel on user interrupt mid-turn
  • Copilot ACP: fs/write_text_file permission and cwd bounds (only reads are tested)
  • Copilot ACP: ACP-native tool_call/tool_call_update rendering in the UI: the fake emits them, but Hermes has no contract for showing them
  • Copilot ACP: Quota/billing error text from the real Copilot CLI: only JSON-RPC codes are modelled
  • All: real vendor endpoints and binaries are never contacted (no paid calls). Fidelity is limited to the published contracts and the offline schema dumps the fakes encode.

Infographic

native-provider-wire-proof

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on f0f5212 — test: gate the ACP crash-retry count on #121467 (merge-order

debug info

CI timings

CI timings · View report · View job

Wall time 7m26s vs 9m17s (-19.9%). 5 job(s) slower, 8 faster, 1 unchanged.

  • OS-specific tests / Windows-only tests: +97.0s
  • OS-specific tests / Windows E2E (real processes): +66.0s
  • Python tests / Run tests: -50.0s
  • Python lints / Windows footguns (blocking): +33.0s
  • Detect affected areas: -22.0s

@arkheioncorp arkheioncorp left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer verdict: COMMENT

Summary: Native provider E2E suites for Bedrock Converse, Gemini, Vertex, Codex app-server, and Copilot ACP wire contracts. Adds shared harness _native_helpers.py and fault/turn test suites.

Findings:

  • The harness design is solid: hermetic fake HOME, env allowlisting (no credential passthrough via _SECRET_ENV_SUFFIXES + _PASSTHROUGH_ENV), real subprocesses, real SQLite reads.
  • HERMES_STATE_DB_GUARD_BYPASS=1 is documented as a test escape hatch — acceptable for E2E.
  • No security concerns in the diff: no hardcoded secrets, no shell injection vectors, no eval/exec.
  • Test design: fault semantics properly parametrize retryable vs terminal faults; xfail markers reference issue numbers for known bugs.
  • The adopt_foreign_durable_rows function in the compression PR (121340) is referenced conceptually here — good integration awareness.

Suggestions (non-blocking):

  • Consider adding a brief note in the harness docstring about what the fakes must NOT do (e.g., must not write to real HOME, must not touch network).
  • The make_home function writes .env with AWS creds; ensure the fake creds in tests/fakes/providers/bedrock_converse.py are clearly non-real placeholders (they appear to be consts like ACCESS_KEY/SECRET_KEY — if those are fake values, fine).

Verdict: COMMENT — design looks sound, no blockers. Ready for merge after CI passes.

@arkheioncorp arkheioncorp left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer verdict: COMMENT

Summary: Native provider E2E suites for Bedrock Converse, Gemini, Vertex, Codex app-server, and Copilot ACP wire contracts. Adds shared harness _native_helpers.py and fault/turn test suites.

Findings:

  • The harness design is solid: hermetic fake HOME, env allowlisting (no credential passthrough via _SECRET_ENV_SUFFIXES + _PASSTHROUGH_ENV), real subprocesses, real SQLite reads.
  • HERMES_STATE_DB_GUARD_BYPASS=1 is documented as a test escape hatch — acceptable for E2E.
  • No security concerns in the diff: no hardcoded secrets, no shell injection vectors, no eval/exec.
  • Test design: fault semantics properly parametrize retryable vs terminal faults; xfail markers reference issue numbers for known bugs.

Suggestions (non-blocking):

  • Consider adding a brief note in the harness docstring about what the fakes must NOT do (e.g., must not write to real HOME, must not touch network).
  • The make_home function writes .env with AWS creds; ensure the fake creds in tests/fakes/providers/bedrock_converse.py are clearly non-real placeholders.

Verdict: COMMENT — design looks sound, no blockers. Ready for merge after CI passes.

@alt-glitch alt-glitch added type/test Test coverage or test infrastructure P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/bedrock AWS Bedrock (boto3, IAM) provider/gemini Google Gemini (AI Studio, Cloud Code) provider/copilot GitHub Copilot (ACP + Chat) provider/openai OpenAI / Codex Responses API labels Sep 24, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 24, 2026
…V4-verified eventstream fake, tools, signed reasoning, resume, compaction, faults)
… Google TLS endpoint + tools/resume, schema, errors, compaction)
… fake, tools, signatures, errors, compaction, stream drop)
…ow, resume, lifecycle, error semantics

Fake ACP agent executable (tests/fakes/providers/copilot_acp.py) validates every client request
against the agent-client-protocol schema and replays scripted turns. Suites drive real hermes chat -q:
tool round trip via the prompt tool bridge, model selection, fail-closed permissions, cwd-bounded fs,
resume = fresh session/new seeded with history, no orphaned agents (SIGTERM-ignoring agent), and
JSON-RPC error / crash retry semantics. Strict xfails: #65788 (late chunk), #121290 (auth remedy).
…(TUI events + nested hermes acp)

Strict xfails for #120550 (tui_gateway gets no reasoning/message delta until the ACP turn ends) and
#101507 (hermes acp over copilot-acp forwards inner chunks only after the inner turn). Timestamps are
compared against the fake agent's own clock; both xfails were proven to XPASS under a local
streaming fix simulation (reverted).
…valid and grounded

Summarizer runs through the ACP agent itself; the post-compaction prompt must carry the summary, the
user's ask and the latest tool result, drop the summarized tool output, and pass schema validation.
…; codex orphan released without a foreign kill

Review fixes for #121345.

- Every strict xfail in tests/e2e/core/providers now uses raises=KnownSymptom (a shared
  AssertionError subclass in _native_helpers) and raises it only at the tracked bug's
  symptom. Preconditions (exit codes, request counts, the approval exchange, the
  descendant pid) are plain asserts, so harness failures - StopIteration from next(...),
  tuple-unpack errors, process death, fixture teardown errors - fail for real instead of
  counting as the known bug. pytest applies xfail to setup/teardown reports too; the
  codex faults file's "5 xfailed" for 4 xfail tests was the teardown error being eaten.
  Verified with --runxfail: all 19 KNOWN cells fail with KnownSymptom.
- Codex fake: the own-session grandchild (#121298) now lives until CodexRun.cleanup()
  drops a release file, so teardown retires the orphan without signalling a process that
  was reparented to init (the live-system guard blocked that kill and leaked a
  sleep(600) per run on dev boxes). Both codex modules carry live_system_guard_bypass
  for the SIGKILL fallback, matching tests/agent/test_shell_hooks_tree_kill.py.
- Copilot ACP resume test no longer asserts session/load is never sent (not a documented
  contract); it keeps the new-agent-process, in-order history and no-duplicate asserts.
- Copilot teardowns tolerate a pid exiting between the liveness check and the kill.
@teknium1
teknium1 force-pushed the tests/e2e3-providers-native branch from 656d938 to bcd2dae Compare September 24, 2026 11:11
…ner's reach

CI's tirith build flags arithmetic expansion (echo MARK-$((6*7))) as HIGH
'Nested expansion' and single-query mode blocks it, so the terminal leg of the
parallel toolUse round trip never ran. The test is about toolUse/toolResult
pairing, not the scanner: use a plain pipeline that still computes the marker
(seq | sed) so a fake echoing the command text cannot satisfy MARK-42.
CI caught crash_every_time_bounded making 4 model calls against a 2-retry budget:
when the CLI's stderr lags its exit the client raises TimeoutError instead of
'exited early', and timeouts retry on another budget. Deterministic with a lagged
stderr pump (5 calls, 126 s). Fix is #121468; until it lands the cell XFAILs on
exactly that signature, passes once fixed, and any other failure stays red.
@teknium1
teknium1 merged commit 9a6108f into main Sep 24, 2026
36 checks passed
teknium1 added a commit that referenced this pull request Sep 24, 2026
…; codex orphan released without a foreign kill

Review fixes for #121345.

- Every strict xfail in tests/e2e/core/providers now uses raises=KnownSymptom (a shared
  AssertionError subclass in _native_helpers) and raises it only at the tracked bug's
  symptom. Preconditions (exit codes, request counts, the approval exchange, the
  descendant pid) are plain asserts, so harness failures - StopIteration from next(...),
  tuple-unpack errors, process death, fixture teardown errors - fail for real instead of
  counting as the known bug. pytest applies xfail to setup/teardown reports too; the
  codex faults file's "5 xfailed" for 4 xfail tests was the teardown error being eaten.
  Verified with --runxfail: all 19 KNOWN cells fail with KnownSymptom.
- Codex fake: the own-session grandchild (#121298) now lives until CodexRun.cleanup()
  drops a release file, so teardown retires the orphan without signalling a process that
  was reparented to init (the live-system guard blocked that kill and leaked a
  sleep(600) per run on dev boxes). Both codex modules carry live_system_guard_bypass
  for the SIGKILL fallback, matching tests/agent/test_shell_hooks_tree_kill.py.
- Copilot ACP resume test no longer asserts session/load is never sent (not a documented
  contract); it keeps the new-agent-process, in-order history and no-duplicate asserts.
- Copilot teardowns tolerate a pid exiting between the liveness check and the kill.
@teknium1
teknium1 deleted the tests/e2e3-providers-native branch September 24, 2026 12:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-reviewed applied to manually approve dangerous changes comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have provider/bedrock AWS Bedrock (boto3, IAM) provider/copilot GitHub Copilot (ACP + Chat) provider/gemini Google Gemini (AI Studio, Cloud Code) provider/openai OpenAI / Codex Responses API type/test Test coverage or test infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: session.close leaks the Codex app-server process tree — AIAgent.close never closes _codex_session

3 participants