Repository navigation
docs(OMN-12432): add merge-queue lint-cancellation diagnosis - #1798
jonahgabriel wants to merge 5 commits into
Conversation
#1739) * fix(OMN-11994): handler_wiring psycopg2 list adaptation for text[] columns (#1730) * fix(OMN-11994): handler_wiring psycopg2 list adaptation for text[] columns Only apply psycopg2.extras.Json() to dict values, not list values. Lists intended for Postgres text[] columns (models_used, machines_used) were being wrapped in Json() producing malformed array literals. * test(OMN-11994): cover psycopg2 text array adaptation * chore(OMN-11994): rerun gates after OCC merge * style(OMN-11994): format handler wiring integration test --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11958): remove duplicate dispatcher for node_architecture_graph_query_effect (#1732) * fix(OMN-11958): deduplicate contract names in discover_contracts to prevent DUPLICATE_REGISTRATION crash When two installed packages (e.g. omnibase_infra and omnimarket) register onex.nodes entry points whose contract.yaml files declare the same `name` field, the auto-wiring engine previously built a manifest containing both contracts. During Phase 2 of wire_from_manifest the second contract would attempt to register dispatcher IDs already registered by the first, raising ONEX_CORE_064_DUPLICATE_REGISTRATION and crash-looping the effects runtime. Fix: track seen contract names in discover_contracts(). When a duplicate name is encountered, the second occurrence is dropped and recorded as a ModelDiscoveryError with a clear message identifying both packages. First occurrence wins. Includes a unit test covering the exact cross-package collision scenario. * test(OMN-11958): cover duplicate contract discovery integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11975): use asyncio.gather for concurrent consumer startup (#1731) * fix(OMN-11975): use asyncio.gather for concurrent consumer startup start_consuming() started 224 consumers serially (18-37min). Replace the for-loop with asyncio.gather, pre-reserving _pending_consumer_keys inside the lock before any concurrent call begins (same reservation pattern as subscribe()). Wall-clock startup time becomes the slowest single consumer rather than the sum of all 224. return_exceptions=True lets every consumer attempt startup; the first failure is re-raised after all have settled. * test(OMN-11975): cover concurrent start_consuming integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * chore(OMN-11513): update pricing manifest to current fleet model IDs (#1733) Replace four stale local model entries (qwen3-coder-30b-a3b, qwen3-14b, qwen3-embedding-8b, deepseek-r1-distill-qwen-32b) with the five current fleet models as documented in OMN-11513 lane map evidence (2026-05-22): - qwen3.6-35b-a3b: vLLM GPTQ-Int4 on .201:8000, 131072 ctx, ~180 t/s - qwen3.6-27b-mtp-iq4xs: llamacpp MTP on .201:8001, 98304 ctx - text-embedding-gte: vLLM gte-Qwen2-1.5B on .201:8002, 8192 ctx - deepseek-v4-flash: ds4.c DeepSeek V4 Flash 284B on .200:8101, 131072 ctx - deepseek-v4-pro: alias for flash with higher reasoning effort on .200:8101 All local models retain zero API cost (LOCAL_ZERO_API_COST_POLICY). Existing cloud model entries (Claude, GPT, Gemini) are unchanged. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11513): replace silent localhost fallbacks with required env vars in LLM path (#1734) 7 os.environ.get("VAR", "http://localhost:...") calls replaced with os.environ["VAR"] (fail-fast on missing config per CLAUDE.md rule #8): - handler_llm_completion.py: LLM_CODER_URL last-resort fallback removed - adapter_code_analysis_enrichment.py: LLM_CODER_URL - adapter_llm_provider_openai.py: LLM_CODER_URL - adapter_code_review_analysis.py: LLM_CODER_FAST_URL - adapter_test_boilerplate_generation.py: LLM_CODER_FAST_URL - adapter_documentation_generation.py: LLM_DEEPSEEK_R1_URL - adapter_summarization_enrichment.py: LLM_QWEN_72B_URL Silent fallbacks to localhost fail on .201 without surfacing an error. All 7 env vars are already set in ~/.omnibase/.env and the generated compose. 322 unit tests pass. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(B9): align pricing manifest keys with vLLM-served model IDs Rename the two local model keys added in PR 1733 to match the exact model IDs returned by /v1/models on the serving endpoints. ModelPricingTable does an exact-match dict lookup, so keys must match what gets written into delegation_events.delegated_to via endpoint_registry.yaml. - qwen3.6-35b-a3b → Qwen3.6-35B-A3B (matches .201:8000 /v1/models) - qwen3.6-27b-mtp-iq4xs → Qwen3.6-27B-MTP-IQ4_XS.gguf (matches .201:8001 /v1/models) - text-embedding-gte, deepseek-v4-flash, deepseek-v4-pro unchanged (already match) 69 pricing-related unit tests pass. * fix(OMN-11996): register DispatchResultApplier for node_delegate_skill_orchestrator (#1736) HandlerDelegateSkill was processing delegation commands and returning ModelDelegateSkillResponse, but without a result applier registered for the contract, the auto-wiring callback silently discarded the handler result. onex.evt.omnimarket.delegate-skill-completed.v1 was never published, causing the CLI adapter to time out on every invocation. Adds DispatchResultApplier for node_delegate_skill_orchestrator with output_topic=onex.evt.omnimarket.delegate-skill-completed.v1 and the allowed_output_topics allowlist covering both completed and failed topics. Follows the same registration pattern as build_loop_orchestrator. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * chore(OMN-11513): refresh receipt gate PR metadata * chore(OMN-11513): baseline service_kernel.py topic literals from PR #1736 PR #1736 added intentional topic literal constants with onex-topic-allow annotations but did not update topic_literal_baseline.txt. The Arch Invariants CI check uses AST scanning and does not read inline annotations, so these two lines are picked up as new violations in the PR merge check. * fix(OMN-11513): remove topic literal baseline suppression * Revert "fix(OMN-11513): remove topic literal baseline suppression" This reverts commit 4e6795c. * fix(OMN-11513): sync llm completion contract * test(OMN-11513): provide required LLM endpoint for adapter tests * chore(OMN-11513): retrigger transient split CI * chore(OMN-11513): retrigger after runner disk cleanup * fix(OMN-11997): remove non-IMMUTABLE ::date functional index from intelligence migrations 024/025 (#1738) * fix(OMN-11997): remove non-IMMUTABLE ::date functional index from intelligence migrations 024/025 PG 16.14 rejects CREATE INDEX on (created_at::date) because the ::date cast is STABLE not IMMUTABLE. Migration 024 created the broken index; migration 025 was supposed to fix it but recreated the same broken form. Fix: remove idx_llm_delegation_call_log_date from 024 entirely (the plain created_at index at line 99 already serves range queries). Update 025 to only DROP the index with no recreation. * ci: trigger receipt gate re-run after OCC SHA update * fix(OMN-11997): move delegate-skill topic literals to topic_constants to fix arch invariant The two raw topic strings added in OMN-11996 (onex.evt.omnimarket.delegate-skill-*) violated the no-hardcoded-topics arch invariant. Move them to topic_constants.py, regenerate the omnimarket enum, and update service_kernel.py to use the constants. * ci: trigger full CI rerun with OCC#1611 deploy evidence * test(OMN-11997): provide LLM_CODER_URL fixture for adapter unit tests test_adapter_llm_provider_openai.py breaks in CI because OMN-11513 PR #1734 replaced the silent localhost fallback in AdapterLlmProviderOpenai.__init__ with os.environ["LLM_CODER_URL"] (required env var). The companion test fix from cb92d56 hasn't merged yet. Apply the same autouse monkeypatch fixture to unblock CI Tests Gate. * chore(OMN-11997): retrigger transient type safety * chore(OMN-11997): retrigger transient runner disk CI * ci: trigger re-run after runner disk-full errors * chore(OMN-11997): retrigger transient dependency fetch --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-11998): add 082_swarm_runs migration for swarm dispatch projection (#1737) * feat(OMN-11998): add 082_swarm_runs migration for swarm dispatch projection table Promotes the swarm_runs table from dev-lane direct SQL to a proper forward migration so migration-gate applies it automatically on all lanes. * chore(OMN-11998): update schema fingerprint for 082_swarm_runs migration New migration file changes the migration set hash; stamp updated via check_schema_fingerprint.py stamp. * fix(OMN-11998): use generated delegate skill topics * fix(OMN-11998): declare delegate skill terminal topics * test(OMN-11513): provide required LLM endpoint for adapter tests * chore(OMN-11998): retrigger after runner disk cleanup * chore(OMN-11998): retrigger after runner network cleanup --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12002): allow auto-wired handlers to publish to all contract-declared topics (#1740) * fix(OMN-12002): prefer handle_async over handle in auto-wired dispatch callbacks Auto-wired dispatch callbacks previously called handler_instance.handle unconditionally. FSM orchestrator handlers (HandlerSwarmDispatchOrchestrator) expose handle_async as the runtime-publish entry point; handle is the no-publish sync/standalone path. Messages dispatched via the event-bus callback loop therefore never triggered the handler's Kafka publishes, silently dropping all sub-command topics. Fix: _make_dispatch_callback uses MRO inspection at wiring time to detect an explicitly-declared handle_async. If found (and callable), it becomes the effective dispatch target, preserving all downstream bus.publish calls. Falls back to handle for handlers that only implement the sync interface. MRO inspection (cls.__dict__ lookup) is required over callable() to avoid false-positives from MagicMock auto-attributes in existing tests. Adds 6 unit tests covering: handle_async preference, side-effect publish execution, multi-topic publish (all fire), sync-only handler fallback, async-handle-only fallback, and non-callable handle_async attribute handling. * fix(OMN-12002): type async dispatch target * fix(OMN-11996): use enum constants for delegate-skill topics, fix arch invariants Replace raw string literals with EnumOmnimarketTopic enum members in service_kernel.py to satisfy the Arch Invariants (OMN-3343) check. Adds TOPIC_DELEGATE_SKILL_COMPLETED and TOPIC_DELEGATE_SKILL_FAILED to topic_constants.py (supplementary source for the enum generator), then regenerates enum_omnimarket_topic.py to include the two new members. * chore(OMN-12002): retrigger after runner disk cleanup --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * chore(OMN-11513): retrigger transient dependency fetch * fix(OMN-11513): remove duplicate delegate skill topic constants * chore(OMN-11513): retrigger transient hatchling fetch --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
…ts (#1741) * fix(OMN-11994): handler_wiring psycopg2 list adaptation for text[] columns (#1730) * fix(OMN-11994): handler_wiring psycopg2 list adaptation for text[] columns Only apply psycopg2.extras.Json() to dict values, not list values. Lists intended for Postgres text[] columns (models_used, machines_used) were being wrapped in Json() producing malformed array literals. * test(OMN-11994): cover psycopg2 text array adaptation * chore(OMN-11994): rerun gates after OCC merge * style(OMN-11994): format handler wiring integration test --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11958): remove duplicate dispatcher for node_architecture_graph_query_effect (#1732) * fix(OMN-11958): deduplicate contract names in discover_contracts to prevent DUPLICATE_REGISTRATION crash When two installed packages (e.g. omnibase_infra and omnimarket) register onex.nodes entry points whose contract.yaml files declare the same `name` field, the auto-wiring engine previously built a manifest containing both contracts. During Phase 2 of wire_from_manifest the second contract would attempt to register dispatcher IDs already registered by the first, raising ONEX_CORE_064_DUPLICATE_REGISTRATION and crash-looping the effects runtime. Fix: track seen contract names in discover_contracts(). When a duplicate name is encountered, the second occurrence is dropped and recorded as a ModelDiscoveryError with a clear message identifying both packages. First occurrence wins. Includes a unit test covering the exact cross-package collision scenario. * test(OMN-11958): cover duplicate contract discovery integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11975): use asyncio.gather for concurrent consumer startup (#1731) * fix(OMN-11975): use asyncio.gather for concurrent consumer startup start_consuming() started 224 consumers serially (18-37min). Replace the for-loop with asyncio.gather, pre-reserving _pending_consumer_keys inside the lock before any concurrent call begins (same reservation pattern as subscribe()). Wall-clock startup time becomes the slowest single consumer rather than the sum of all 224. return_exceptions=True lets every consumer attempt startup; the first failure is re-raised after all have settled. * test(OMN-11975): cover concurrent start_consuming integration --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * chore(OMN-11513): update pricing manifest to current fleet model IDs (#1733) Replace four stale local model entries (qwen3-coder-30b-a3b, qwen3-14b, qwen3-embedding-8b, deepseek-r1-distill-qwen-32b) with the five current fleet models as documented in OMN-11513 lane map evidence (2026-05-22): - qwen3.6-35b-a3b: vLLM GPTQ-Int4 on .201:8000, 131072 ctx, ~180 t/s - qwen3.6-27b-mtp-iq4xs: llamacpp MTP on .201:8001, 98304 ctx - text-embedding-gte: vLLM gte-Qwen2-1.5B on .201:8002, 8192 ctx - deepseek-v4-flash: ds4.c DeepSeek V4 Flash 284B on .200:8101, 131072 ctx - deepseek-v4-pro: alias for flash with higher reasoning effort on .200:8101 All local models retain zero API cost (LOCAL_ZERO_API_COST_POLICY). Existing cloud model entries (Claude, GPT, Gemini) are unchanged. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11513): replace silent localhost fallbacks with required env vars in LLM path (#1734) 7 os.environ.get("VAR", "http://localhost:...") calls replaced with os.environ["VAR"] (fail-fast on missing config per CLAUDE.md rule #8): - handler_llm_completion.py: LLM_CODER_URL last-resort fallback removed - adapter_code_analysis_enrichment.py: LLM_CODER_URL - adapter_llm_provider_openai.py: LLM_CODER_URL - adapter_code_review_analysis.py: LLM_CODER_FAST_URL - adapter_test_boilerplate_generation.py: LLM_CODER_FAST_URL - adapter_documentation_generation.py: LLM_DEEPSEEK_R1_URL - adapter_summarization_enrichment.py: LLM_QWEN_72B_URL Silent fallbacks to localhost fail on .201 without surfacing an error. All 7 env vars are already set in ~/.omnibase/.env and the generated compose. 322 unit tests pass. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11996): register DispatchResultApplier for node_delegate_skill_orchestrator (#1736) HandlerDelegateSkill was processing delegation commands and returning ModelDelegateSkillResponse, but without a result applier registered for the contract, the auto-wiring callback silently discarded the handler result. onex.evt.omnimarket.delegate-skill-completed.v1 was never published, causing the CLI adapter to time out on every invocation. Adds DispatchResultApplier for node_delegate_skill_orchestrator with output_topic=onex.evt.omnimarket.delegate-skill-completed.v1 and the allowed_output_topics allowlist covering both completed and failed topics. Follows the same registration pattern as build_loop_orchestrator. Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11996): use enum constants for delegate-skill topics, fix arch invariants Replace raw string literals with EnumOmnimarketTopic enum members in service_kernel.py to satisfy the Arch Invariants (OMN-3343) check. Adds TOPIC_DELEGATE_SKILL_COMPLETED and TOPIC_DELEGATE_SKILL_FAILED to topic_constants.py (supplementary source for the enum generator), then regenerates enum_omnimarket_topic.py to include the two new members. * chore: retrigger CI after OCC receipts added * chore: retrigger CI after OCC receipt PR-binding fix * fix(OMN-11996): bump node_llm_completion_effect contract version for handler fail-fast change Handler now requires LLM_CODER_URL to be set (no localhost fallback); bump contract patch version to satisfy contract-sync gate. * chore(OMN-11996): retrigger transient CI checks * chore(OMN-11996): retrigger transient migration CI * chore(OMN-11996): retrigger transient dependency fetch * fix(OMN-11997): remove non-IMMUTABLE ::date functional index from intelligence migrations 024/025 (#1738) * fix(OMN-11997): remove non-IMMUTABLE ::date functional index from intelligence migrations 024/025 PG 16.14 rejects CREATE INDEX on (created_at::date) because the ::date cast is STABLE not IMMUTABLE. Migration 024 created the broken index; migration 025 was supposed to fix it but recreated the same broken form. Fix: remove idx_llm_delegation_call_log_date from 024 entirely (the plain created_at index at line 99 already serves range queries). Update 025 to only DROP the index with no recreation. * ci: trigger receipt gate re-run after OCC SHA update * fix(OMN-11997): move delegate-skill topic literals to topic_constants to fix arch invariant The two raw topic strings added in OMN-11996 (onex.evt.omnimarket.delegate-skill-*) violated the no-hardcoded-topics arch invariant. Move them to topic_constants.py, regenerate the omnimarket enum, and update service_kernel.py to use the constants. * ci: trigger full CI rerun with OCC#1611 deploy evidence * test(OMN-11997): provide LLM_CODER_URL fixture for adapter unit tests test_adapter_llm_provider_openai.py breaks in CI because OMN-11513 PR #1734 replaced the silent localhost fallback in AdapterLlmProviderOpenai.__init__ with os.environ["LLM_CODER_URL"] (required env var). The companion test fix from cb92d56 hasn't merged yet. Apply the same autouse monkeypatch fixture to unblock CI Tests Gate. * chore(OMN-11997): retrigger transient type safety * chore(OMN-11997): retrigger transient runner disk CI * ci: trigger re-run after runner disk-full errors * chore(OMN-11997): retrigger transient dependency fetch --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * feat(OMN-11998): add 082_swarm_runs migration for swarm dispatch projection (#1737) * feat(OMN-11998): add 082_swarm_runs migration for swarm dispatch projection table Promotes the swarm_runs table from dev-lane direct SQL to a proper forward migration so migration-gate applies it automatically on all lanes. * chore(OMN-11998): update schema fingerprint for 082_swarm_runs migration New migration file changes the migration set hash; stamp updated via check_schema_fingerprint.py stamp. * fix(OMN-11998): use generated delegate skill topics * fix(OMN-11998): declare delegate skill terminal topics * test(OMN-11513): provide required LLM endpoint for adapter tests * chore(OMN-11998): retrigger after runner disk cleanup * chore(OMN-11998): retrigger after runner network cleanup --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-12002): allow auto-wired handlers to publish to all contract-declared topics (#1740) * fix(OMN-12002): prefer handle_async over handle in auto-wired dispatch callbacks Auto-wired dispatch callbacks previously called handler_instance.handle unconditionally. FSM orchestrator handlers (HandlerSwarmDispatchOrchestrator) expose handle_async as the runtime-publish entry point; handle is the no-publish sync/standalone path. Messages dispatched via the event-bus callback loop therefore never triggered the handler's Kafka publishes, silently dropping all sub-command topics. Fix: _make_dispatch_callback uses MRO inspection at wiring time to detect an explicitly-declared handle_async. If found (and callable), it becomes the effective dispatch target, preserving all downstream bus.publish calls. Falls back to handle for handlers that only implement the sync interface. MRO inspection (cls.__dict__ lookup) is required over callable() to avoid false-positives from MagicMock auto-attributes in existing tests. Adds 6 unit tests covering: handle_async preference, side-effect publish execution, multi-topic publish (all fire), sync-only handler fallback, async-handle-only fallback, and non-callable handle_async attribute handling. * fix(OMN-12002): type async dispatch target * fix(OMN-11996): use enum constants for delegate-skill topics, fix arch invariants Replace raw string literals with EnumOmnimarketTopic enum members in service_kernel.py to satisfy the Arch Invariants (OMN-3343) check. Adds TOPIC_DELEGATE_SKILL_COMPLETED and TOPIC_DELEGATE_SKILL_FAILED to topic_constants.py (supplementary source for the enum generator), then regenerates enum_omnimarket_topic.py to include the two new members. * chore(OMN-12002): retrigger after runner disk cleanup --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com> * fix(OMN-11996): remove duplicate delegate skill topic constants * chore(OMN-11996): retrigger transient dependency fetch --------- Co-authored-by: jonahgabriel <jonahgabriel@users.noreply.github.com>
|
Warning Review limit reached
More reviews will be available in 28 minutes and 21 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (4)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
OMN-12432 - merge-queue lint cancellation diagnosis
Adds docs/diagnosis-merge-queue-lint-cancellation.md, the diagnosis written while investigating the omnibase_infra dev merge-queue stall (OMN-12432 #1789, OMN-12416 #1781, OMN-12421 #1782, OMN-12433 #1792). It documents the root-cause chain: cancelled Lint jobs -> CI Summary never posts success -> queue refuses enqueuePullRequest.
This was an untracked file in the canonical clone; moved to a worktree and preserved during the 2026-05-30 cleanup rather than discarded.
dod_evidence (OMN-12432)
Evidence-Source: bd187f8da0dfe92c0ba8aa42ad6df24638a8e5d8
Evidence-Ticket: OMN-12432