feat: bkit v2.0.5 — Gemini CLI v0.39.0 migration + SessionStart slim + baseline runner repair - #23
Merged
Conversation
…r repair
Three logical layers bundled per user request:
1. Backlog PDCA artifacts (v0.37.1 → v0.38.2)
- 5 plans, 7 research docs (incl. context-performance + post-v0.38.2),
5 impact analyses, 5 reports, 1 v2.1.0 plan/design.
- Carry-over from prior unmerged cycles.
2. v0.39.0 migration cycle (Strategy B' — Spot Validation + Defensive Test)
- bkit.config.json: testedVersions += 0.38.0/0.38.1/0.38.2/0.39.0
- lib/gemini/version.js: 7 new v0.39+ feature flags
(hasInvokeAgent, hasContextManagerSidecar, hasMcpAuthBlock,
hasToolControlledDisplay, hasMemoryInbox, hasGeminiPlansDirEnv,
hasUseAgentStream) with v0.38.2 boundary guard verified
- tests/suites/tc113-session-start-duplication-defense.js: NEW
8 tests, defensive carrier for upstream Issue #25655 (fix PR #25827 OPEN).
Tc107 was occupied (modes-migration), renamed to tc113.
- tests/run-all.js: tc113 registered (Sprint 9)
- docs/01-plan/features/v2.1.0-context-optimization.plan.md: v0.40.0 cycle
re-entry hint embedded (ContextManager+Sidecar #24752 + MemoryManager 4-tier
#25716 + autoMemory split #25601 trigger combined re-design).
- Plan/Research/Impact/Do.Analysis/Report all under docs/.
3. Baseline test-runner repair (side-effect of Wave 3 runtime verification)
- tests/suites/tc80-architecture-v200.js + tc95-architecture-migration.js:
guarded process.exit(failed > 0 ? 1 : 0) with `if (require.main === module)`.
Previously these suites silently aborted run-all.js, masking sprint 5–9
(TC-81 through TC-113) aggregation entirely.
- hooks/scripts/session-start.js: guarded main() so requiring as a module
(e.g. tc100 COMP-07) no longer kills the parent process. Function
exports added for module consumers.
- tests/test-utils.js: runSuite now aggregates `mod.precomputed` results
emitted by guarded suites.
- Iterate quick wins (count → semantic assertion pattern):
* tc18 V156-14 (was: 48 keys magic number)
* tc82 VER-01/02/03 (was: 19/14/14 magic numbers)
* tc88 SS-24 (was: literal 'v2.0.3' string)
Results
- v0.39.0 direct AC: 9/9 = 100% (Plan §9 P0 + P1 all green)
- Self-regression: 0 (V156-14 self-fixed in iterate phase)
- Full runner: 2018 tests / 1914 pass / 80 fail / 24 skip / 94.8%
Pre-existing baseline issues newly visible after runner-abort wall removal.
PDCA-* cluster (35) traces to phantom API in lib/pdca/status.js — separate
cycle (`bkit-baseline-stabilization`) recommended.
- Hard-Gate: Plan §3 NG list 13 items 100% honored. Both sharp decisions
(ContextManager timing → Option B; #25655 strategy → Option X) followed.
Documents
- docs/01-plan/features/gemini-cli-v0.39.0-migration.plan.md
- docs/01-plan/research/gemini-cli-v0.39.0-research.md
- docs/01-plan/research/gemini-cli-v0.40.0-preview-research.md
- docs/03-analysis/gemini-cli-v0.39.0-impact.analysis.md
- docs/03-analysis/gemini-cli-v0.39.0-do.analysis.md
- docs/04-report/gemini-cli-v0.39.0-migration.report.md (Do/Check/Iterate §10)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… bump to 2.0.5 PDCA full cycle (Plan→Design→Do→Iterate→QA→Report) under L4 Full-Auto mode. Plan §1 G1~G7 = 7/7 = 100%. Self-regression: 0. ## Changes ### Manifest & version sync (v2.0.4 → v2.0.5) - gemini-extension.json, bkit.config.json: version → 2.0.5 - hooks/scripts/session-start.js: 'v2.0.4' / metadata '2.0.3' → 2.0.5 - GEMINI.md: header + footer → v2.0.5 ### SessionStart slim default + verbose env var (mitigates Issue #25655) - hooks/scripts/session-start.js#generateDynamicContext: default returns single-line activation header only. - BKIT_SESSION_START_VERBOSE=true restores legacy multi-section body. - GEMINI.md absorbs PDCA Core Rules / Agent Auto-Triggers / Natural Language Feature Request sections (auto-loaded by Gemini CLI on every session). - Visual impact of Issue #25655 (CLI v0.38.x+ duplicates systemMessage) reduced ~98% (60 lines × 2 → 1 line × 2). Hook contract unchanged. ### MCP list_agents 16 → 21 - mcp/bkit-server.js#AGENTS: PM Agent Team added (pm-lead, pm-discovery, pm-strategy, pm-research, pm-prd) — were defined in agents/ but missing from registry, causing list_agents to return only 16/21. ### Defensive test (new + iterate) - tests/suites/tc114-session-start-slim-mode.js: NEW 6 tests, all PASS. Asserts slim default + verbose mode + manifest cross-reference + Issue #25655 hook contract preservation. - tests/suites/tc113-...js: 8/8 PASS retained (no contract drift). ### Iterate quick wins (count → semantic assertion refactor) Self-introduced regressions identified post-Wave-3 and recovered in iter 1: - tc07 CFG-02/03, tc19 V156-51: hardcoded version → manifest cross-ref - tc80 TC80-03/04/06: hardcoded agent counts (16/10/4) → semantic floor (≥) - tc94 TC94-67, tc98 TC98-02/05: line count limits raised for slim/verbose branch absorption (GEMINI.md 30→100, session-start.js 500→600) - tc01 HOOK-07/08, tc08 CTX-09, tc10 AF-01, tc22 TC-22-01/02: opt into BKIT_SESSION_START_VERBOSE=true for tests that asserted body markers ## Results - Self-regression: 0 (comm -13 iter1 iter2 = empty) - Pre-existing baseline recovery: +3 (81 → 78 fail) - Full runner: 2024 tests / 1917 pass / 78 fail / 24 skip / 94.7% - Hard-Gate: Plan §3 NG list 100% honored (NG1 wrong-layer maintained) - Risk: LOW — verbose env var preserves backward compat, GEMINI.md absorbs body so no information lost Documents - docs/01-plan/features/bkit-v2.0.5-finalization.plan.md - docs/02-design/features/bkit-v2.0.5-finalization.design.md - docs/04-report/bkit-v2.0.5-finalization.report.md Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two-cycle release that completes the Gemini CLI v0.37.x → v0.39.0 migration and ships a UX-focused finalization (SessionStart slim default, MCP
list_agentsfix, version bump to v2.0.5). Self-introduced regressions: 0. Hard-Gate (Plan §3 NG list) honored throughout.d031e1c): Gemini CLI v0.37.1 → v0.38.2 → v0.39.0 migration withtc113defensive carrier for upstream Issue #25655 + baseline test runner repair (tc80/tc95/session-start.js process.exit guards) that exposed the true baseline pass rate for the first time.a6acd11): bkit v2.0.5 finalization — SessionStartsystemMessageslim default,mcp_bkit_list_agents16 → 21 (PM Agent Team added), manifest/code/docs version sync to v2.0.5.What's New
🎯 Issue #25655 mitigation (no upstream change required)
Gemini CLI v0.38.x+ duplicates the SessionStart
systemMessagepayload (renderer-level bug, google-gemini/gemini-cli#25655, fix PR #25827 OPEN). bkit hooks emit exactly one JSON line — proven bytc113(8/8 PASS) — so the duplication is purely cosmetic. v2.0.5 collapses the activation message to a single line by default, reducing the visible noise by ~98% even when the duplication still occurs.🤖 Full Agent visibility via MCP
mcp_bkit_list_agentsnow exposes all 21 agents (was 16). The PM Agent Team —pm-lead,pm-discovery,pm-strategy,pm-research,pm-prd— was defined underagents/but missing from the MCP registry. PM workflows can now be orchestrated through the MCP surface end-to-end.🏷 7 new feature flags for Gemini CLI v0.39.0
lib/gemini/version.jsadds boundary-guarded gates:All flags resolve
trueat v0.39.0 andfalseat v0.38.2 (verified).🛠 Baseline test runner repair
Three pre-existing scripts (
tc80-architecture-v200.js,tc95-architecture-migration.js,hooks/scripts/session-start.js) calledprocess.exit()unconditionally, which silently aborted the parent runner and masked sprint 5–9 results. Guarded withif (require.main === module)+module.exports = { tests: [], precomputed }. The full runner now reaches its terminal aggregator for the first time, exposing a real baseline pass rate of 94.7% (1917/2024) where it was previously invisible behind the abort wall.📚 SessionStart context migration to GEMINI.md
The PDCA Core Rules / Agent Auto-Triggers / Natural Language Feature Request sections (~58 lines previously printed every session) now live in
GEMINI.md, which the Gemini CLI auto-loads on every session. No information is lost. The only change is where the user sees it.🆘 Backward compatibility
BKIT_SESSION_START_VERBOSE=truerestores the legacy v2.0.4 multi-section body. Users who relied on the verbose output have a one-line opt-in.User Experience Changes
GEMINI.md(auto-imported by CLI on every session)/mcp list(bkit server)pm-lead,pm-discovery,pm-strategy,pm-research,pm-prdall callableBKIT_SESSION_START_VERBOSE=trueBreaking Changes
'2.0.4',19keys,48keys,16agents,10/4tier counts) were refactored to manifest cross-reference / semantic floor so future minor releases will not break tests. Test-only change; no runtime impact.Migration Guide
For most users — none required. Update via your usual extension install path; v2.0.5 will be activated on the next
geminisession.If you depended on the verbose SessionStart body being printed every session:
export BKIT_SESSION_START_VERBOSE=trueAdd to your shell rc to make it permanent.
Test Results
tc113(Issue #25655 hook contract carrier)tc114(slim default + verbose env var, NEW)comm -13 iter1 iter2empty)Files Changed
42 files in cycle 1 + 21 files in cycle 2 = 63 file-level changes across 2 commits. Highlights:
gemini-extension.json,bkit.config.json,GEMINI.md— version sync to 2.0.5hooks/scripts/session-start.js— slim default + verbose env var +require.mainguardmcp/bkit-server.js— PM Agent Team added (16 → 21)lib/gemini/version.js— 7 new v0.39+ feature flags with boundary guardstests/suites/tc113-...-defense.js(NEW) — Issue #25655 hook contract carriertests/suites/tc114-...-slim-mode.js(NEW) — slim/verbose mode unit teststests/suites/tc80-architecture-v200.js,tc95-architecture-migration.js— process.exit guardstests/test-utils.js—precomputedaggregation supportdocs/01-plan/,docs/02-design/,docs/03-analysis/,docs/04-report/,docs/05-qa/— full PDCA paper trail (Plan + Design + Impact + Do.Analysis + Report + QA)Related
spawn_agentMCP design)Next Steps (separate cycles)
bkit-baseline-stabilization— recover the 78 pre-existing failures (PDCA-* phantom API inlib/pdca/status.jsis the largest cluster)gemini-cli-v0.40.0-migration— when v0.40.0 stable ships, re-enter via the v2.1.0-context-optimization plan (re-entry hint embedded)🤖 Generated with Claude Code