Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
115 commits
Select commit Hold shift + click to select a range
b4a07e9
feat: support matview_refresh_interval "off" to disable logstore matv…
jeremym-tanium Jul 30, 2026
ed52d94
V2.0.0 (#4365)
akshaydeo Jul 30, 2026
3ef6c5e
gomod fixes (#5731)
akshaydeo Jul 30, 2026
d7e9ae5
third party notice (#5735)
akshaydeo Jul 31, 2026
836bd94
feat(mcp-guardrails): add MCP log redaction changes (#5744)
Madhuvod Jul 31, 2026
67898dc
feat(mcp-guardrails): ui changes (#5745)
Madhuvod Jul 31, 2026
c857340
added plugin logs in mcp logs (#5746)
Madhuvod Jul 31, 2026
ac8d07d
mcp guardrails : config,helm and docs changes (#5758)
Madhuvod Aug 2, 2026
e7c62fc
adds first time setup token to avoid opening new setup to the world (…
akshaydeo Aug 3, 2026
e08de39
brings back onboarding widget (#5784)
akshaydeo Aug 3, 2026
720fb26
path normalization auth bypass (#5763)
akshaydeo Aug 3, 2026
d22edce
[fix]: clear stuck entity-assignment validation on virtual key sheet …
CMWR421 Aug 3, 2026
b8bc2c8
[fix]: transcription - carry multipart filename through ingress so co…
AdityaPainuli Aug 4, 2026
5eea53d
fix: fixes normalization and whitespace trimming for metrics exportin…
roroghost17 Aug 4, 2026
5e8b41b
feat: add team / customer / bu ids and names to OTEL metrics (#5848)
roroghost17 Aug 4, 2026
56de452
feat: adds service instance id to OTEL attributes (#5849)
roroghost17 Aug 4, 2026
bc1345e
docs: add user scopes, pagination, and extended pricing fields to pri…
Pratham-Mishra04 Aug 5, 2026
ae0888c
docs: add user scopes, fast-mode costs, OCR fields, and pagination to…
Pratham-Mishra04 Aug 5, 2026
859e34c
feat: add pricing fields for >2048px and >4096px image output cost ov…
Pratham-Mishra04 Aug 5, 2026
d4edcb6
mod fix (#5864)
akshaydeo Aug 5, 2026
ac9b4a1
strip prefix check for openai reasoning models check (to accmodate ma…
akshaydeo Aug 5, 2026
e575bab
fix(vertex): support API-key/context-header auth in cached content me…
TransactCharlie Aug 5, 2026
5453ef2
docs: adds a subsection for mcp guardrails (#5869)
Madhuvod Aug 5, 2026
4afcdd5
feat: send HTTP/2 PING keepalives on the Bedrock provider (#5213)
jeremym-tanium Aug 5, 2026
2fb58b7
fixes generateContent to keep safetyRatings and avgLogprobs (#5877)
akshaydeo Aug 5, 2026
dafc587
fix(bedrock): preserve encrypted reasoning replay signatures (#5879)
zachgersh Aug 5, 2026
d713f15
nil delta with non-nil finish reason doesnt bail out anymore (#5878)
akshaydeo Aug 6, 2026
07c4d8f
perf(anthropic): avoid copying known request fields (#5809)
zachgersh Aug 6, 2026
5613ead
dependabot fixes (#5889)
akshaydeo Aug 6, 2026
041817a
fix(core): carry tool-result is_error through chat completions (#5450)
AidanAllchin Aug 6, 2026
4a51039
fix: base provider resolution (#5897)
TejasGhatte Aug 6, 2026
f3f3b8e
fix(transports): drop trailing blank line from SSE heartbeat frame (#…
jeremym-tanium Aug 6, 2026
155c0cd
harness tests (#5880)
akshaydeo Aug 6, 2026
da60e05
file type mapping bug fix (#5884)
akshaydeo Aug 6, 2026
8ab75f3
keep pricing objects in anthropic path for ai-sdk compatibility (#5886)
akshaydeo Aug 6, 2026
5d6b68d
deepseek thinking fixes (#5888)
akshaydeo Aug 6, 2026
4b17b6f
normalize tool_search_tool_* on the Responses path (#5891)
akshaydeo Aug 6, 2026
2c6f132
fail soft on invalid_encrypted_content by stripping replayed reasonin…
akshaydeo Aug 6, 2026
06eca74
mcp tool call with error handling (#5894)
akshaydeo Aug 6, 2026
9303c56
fixed message start frame (#5907)
akshaydeo Aug 6, 2026
9fbbacf
pr resolve skill (#5922)
akshaydeo Aug 6, 2026
2dccfa8
service_teir patches (#5928)
akshaydeo Aug 7, 2026
0b3fa58
streamreader locks to avoid concurrent stream writes (#5927)
akshaydeo Aug 7, 2026
7b1149e
fix(bedrock): map ConverseStream stopReason for tool-use turns (#5209)
axelray-dev Aug 7, 2026
55871ff
last_reset and current_usage override bug fix (#5932)
akshaydeo Aug 7, 2026
a4a2c97
cache control fixes for anthropic and bedrock (#5931)
akshaydeo Aug 7, 2026
6af0a32
[fix]: tracing - concurrent requests sharing a W3C traceID no longer …
AdityaPainuli Aug 7, 2026
b3189ef
[fix]: Anthropic - include_server_side_tool_invocations now reaches t…
AdityaPainuli Aug 7, 2026
cc4a0d6
adds harness run test (#5933)
akshaydeo Aug 7, 2026
4742c6c
adaptive thinking support for passthrough mode (#5934)
akshaydeo Aug 7, 2026
846bb8a
feat: add w3c trace id to context (#5945)
roroghost17 Aug 7, 2026
d1a9104
fix: align budget override validity with calendar boundaries (#5962)
danpiths Aug 7, 2026
25b094c
[fix]: schemas- omit absent tool-call function name on streaming delt…
AdityaPainuli Aug 8, 2026
55c5e60
fix: guard nil ConfigStore, propagate resource, and surface pending-b…
Pratham-Mishra04 May 29, 2026
8ee4a70
fix: address PR review findings - client-scoped test id, popup-blocke…
Pratham-Mishra04 May 29, 2026
095e2e4
feat: add per user oauth mcp support for config.json
Pratham-Mishra04 May 29, 2026
4569814
fix: close verify-headers double-submit race, preserve TLS/timeout/pe…
Pratham-Mishra04 May 29, 2026
e0c116c
fix: always label the bootstrap action "Authorize" regardless of auth…
Pratham-Mishra04 May 29, 2026
85d8cb3
fix: seed linked oauth_configs/config_mcp_clients rows for shared-tok…
Pratham-Mishra04 Jul 28, 2026
094e0cc
fix: close leaked sqlDB in flows-table perf setup, make OAuth flow cl…
Pratham-Mishra04 Jul 28, 2026
0f10fb5
fix: reject inactive tokens in ValidateToken, document shared/per-ide…
Pratham-Mishra04 Jul 29, 2026
7c63205
feat: generalize TokenRefreshWorker's auth-mode scope
Pratham-Mishra04 Jul 29, 2026
b691a45
fix: gate SSE OnConnectionLost on connection identity, preserve Needs…
Pratham-Mishra04 Jul 29, 2026
304f82c
fix: repair shared connections regardless of destructive hint, fail c…
Pratham-Mishra04 Jul 29, 2026
66e28dd
fix: remove 'View sessions' link from the MCP client edit sheet
Pratham-Mishra04 Jul 29, 2026
d52d849
fix: don't silently drop stored oauth scopes on decode failure, skip …
Pratham-Mishra04 Jul 29, 2026
c23107b
fix: restrict Reauthorize to shared OAuth clients, show loading state…
Pratham-Mishra04 Jul 29, 2026
db4dded
test: seed opposite-auth_mode token in TestAccessToken_MissingToken_R…
Pratham-Mishra04 Jul 29, 2026
a446015
test: cover every field PromoteSharedOauthTokenToAdmin transfers and …
Pratham-Mishra04 Jul 29, 2026
a16e180
fix: add per-entry version so a rejected stale Get can't evict a conc…
Pratham-Mishra04 Jul 30, 2026
e9cf929
fix: propagate ctx through userTokenCache.Fill so a canceled request …
Pratham-Mishra04 Jul 30, 2026
9e6703d
fix: propagate ctx through headerCredentialCache.Fill so a canceled r…
Pratham-Mishra04 Jul 30, 2026
e6db00c
fix: carry the ConfigHash-checkpoint regression test forward across O…
Pratham-Mishra04 Jul 2, 2026
062bb17
docs: mcp oauth and per user types config json support docs update
Pratham-Mishra04 May 29, 2026
e8c1c4b
docs: note pending_verification state as a 400 trigger for initiate-v…
Pratham-Mishra04 Jul 2, 2026
65fab08
docs: add secret var support to oauth client_id and client_secret docs
Pratham-Mishra04 Jul 2, 2026
a3b1248
docs: fix tool_sync_interval nanosecond claim, fix disable_vk_identit…
Pratham-Mishra04 Jul 30, 2026
9022abf
fix: wire tool_execution_timeout on MCP client creation, fix reauthor…
Pratham-Mishra04 Jul 30, 2026
2b4c10e
fix: select the pending-verification message from whether verificatio…
Pratham-Mishra04 Jul 30, 2026
daf56ec
fix: route pending token_exchange clients through the verify-exchange…
Pratham-Mishra04 Jul 30, 2026
b8c8414
fix: retain completed inflightClientOp in test double, fix idempotent…
Pratham-Mishra04 Jul 30, 2026
942cedf
fix: bind MCP connect attempts to entry identity, guard AddClient's dial
Pratham-Mishra04 Jul 30, 2026
133f891
docs: document make-before-break reconnection and shared-client retry…
Pratham-Mishra04 Jul 30, 2026
088e76f
fix: enforce non-empty audience/client_id and reject token_exchange f…
Pratham-Mishra04 Jul 30, 2026
de5b186
fix: configure bounded http.Server timeouts and request-body limit fo…
Pratham-Mishra04 Jul 30, 2026
61583ed
docs: scope needs_reauth's applicable auth types instead of claiming …
Pratham-Mishra04 Aug 2, 2026
959dfc8
fix: correct branch-resolution, backup-safety, conflict-marker, and v…
Pratham-Mishra04 Aug 2, 2026
e08d12c
feat: allow gating OAuthTokenRefreshWorker sweeps
Pratham-Mishra04 Aug 3, 2026
29fe665
fix: render MCP client state badges with spaces instead of underscores
Pratham-Mishra04 Aug 3, 2026
2e91d16
fix: resolve MCP client verify UX gaps
Pratham-Mishra04 Aug 5, 2026
372e300
feat: adds jwt option in sample mcp client
Pratham-Mishra04 Aug 5, 2026
f6c6b39
fix: preserve last-known tool maps across close-first reconnects, che…
Pratham-Mishra04 Aug 6, 2026
ad49912
fix: rebuild ephemeral client fresh across the whole connect+init ret…
Pratham-Mishra04 Aug 6, 2026
abd9075
feat: redesign MCP client edit sheet into tabbed layout with accessib…
Pratham-Mishra04 Aug 6, 2026
9c9f121
feat: apply MCP edit sheet's design language to the create sheet
Pratham-Mishra04 Aug 6, 2026
a464d9c
fix: break lock-order inversion in ConnectionCheckerManager, close da…
Pratham-Mishra04 Aug 6, 2026
cddddfc
fix: pin needs_session_stickiness across config.json reconciliation s…
Pratham-Mishra04 Aug 6, 2026
7379092
fix: use kebab-case data-testid values in MCP sessions filter sidebar
Pratham-Mishra04 Aug 6, 2026
83e5058
feat: replace MCP oauth grants' top filter bar with a side filter panel
Pratham-Mishra04 Aug 6, 2026
839bfd0
feat: add MCP server filter to Auth Sessions sidebar, scoped to per-u…
Pratham-Mishra04 Aug 6, 2026
6adc398
fix: don't treat a CAS loss to a still-active concurrent refresh as a…
Pratham-Mishra04 Aug 6, 2026
35163d1
docs: document that stateChangeCallback has no cross-transition order…
Pratham-Mishra04 Aug 6, 2026
53addcf
feat: add Virtual Key and Users filters to MCP Auth Sessions sidebar
Pratham-Mishra04 Aug 6, 2026
4fba9d7
feat: add VK and Users filters to OAuth Grants sidebar
Pratham-Mishra04 Aug 6, 2026
95e79e4
fix: discover tools synchronously for per-call MCP clients, fix share…
Pratham-Mishra04 Aug 7, 2026
d5e43ce
feat: persist and resync MCP tool discoveries uniformly across all cl…
Pratham-Mishra04 Aug 7, 2026
2d06da6
fix: persist discovered tools before triggering cluster propagation i…
Pratham-Mishra04 Aug 8, 2026
8eeee78
fix: correct per-user MCP state-projection and reauthorize completion…
Pratham-Mishra04 Aug 8, 2026
52d906b
docs: document needs_session_stickiness across Web UI, API, and conf…
Pratham-Mishra04 Aug 8, 2026
c3d2ca7
docs: add canonical MCP connections/states/lifecycles page, fix stale…
Pratham-Mishra04 Aug 8, 2026
ca34a32
docs(openapi): fix stale MCP connection-state enum, document prematu…
Pratham-Mishra04 Aug 8, 2026
c077583
docs: fix remaining reconnect-400 wording and tool-persistence gaps …
Pratham-Mishra04 Aug 8, 2026
0cf72ca
[fix]: gemini- map truncated responses to MAX_TOKENS finish reason (#…
AdityaPainuli Aug 9, 2026
4c208bb
fix(ui): skip password validation for redacted credential (#5953)
G-XD Aug 9, 2026
c1bd494
feat: logs chain support in logstore (backend)
impoiler Jul 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
226 changes: 213 additions & 13 deletions .claude/skills/investigate-issue/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,10 +14,17 @@ Fetch a GitHub issue, analyze the report, search the codebase for relevant code,
3. Codebase Analysis + Documentation Research (from Step 3, including sub-step 3e)
4. Impact Analysis (from Step 4)
5. Test Plan (from Step 5)
6. Full Presentation (Step 6 template)
6. Regression Rerun Scope (from Step 5e) -- which existing tests must rerun, and why
7. TDD: Failing Tests (from Step 5f) -- Bug issues only: the test source AND the actual
red output from the applicable Makefile recipe, pasted verbatim, BEFORE asking to
implement
8. Full Presentation (Step 6 template)

If any section is missing, go back and complete it before presenting the report.

For Bug issues the approval gate is NOT "may I write a test?" -- the failing test is already
written and shown to be red. The gate is "may I write the fix?"

## Usage

```
Expand All @@ -32,8 +39,13 @@ If any section is missing, go back and complete it before presenting the report.
3. **Search the codebase and research docs** -- Find relevant code, then research the libraries it depends on via Context7 and WebSearch
4. **Analyze impact** -- Cross-reference codebase findings with documentation to identify side effects, dependencies, and breaking changes
5. **Suggest tests** -- If changes touch `core/`, recommend specific LLM and MCP test additions
6. **Present the plan** -- Show findings and recommended changes to the user
7. **Implement with approval** -- After plan approval, make changes one at a time with user confirmation
5b. **Scope the reruns** -- Use coverage to determine which existing tests exercise the lines
you plan to change, and tier them by necessity
5c. **Go red (Bug issues)** -- Write the regression test(s) and run them to confirm failure for
the right reason, BEFORE presenting the plan
6. **Present the plan** -- Show findings, the red test output, the rerun scope, and the
recommended fix
7. **Implement with approval** -- After approval, apply the fix and show red-to-green

## Step 1: Fetch the Issue

Expand Down Expand Up @@ -410,6 +422,127 @@ If UI changes are involved, recommend E2E test updates following the patterns in
make run-e2e FLOW=<feature>
```

### 5e. Determine Rerun Scope from Coverage

Do not guess which tests are affected -- measure it. Package-level coverage is too coarse:
it proves the package is tested, not that any test executes the specific lines you are
changing. Attribute coverage per test.

**Step 1 -- Establish the changed-line set.** For each file in the Implementation Plan, list
the functions/methods you will modify AND the line ranges within them. Step 3 filters the
coverage profile by line range, so symbol names alone are not enough.

**Step 2 -- Find candidate tests.** Every test in the changed file's package, plus every test
in the packages of its callers (from the Step 4b `grep` for callers):

Enumerate by PACKAGE, not by file or by symbol name. All `_test.go` files in a directory
compile into one test binary, and many tests reach the target through helpers or fixtures
rather than naming it -- so a textual grep for `<TargetSymbol>` under-reports badly (in
`core/mcp`, 60 test functions exist but only one file names `MCPManager`).

```bash
# 1. Changed file's own package -- every test file, not just the matching one
grep -hn "^func Test" <pkg>/*_test.go

# 2. Caller packages -- map Step 4b's caller hits to directories, then enumerate each
grep -rl "<TargetSymbol>" --include='*.go' core/ framework/ transports/ plugins/ \
| xargs -n1 dirname | sort -u \
| while read -r d; do grep -Hn "^func Test" "$d"/*_test.go 2>/dev/null; done
```

Over-inclusion is fine here -- Step 3 narrows the list by measurement. Omission is not:
a dropped candidate never gets coverage-checked and silently disappears from the report.

**Step 3 -- Attribute coverage per test.** Run each candidate with its OWN profile and check
whether it actually reaches the target symbol. A profile from the whole package cannot tell
you this.

Coverage attribution is the ONE place raw `go test` is allowed: no Makefile recipe accepts
`-coverprofile`. This profile is a measurement tool only -- it is never the red/green
evidence in the report. Red and green validation always runs through the Make target
(Step 5f).

```bash
cd <module>
# -coverpkg is REQUIRED: by default a test only instruments its OWN package, so a
# caller-package test would show zero blocks for the target and be wrongly dropped.
go test ./<candidate-pkg>/... -run '^<TestName>$' -covermode=count \
-coverpkg=./... -coverprofile=<scratchpad>/cover-<TestName>.out

# Filter the raw profile to the changed line range. Do NOT use `go tool cover -func`:
# it reports a per-FUNCTION percentage, so a test covering another branch of the same
# function reads as "covered" while your changed line has count=0.
awk -v file="<changed-file>" -v lo=<startLine> -v hi=<endLine> 'NR>1 {
split($1, a, ":"); split(a[2], r, ","); split(r[1], s, "."); split(r[2], e, ".");
if (a[1] ~ file && s[1] <= hi && e[1] >= lo)
printf " lines %s-%s count=%s\n", s[1], e[1], $3
}' <scratchpad>/cover-<TestName>.out
```

Profile line grammar: `<import-path>/file.go:startLine.startCol,endLine.endCol numStmt count`.
Pass `<changed-file>` to `file` exactly as it appears in the diff, with no extra extension:
`~` is a regex match against that import-path-qualified field, so a repo-relative path such as
`bifrost-http/lib/streamreader.go` is the right granularity. A bare `streamreader.go` would
also match a same-named file in any other package `-coverpkg=./...` pulled in, and a doubled
extension (`streamreader.go.go`) matches nothing -- which is indistinguishable from "no test
covers this change" and silently drops every required rerun. The last
field is the execution count. A non-zero count on ANY block overlapping your changed range
means that test executes the code you are changing -- it is a required rerun. All-zero (or
no overlapping block) means it does not.

For UI changes, use the framework's own dependency graph instead of Go coverage:
```bash
cd ui && ./node_modules/.bin/vitest related <changed-file> --run
```
Use the lockfile-pinned binary, not `npx vitest`: with a non-TTY stdin (which is every agent
run) npm assumes `--yes` and silently installs whatever version the registry currently calls
latest when `ui/node_modules` is missing, bypassing `ui/package-lock.json`.

**Step 4 -- Tier the results.** Report reruns in three tiers so the user can see what is
mandatory versus precautionary:

| Tier | Definition | Obligation |
|------|-----------|------------|
| **Tier 1 -- Must rerun** | Coverage shows the test executes a changed line | Blocking. A failure here is caused by your change |
| **Tier 2 -- Should rerun** | Same package, or asserts a contract your change touches, but does not cover the changed line | Run before handing off. Cheap and catches contract drift |
| **Tier 3 -- Downstream** | Tests in packages of callers/dependents identified in Step 4b | Run if Tier 1 or 2 revealed anything, or if the change altered an exported signature or shared invariant |

Explicitly list any test you are NOT rerunning and why (e.g. requires live provider
credentials, unrelated subsystem). Silent omission reads as "everything passed."

Note that provider-gated suites (`make test-core PROVIDER=...`) hit live APIs. If a Tier 1
test is provider-gated, say so -- the user decides whether to spend the call.

### 5f. Write and Run the Failing Tests -- Red (Bug issues only)

For issues classified **Bug** in Step 2a, write the regression test(s) and confirm red
*before* presenting the plan. This is a change from asking permission first: a plan whose
test has not been run is a hypothesis, and reviewers cannot distinguish a real reproduction
from a plausible one.

1. Write the test(s) from the Step 5 Test Plan, targeting the actual root cause from Step 3/4.
Write them against the CURRENT code -- do not write the fix.
2. Run them with the correct Makefile recipe (never raw `go test` for suites that have one).
3. Confirm the failure is for the **expected reason**: the missing behavior, wrong bytes, or
absent field. A compile error, typo, nil panic, or unrelated assertion is NOT a valid red
-- fix the test and rerun until the failure is the real defect.
4. Capture the verbatim failure output. Do not paraphrase or summarize it.
5. **Redact before it leaves the scratchpad.** Provider-gated suites hit live APIs with real
credentials (`OPENAI_API_KEY`, `AWS_SECRET_ACCESS_KEY`, and ~20 others), and upstream
401/403 bodies echoed through `bifrostErr.Error.Message` can carry the key or auth header.
Keep the raw output in `<scratchpad>/` -- never commit it, never paste it. Produce a
redacted transcript for the report, replacing each secret with `[REDACTED: <what>]` so the
removal is visible. Redaction is narrow: it MUST preserve the command, the failing
assertion with actual-vs-expected values, and the stack trace. If redacting would remove
the evidence itself, say so rather than trimming the assertion.

If the test unexpectedly PASSES, stop. The diagnosis in Step 3/4 is wrong or incomplete.
Return to Step 3 and say so plainly in the report rather than adjusting the test until it
fails.

Creating test files is the one write action permitted before plan approval, because the test
is itself the evidence. Do not touch any non-test source file at this stage.

## Step 6: Present Findings

Present everything to the user in this structured format:
Expand Down Expand Up @@ -485,6 +618,41 @@ Copy the research table from Step 3e here. If you skipped Step 3e, go back and d
|------|------|---------------|
| `TestNewScenario` | `<path>` | <scenario description> |

### TDD: Failing Tests (Red)

<Bug issues only. If the issue is a Feature or Docs, write "N/A -- <type> issue, red-before-green
does not apply per AGENTS.md" and omit the rest of this section.>

#### Test Source
<The full source of each new/extended test, as written. Not a description -- the actual code.>

#### Red Output (verbatim, redacted)
```
<paste the failing test output, including the command that produced it. Secrets replaced
with [REDACTED: <what>] -- nothing else altered. Raw output stays in <scratchpad>/.>
```

#### Why This Failure Is the Real Defect
<Explain how the assertion that failed maps to the root cause in the Codebase Analysis. State
explicitly that this is not a compile error, typo, or unrelated panic.>

### Regression Rerun Scope

<From Step 5e. Coverage-attributed, not guessed.>

| Tier | Test | File | Covers Changed Line? | Why It Reruns |
|------|------|------|---------------------|---------------|
| 1 | `TestX` | `<path>` | Yes -- `<symbol>` count N | <what it guards> |
| 2 | `TestY` | `<path>` | No -- same package | <contract it asserts> |
| 3 | `TestZ` | `<path>` | No -- caller package | <downstream invariant> |

**Rerun commands:**
```bash
<exact commands, tier by tier>
```

**Not rerunning:** <list with reasons, or "None -- full scope covered above">

<If changes touch core/>
#### LLM Test Additions
<Specific LLM test recommendations per Section 5b format>
Expand All @@ -504,25 +672,57 @@ Copy the research table from Step 3e here. If you skipped Step 3e, go back and d

---

**Proceed with implementation?** (yes / no / modify plan)
<If Bug>
The failing test above is committed to disk and confirmed red. No source file has been
modified.
</If>
<If Feature or Docs>
No source file has been modified. Red-before-green does not apply to <type> issues per
AGENTS.md; tests are written alongside the implementation in Step 7.
</If>

**Implement the fix now?** (yes / no / modify plan)
```

## Step 7: Implement with Per-Change Approval

Once the user approves the plan:

### 7a. Create a Todo List
### 7a. Bug issues: the red already happened (per AGENTS.md "Testing" section)

For **Bug** issues the failing test was written and confirmed red in Step 5f, before the
approval gate. Step 7 therefore starts at the fix, not the test:

1. Apply the fix changes from the plan (Step 7c below).
2. Re-run the exact same command from the Step 6 "Red Output" block and confirm it now passes.
3. Present red-then-green as a pair: the command, the prior failure, the new pass.
4. Run the Tier 1 and Tier 2 reruns from the Regression Rerun Scope. Report every result,
including any pre-existing failure unrelated to this change -- say so explicitly rather
than omitting it.

Do not modify the test while implementing the fix. If the test needs to change to pass, the
diagnosis was wrong: stop and tell the user rather than editing the test to match the code.

This does not apply to Feature or Docs issues -- AGENTS.md scopes "red before green" to bug
fixes. For those, tests are added per the Test Plan after the implementation, and Step 5f is
skipped.

### 7b. Create a Todo List

Create a todo item for each change in the plan:
Create a todo item for each change in the plan. For Bug issues the regression test already
exists and was confirmed red in Step 5f, before the approval gate -- so it enters the list
already completed and the first pending task is the fix. Do not re-write or modify it (see 7a):
```
1. Change 1: <description> -- pending
2. Change 2: <description> -- pending
3. Update test: <description> -- pending
4. Add new test: <description> -- pending
1. Write failing test: <description> -- completed (Bug issues only; confirmed red in Step 5f)
2. Change 1: <description> -- pending
3. Change 2: <description> -- pending
4. Update existing test: <description> -- pending
5. Verify all tests pass -- pending
```
For Feature and Docs issues, omit item 1 entirely: Step 5f is skipped for those, and the new
tests from the Test Plan are added after the implementation, as their own pending items.

### 7b. For Each Change
### 7c. For Each Change

Before making any edit, present the change to the user:

Expand All @@ -542,7 +742,7 @@ Before making any edit, present the change to the user:

Wait for user approval before applying. If user says "no", skip and move to the next change. If user says "modify", discuss and adjust.

### 7c. After All Changes
### 7d. After All Changes

Once all approved changes are applied:

Expand All @@ -561,7 +761,7 @@ Once all approved changes are applied:
make run-e2e FLOW=<feature>
```

2. Report results to the user
2. Report results to the user, including the red-then-green transcript/summary for Bug issues (what failed before, what passes now)
3. If tests fail, investigate and propose fixes (with approval)

## Error Handling
Expand Down
Loading
Loading