evidence(ai-search): controlled entity-and-offer re-run on 2026-08-12 (Bing, DuckDuckGo) - #128
evidence(ai-search): controlled entity-and-offer re-run on 2026-08-12 (Bing, DuckDuckGo)#128nish3451 wants to merge 1 commit into
Conversation
…ng, DuckDuckGo) on 2026-08-12 Fresh 2026-08-12 captures for the runnable pairs: DuckDuckGo q1/q2/q5 and Bing q1/q5/q7. All six are Wrong or Absent — no Found transition, so the backlog item honestly stays open. The DuckDuckGo q5 answer names tinystudio.io but describes the Mac subtitling app; the Bing and DuckDuckGo organic listings still carry the retired "TinyStudio Agent Desk" title. Google could not be re-run from this host (CAPTCHA on every surface), so the 2026-08-06 Google captures are retained with their original tested dates; nothing was relabelled and no Google run was fabricated. The audit page bundle is regenerated; check, tests and diff --check pass. Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
There was a problem hiding this comment.
nish3451 has reached the 50-credit limit for trial accounts. To continue receiving code reviews, upgrade your plan.
📝 WalkthroughWalkthroughThe pull request refreshes AI-search evidence for the 2026-08-12 controlled rerun. It updates DuckDuckGo and Bing captures, documents blocked runs and validation results, removes superseded fixture records, and synchronizes the embedded audit artifact. ChangesAI-search evidence rerun
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/evidence/ai-search/2026-08-12-controlled-rerun.md`:
- Around line 119-123: Update the verification command fence in the controlled
rerun document to declare the shell language as sh, while preserving the
existing commands unchanged.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 5d1b341b-2d46-43bf-b2f6-dae59694089a
📒 Files selected for processing (3)
docs/evidence/ai-search/2026-08-12-controlled-rerun.mdevidence-fixtures/ai-search/evidence.jsonpublic/audit.html
| ``` | ||
| npm run check | ||
| npm test | ||
| git diff --check | ||
| ``` |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Add a language to the verification fence.
markdownlint-cli2 reports MD040 for this fence. Mark it as sh.
Proposed fix
-```
+```sh
npm run check
npm test
git diff --check</details>
</review_comment>
<!-- suggestion_start -->
<details>
<summary>📝 Committable suggestion</summary>
> ‼️ **IMPORTANT**
> Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
```suggestion
🧰 Tools
🪛 markdownlint-cli2 (0.23.2)
[warning] 119-119: Fenced code blocks should have a language specified
(MD040, fenced-code-language)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/evidence/ai-search/2026-08-12-controlled-rerun.md` around lines 119 -
123, Update the verification command fence in the controlled rerun document to
declare the shell language as sh, while preserving the existing commands
unchanged.
Source: Linters/SAST tools
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ae19de543c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "questionId": "q1-what-tinystudio-does", | ||
| "engine": "duckduckgo", | ||
| "state": "wrong", | ||
| "testedAt": "2026-08-12", |
There was a problem hiding this comment.
Use the actual capture date for the new runs
The commit was created at 2026-08-11 22:08 UTC, while the review environment is still on 2026-08-11 UTC, but all six new runs are marked as captured on 2026-08-12 and the receipt says they were executed then. This publishes future evidence as already verified on /audit; record the actual execution date consistently in the fixture, embedded bundle, and receipt, or perform the rerun on the stated date first.
AGENTS.md reference: AGENTS.md:L2-L3
Useful? React with 👍 / 👎.
| { | ||
| "schema": "ai-search-evidence/v1", | ||
| "testedOn": "2026-08-06", | ||
| "testedOn": "2026-08-12", |
There was a problem hiding this comment.
Expose the retained runs' individual test dates
Only 6 of the 15 runs use the new date; the seven Google runs and two not-tested runs remain dated 2026-08-06. However, public/audit.js renders this field as Controlled test, 2026-08-12 and never renders run.testedAt, so the public artifact presents all 15 records as one August 12 test. Render each run's date or describe this field as a last-updated date with the retained date range instead of advancing the single displayed test date.
AGENTS.md reference: AGENTS.md:L2-L3
Useful? React with 👍 / 👎.
Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
…#128) on current main (#211) * reconcile(evidence): land the stranded 2026-08-12 AI-search re-run on current main PR #128 (lane1/ai-search-rerun-entity-offer-20260812) carried the only fresh 2026-08-12 AI-search captures but went CONFLICTING against main and was stranded unowned, so the live /audit panel kept rendering the 2026-08-09 record. This re-lands the measurement on current main: - evidence-fixtures/ai-search/evidence.json: the 2026-08-12 DuckDuckGo (q1/q2/q5) and Bing (q1/q5/q7) captures, all Wrong or Absent, with testedOn moved to 2026-08-12; the 2026-08-06 Google captures and the not-tested runs are retained. - controlled-questions.json: unchanged - main's q5 ground truth stands (the retired Agent Desk wording PR #128 carried predates PR #43 and must not resurrect). - public/audit.html: embedded AI-search bundle regenerated from fixtures (drift guard passes). - scripts/test-agent-ui.mjs: strict-state set now expects absent, which the 2026-08-12 Bing runs reintroduced. - docs/evidence/ai-search/: original 2026-08-12 receipt with a reconciliation note, plus this pass's reconciliation receipt. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * docs(lane1): add lane report for the PR #128 AI-search reconciliation Co-authored-by: CommandCodeBot <noreply@commandcode.ai> --------- Co-authored-by: nish3451 <nish3451@users.noreply.github.com> Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
|
Closing as already-landed. The 2026-08-12 AI-search controlled re-run was re-landed on current main via commit a654ab4 'reconcile(evidence): land the stranded 2026-08-12 AI-search re-run (PR #128) on current main (#211)'. The reconciliation note is recorded in docs/evidence/ai-search/2026-08-12-controlled-rerun.md. Nothing further to land. |
…intake response (rebase of #116) (#245) * fix(worker): honor the six-a-month intake cap with a truthful closed-intake response Rebase of PR #116 (fix/signup-monthly-cap-lane1) onto current main, resolving conflicts with the daily rate-limit and storage-failure tests added since the branch was cut. The cap counter write fails closed to 503 when D1 is unavailable, so the existing storage-failure honesty tests still pass. - src/worker.js: MAX_APPRAISALS_PER_MONTH constant, monthly cap check in signupResponse, closedIntakeResponse() self-contained 409 page. - scripts/check-site.mjs: static source guards for the cap. - scripts/test-agent-worker.mjs: CountingFakeDB/CountingFakeStatement + 3 tests (JSON cap, browser cap, invalid-email-no-slot). * docs(evidence): triage closeout for parked PRs #116/#128/#137 (2026-08-18) --------- Co-authored-by: minimax-vps <minimax-vps@fleet.local>
…26-08-21 docs(evidence): re-verify the 2026-08-12 AI-search re-run (PR #128) is landed and live (2026-08-21, lane 1)
What this does
Re-establishes the AI-search evidence fixture's verified record on main with fresh controlled captures from 2026-08-12 for every engine reachable from the run environment:
wrongcaptures (verbatim answers + cited URLs)wrong, q5/q7absent)All six fresh runs are
WrongorAbsent— noFoundtransition, so the backlog item "Re-establish verified AI-search entity and offer understanding after the 2026-08-08 15:04 live recheck" honestly stays open.Key findings (2026-08-12)
/sorry/CAPTCHA for the run host. No Google run was fabricated; the last verified Google captures (2026-08-06) are retained with their original tested dates.Files
evidence-fixtures/ai-search/evidence.json— fresh 2026-08-12 runs (DuckDuckGo q1/q2/q5, Bing q1/q5/q7); Google + not-tested runs unchangedpublic/audit.html— embedded AI-search bundle regenerated from fixturesdocs/evidence/ai-search/2026-08-12-controlled-rerun.md— receiptVerification
npm run checkpassesnpm testpasses (92 tests)git diff --checkcleanHonest limits
Summary by CodeRabbit