Skip to content

Public-safe AIOS Habit MVP hardening - #1

Merged
Nakazasen merged 5 commits into
mainfrom
public-mvp-hardening
Jun 20, 2026
Merged

Nakazasen merged 5 commits into
mainfrom
public-mvp-hardening

Conversation

@Nakazasen

@Nakazasen Nakazasen commented Jun 20, 2026 •

Copy link
Copy Markdown
Owner

Hardens AIOS Habit public MVP with privacy guardrails, real phase gates, tests, docs, CI, and Python source formatting integrity.

Summary by CodeRabbit

  • New Features
    • Added the aios-habit CLI (status, discovery, evidence/memory workflows, audits, phase validation, and handover generation).
    • Launched AIOS Habit Studio and AIOS Case Cockpit for managing the local workflow, exports, and audits.
  • Bug Fixes
    • Strengthened privacy/export safety (redaction + local-only handling), including safer asset uploads and more robust graph label rendering.
  • Tests
    • Added/expanded pytest coverage and a CI workflow to run tests on push/PR.
  • Documentation
    • Published governance, architecture, privacy model, runbooks, roadmap, and numerous schemas/templates.

@coderabbitai

coderabbitai Bot commented Jun 20, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

This PR introduces the complete AIOS Habit platform—a local-first, evidence-based personal memory system. It establishes governance and architectural foundations, repository structure with discovery and audit logs, JSON schemas and markdown templates, a core Python package with domain services and validation logic, two Streamlit applications (Studio for memory management and Case Cockpit for case-driven evidence capture), a multi-command CLI, comprehensive test coverage, and GitHub Actions CI.

Changes

AIOS Habit Complete Platform Launch

Layer / File(s) Summary
Governance and design foundations
CONSTITUTION.md, ARCHITECTURE.md, 01_design/*, README.md, docs/CLI_REFERENCE.md, docs/PRIVACY_MODEL.md, docs/RECOVERY_GUIDE.md, docs/OPERATOR_RUNBOOK.md, docs/DATA_MODEL.md
Establishes operating mission, mandatory development rules, and layered architecture; defines local-first governance, evidence-first extraction, PASS/FAIL validation criteria, anti lock-in policy, and design terminology; provides operator guidance for CLI usage, privacy handling, and recovery procedures.
Strategy, roadmap, and inheritance audit
ROADMAP.md, CHANGELOG.md, MANIFEST.md, MASTER_*.md, PROJECT_HANDOVER.md, PRODUCT_NORTH_STAR.md, MIGRATION_POLICY.md, INTEGRATION_DECISION_LOG.md, DEPRECATION_PAUSE_LIST.md, HARVEST_BACKLOG.md, MONDAY_PILOT_CHECKLIST.md, WORKLENS_ARCHITECTURE.md, WORKLENS_MASTER_ROADMAP.md, REPO_INHERITANCE_MAP.md, CASE_COCKPIT_ACCEPTANCE_CRITERIA.md, docs/inheritance_audit/*
Records public-safe MVP roadmap with phase results, changelog milestones, manifest inventory, master profiles, handover artifact; defines product north star, migration tenets, integration decisions, deprecation strategy, harvest backlog with complexity assessment, pilot acceptance criteria, target architecture, master 11-phase roadmap, repository inheritance strategy, and legacy audit findings.
Repository structure and governance artifacts
02_sources/, 03_evidence_registry/, 04_extraction_workspace/, 05_memory_vault/, 06_workflow_library/, 07_ai_export_packs/, 08_audit/, 09_handover/, 12_tools/, _archive/
Populates directory structure with source/project inventories (JSON and markdown), evidence registry README and index, extraction workspace conflict log, memory vault/workflow library/export pack/audit/handover READMEs, phase validation reports (phases 0-9), open issue tracking, validation/rollback/audit logs, phase-0 handover documentation, tool roadmap, and archive guidance.
JSON schemas and markdown templates
10_schemas/*.json, 11_templates/*.md
Introduces JSON Schema contracts (Draft 2020-12) for EvidenceRecord, MemoryUnit, ProjectCard, WorkflowCard, DecisionPattern, and PhaseRecord with field validation and enum constraints; provides reusable markdown templates for audit reports, decision records, evidence records, extraction reports, handovers, memory cards, project cards, and workflow cards with structured placeholders.
Core Python package infrastructure
pyproject.toml, src/aios_habit/__init__.py, src/aios_habit/core.py, src/aios_habit/storage.py, src/aios_habit/models.py, src/aios_habit/paths.py, src/aios_habit/workflow.py
Defines package build config with setuptools and pytest; introduces all five dataclasses (EvidenceRecord, MemoryUnit, ProjectCard, WorkflowCard, DecisionPattern) with required field validation and status/type enums; provides JSONL/JSON persistence helpers, path constants, and module re-exports for public API.
Domain service modules
src/aios_habit/discovery.py, src/aios_habit/evidence.py, src/aios_habit/extraction.py, src/aios_habit/memory.py, src/aios_habit/profiles.py, src/aios_habit/export_pack.py, src/aios_habit/handover.py
Implements filesystem project discovery with depth-bounded walk and ignore-dir filtering; implements evidence record construction with hash computation and local-only export control; implements markdown extraction for candidate memory; implements memory-to-evidence cross-validation enforcing verified memories reference evidence; implements profile markdown generation filtering verified/exportable memory; implements secret redaction and raw-transcript detection for safe exports; implements handover markdown building with evidence/memory counts.
Audit and phase validation
src/aios_habit/audit.py, src/aios_habit/phase_gate.py
Implements audit_repo() with text-file scanning for secret patterns and raw-content markers within export-pack paths, JSONL loading, and verified-memory evidence linkage validation; implements phase_validate() dispatcher with 10 per-phase validators checking file/module presence, schema parsability, model validation behavior, and phase-8 self-test audit execution to ensure audit logic itself works.
CLI command orchestration
src/aios_habit/cli.py
Implements full aios-habit CLI with subcommands: status (counts), discover (project scan), evidence (add/list/validate), memory (add/list/validate/export), extract (markdown-to-candidate), workflow (add/list/validate), decision (add/list/validate), profile build, export (pack generation), audit (repo validation), phase validate, handover build; all routed through argparse with func dispatch and optional dry-run behavior.
Repository safety and continuous integration
.gitignore, .github/workflows/test.yml, tests/test_aios_habit.py
Adds comprehensive .gitignore patterns for Python artifacts, secrets, runtime data (JSONL exports), and Case Cockpit local assets; adds GitHub Actions workflow to run pytest on Windows with Python 3.11; adds pytest test suite covering schema parsing, project discovery, model validation, memory-evidence linkage, audit detection, phase-gate behavior, CLI smoke tests, .gitignore presence checks, and module import verification.
Case data models and local storage
src/aios_habit/case_models.py, src/aios_habit/case_store.py, src/aios_habit/case_ingest.py, src/aios_habit/case_actions.py
Introduces Case and EvidenceItem dataclasses with status, priority, privacy, and timeline tracking; implements JSONL-backed store with init/load/save/upsert and per-case asset directories; implements safe file-name sanitization with traversal prevention and collision avoidance; implements pandas-based Excel/CSV ingestion with row-limiting and error fallback; implements rule-based next-action generation from case state and evidence types.
Case Cockpit UI and hardening tests
src/aios_habit/case_cockpit.py, src/aios_habit/case_graph.py, src/aios_habit/case_audit.py, src/aios_habit/case_prompt.py, tests/test_case_cockpit.py, tests/test_case_cockpit_hardening.py
Implements multi-page Streamlit Case Cockpit with pages for case CRUD, tabbed evidence ingestion (spreadsheet/image/paste/manual), case-map visualization (Mermaid with text fallback), rule-based next-action generation, prompt-pack composition with target-specific filtering, handover markdown preview/edit, and local audit validation; adds mermaid label sanitization, case-audit validation of case/evidence fields with privacy-leakage detection on local-only evidence, target-specific evidence filtering with exclusion notices, and handover markdown assembly; includes pytest coverage for model defaults, CSV ingestion, action generation, launcher verification, prompt privacy rules, Excel ingestion, audit case validation, mermaid character escaping, handover content, storage round-trip, and filename sanitization with collision avoidance.
Studio orchestration and launch infrastructure
src/aios_habit/studio.py, RUN_AIOS_HABIT_STUDIO.bat, RUN_AIOS_CASE_COCKPIT.bat, scripts/run_studio.ps1, scripts/run_case_cockpit.ps1, docs/STUDIO_UI.md, docs/CASE_COCKPIT.md, docs/INSTALL.md
Implements Studio multi-page Streamlit app with pages for dashboard, project discovery, evidence registry (add/list), memory vault (add/list/validate), review queue (candidate approval workflow), master profiles (build and preview), export packs (redact/validate/write), system audit (repo validation and phase reports), handover generation, and settings; provides Windows batch and PowerShell launcher scripts with dependency checking and editable install fallback; documents Studio and Case Cockpit features, launch paths, privacy hardening, install prerequisites, and operator workflows.
```mermaid sequenceDiagram participant Source as Sources participant Studio as Studio UI participant CLI as CLI Commands participant Vault as Memory Vault participant Export as Export Packs participant Audit as Audit &
Phase Gate

Source->>Studio: Add Evidence
Studio->>Vault: Store/Validate
Vault->>Audit: Check linkage
Audit-->>Vault: PASS/FAIL

Vault->>Studio: Build Profiles
Studio->>Export: Render Packs
Export->>Audit: Redact & Scan
Audit-->>Export: Safe/Unsafe

CLI->>Vault: Memory Commands
Vault->>Audit: Validate State
Audit-->>CLI: Report Status

</diagram>

## Estimated code review effort

🎯 4 (Complex) | ⏱️ ~80 minutes

## Poem

> 🐇 *A rabbit built a vault so grand,*
> *With schemas, phases, rules so planned,*
> *No raw transcripts here—just evidence clean,*
> *A Constitution guides what has been seen,*
> *Two studios sing: one for habits, one for cases,*
> *Phase-gated audits lock down all traces,*
> *Local-first forever, no vendor in sight,*
> *PASS before push—the rabbit got it right!* ✨🥕

</details>

<!-- walkthrough_end -->
<!-- pre_merge_checks_walkthrough_start -->

<details>
<summary>🚥 Pre-merge checks | ✅ 4 | ❌ 1</summary>

### ❌ Failed checks (1 warning)

|     Check name     | Status     | Explanation                                                                          | Resolution                                                                         |
| :----------------: | :--------- | :----------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------- |
| Docstring Coverage | ⚠️ Warning | Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. | Write docstrings for the functions missing them to satisfy the coverage threshold. |

<details>
<summary>✅ Passed checks (4 passed)</summary>

|         Check name         | Status   | Explanation                                                                                                                                                                                          |
| :------------------------: | :------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|      Description Check     | ✅ Passed | Check skipped - CodeRabbit’s high-level summary is enabled.                                                                                                                                          |
|         Title check        | ✅ Passed | The PR title accurately summarizes the main change: a comprehensive hardening of the AIOS Habit MVP for public release, focusing on security, testing, documentation, and code quality improvements. |
|     Linked Issues check    | ✅ Passed | Check skipped because no linked issues were found for this pull request.                                                                                                                             |
| Out of Scope Changes check | ✅ Passed | Check skipped because no linked issues were found for this pull request.                                                                                                                             |

</details>

<sub>✏️ Tip: You can configure your own custom pre-merge checks in the settings.</sub>

</details>

<!-- pre_merge_checks_walkthrough_end -->
<!-- finishing_touch_checkbox_start -->

<details>
<summary>✨ Finishing Touches</summary>

<details>
<summary>📝 Generate docstrings</summary>

- [ ] <!-- {"checkboxId": "7962f53c-55bc-4827-bfbf-6a18da830691"} --> Create stacked PR
- [ ] <!-- {"checkboxId": "3e1879ae-f29b-4d0d-8e06-d12b7ba33d98"} --> Commit on current branch

</details>
<details>
<summary>🧪 Generate unit tests (beta)</summary>

- [ ] <!-- {"checkboxId": "f47ac10b-58cc-4372-a567-0e02b2c3d479", "radioGroupId": "utg-output-choice-group-unknown_comment_id"} -->   Create PR with unit tests
- [ ] <!-- {"checkboxId": "6ba7b810-9dad-11d1-80b4-00c04fd430c8", "radioGroupId": "utg-output-choice-group-unknown_comment_id"} -->   Commit unit tests in branch `public-mvp-hardening`

</details>

</details>

<!-- finishing_touch_checkbox_end -->
<!-- tips_start -->

---

Thanks for using [CodeRabbit](https://coderabbit.ai?utm_source=oss&utm_medium=github&utm_campaign=Nakazasen/AIOS_habbit&utm_content=1)! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

<details>
<summary>❤️ Share</summary>

- [X](https://twitter.com/intent/tweet?text=I%20just%20used%20%40coderabbitai%20for%20my%20code%20review%2C%20and%20it%27s%20fantastic%21%20It%27s%20free%20for%20OSS%20and%20offers%20a%20free%20trial%20for%20the%20proprietary%20code.%20Check%20it%20out%3A&url=https%3A//coderabbit.ai)
- [Mastodon](https://mastodon.social/share?text=I%20just%20used%20%40coderabbitai%20for%20my%20code%20review%2C%20and%20it%27s%20fantastic%21%20It%27s%20free%20for%20OSS%20and%20offers%20a%20free%20trial%20for%20the%20proprietary%20code.%20Check%20it%20out%3A%20https%3A%2F%2Fcoderabbit.ai)
- [Reddit](https://www.reddit.com/submit?title=Great%20tool%20for%20code%20review%20-%20CodeRabbit&text=I%20just%20used%20CodeRabbit%20for%20my%20code%20review%2C%20and%20it%27s%20fantastic%21%20It%27s%20free%20for%20OSS%20and%20offers%20a%20free%20trial%20for%20proprietary%20code.%20Check%20it%20out%3A%20https%3A//coderabbit.ai)
- [LinkedIn](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fcoderabbit.ai&mini=true&title=Great%20tool%20for%20code%20review%20-%20CodeRabbit&summary=I%20just%20used%20CodeRabbit%20for%20my%20code%20review%2C%20and%20it%27s%20fantastic%21%20It%27s%20free%20for%20OSS%20and%20offers%20a%20free%20trial%20for%20proprietary%20code)

</details>


<sub>Comment `@coderabbitai help` to get the list of available commands and usage tips.</sub>

<!-- tips_end -->

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 15

Note

Due to the large number of review comments, Critical, Major severity comments were prioritized as inline comments.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (4)
MASTER_PROJECT_INDEX.md (1)

5-70: ⚠️ Potential issue | 🔴 Critical | ⚡ Quick win

CRITICAL: Character encoding corruption in Vietnamese text renders document unreadable.

Lines 5, 9, 15-21, 24-25, 39-40 and throughout contain mojibake (corrupted character encoding). Vietnamese text like "Lニーu danh m盻・c" should be readable Vietnamese. The file appears to have been saved or processed with incorrect UTF-8 encoding (likely interpreted as Latin-1 and re-saved).

This is a data integrity violation per CONSTITUTION.md §6 (raw data policy) and violates source-of-truth requirements. The document is unreadable for Vietnamese speakers and will break any validation expecting valid UTF-8.

🔧 Recommended fix: Re-save file in proper UTF-8 encoding
  1. Open MASTER_PROJECT_INDEX.md in a text editor (e.g., VS Code, Notepad++, Sublime)
  2. Ensure encoding is set to UTF-8 (without BOM) or UTF-8 with BOM (consistent with other files)
  3. Re-save the file
  4. Verify: All Vietnamese text should render correctly (e.g., "Lưu danh mục project được biết" instead of "Lニーu danh m盻・c")

Alternatively, if the source Vietnamese text is available, restore it to this file.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@MASTER_PROJECT_INDEX.md` around lines 5 - 70, The MASTER_PROJECT_INDEX.md
file contains character encoding corruption affecting Vietnamese text throughout
(visible as mojibake like "Lニーu danh m盻・c" instead of proper Vietnamese
characters). Open the file in a text editor, change the file encoding to UTF-8
(without BOM), and re-save it to ensure all Vietnamese text renders correctly
and the document becomes readable. Verify after saving that the text displays
properly (for example, "Lưu danh mục project được biết" should appear instead of
the corrupted version).
08_audit/README.md (1)

1-6: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Translate or remove Vietnamese text for public release consistency.

Lines 3 and 5 contain untranslated Vietnamese:

  • Line 3: "Lưu issue, validation log và rollback log."
  • Line 5: "Audit là bắt buộc trước fix trong các workflow kỹ thuật."

For a public English-primary repository, this harms clarity and maintainability. Translate to English or remove entirely.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@08_audit/README.md` around lines 1 - 6, The README.md file in the 08_audit
directory contains untranslated Vietnamese text on lines 3 and 5 that should be
translated to English for consistency in a public-facing repository. Replace the
Vietnamese text "Lưu issue, validation log và rollback log." and "Audit là bắt
buộc trước fix trong các workflow kỹ thuật." with their English equivalents to
ensure the documentation is clear and maintainable for all users. Alternatively,
remove these lines entirely if they are not essential to understanding the Audit
feature.
07_ai_export_packs/README.md (1)

1-23: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove Vietnamese text for consistency in public release.

Line 3 contains untranslated Vietnamese: "Thư mục này chứa bản chuyển đổi profile/memory cho từng AI." The repository is English-primary; mixing languages without translation harms discoverability and maintenance for the public MVP release.

Translate to English or remove the Vietnamese line entirely.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@07_ai_export_packs/README.md` around lines 1 - 23, The README.md file
contains Vietnamese text on the line following the "AI Export Packs" heading
that is not translated to English. To maintain consistency with the
English-primary repository for the public MVP release, either provide a complete
English translation of the Vietnamese phrase "Thư mục này chứa bản chuyển đổi
profile/memory cho từng AI." or remove the Vietnamese line entirely. Choose
whichever approach best fits the documentation style and ensures the README is
fully comprehensible to English-speaking users.
08_audit/open_issues.md (1)

1-6: ⚠️ Potential issue | 🟡 Minor

Resolve or re-prioritize ISS-0001 before public release.

ISS-0001 is marked OPEN with status Pending review, while phase_0_report.md shows PASS. The technical Phase 0 validation (required files and schemas) passes, but the user review gate remains open. For a public MVP, clarify the status: either complete the user review, defer ISS-0001 to a later phase with explicit target, or re-assess severity if it's truly non-blocking.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@08_audit/open_issues.md` around lines 1 - 6, ISS-0001 in the open_issues.md
table is marked as OPEN with Pending review status, but the Phase 0 validation
has passed, creating ambiguity for public release. Update the ISS-0001 row to
resolve this: either complete the user review and update Status to reflect
closure, or change the Status to explicitly defer ISS-0001 to a future phase
(e.g., Phase 1) with a target date in the Resolution column, or re-assess and
lower the Severity if this is truly non-blocking for the MVP release. Choose the
appropriate action based on product requirements and ensure the table clearly
communicates the decision for stakeholders.
🟡 Minor comments (13)
src/aios_habit/profiles.py-8-10 (1)

8-10: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

KeyError risk: missing guard for 'statement' key.

Line 10 accesses memory['statement'] directly without a guard, which will raise KeyError if the dictionary is missing the 'statement' key. While the CLI pre-filters memory items, this function is a reusable utility that should be defensive.

🛡️ Proposed fix: add defensive guard
     for memory in items:
         evidence = ", ".join(memory.get("evidence_ids", []))
-        body += f"- {memory['statement']} Evidence: {evidence}\n"
+        body += f"- {memory.get('statement', 'UNKNOWN')} Evidence: {evidence}\n"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/profiles.py` around lines 8 - 10, The memory dictionary access
in the loop over items is not defensive. Replace the direct dictionary access
`memory['statement']` with the safe `.get()` method using a default value
(similar to how `memory.get("evidence_ids", [])` is already done on the previous
line) to prevent KeyError when the 'statement' key is missing from the memory
dictionary. This ensures the utility function remains defensive even though the
CLI may pre-filter the data.
src/aios_habit/paths.py-7-19 (1)

7-19: ⚠️ Potential issue | 🟡 Minor

Remove unused IGNORE_DIRS from paths.py.

IGNORE_DIRS is defined in both paths.py (lines 7-19) and discovery.py (lines 5-18) with identical content. However, paths.IGNORE_DIRS is never imported or used anywhere in the codebase, while discovery.IGNORE_DIRS is actively used on lines 39 and 60 of discovery.py. Remove the unused definition from paths.py to eliminate dead code and reduce maintenance burden.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/paths.py` around lines 7 - 19, Remove the unused IGNORE_DIRS
constant definition from paths.py (lines 7-19) since it is never imported or
used anywhere in the codebase and an identical definition already exists and is
actively used in discovery.py. Simply delete the entire IGNORE_DIRS set
definition block along with any associated whitespace to eliminate the dead
code.
src/aios_habit/evidence.py-26-26 (1)

26-26: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Add separator when concatenating strings for hashing.

Concatenating source_path and summary without a delimiter can cause hash collisions when the same characters are split differently. For example, path="abc" + summary="def" produces the same hash as path="ab" + summary="cdef".

🔧 Proposed fix with delimiter
-        hash=sha_text((source_path or "") + (summary or "")),
+        hash=sha_text((source_path or "") + "\n" + (summary or "")),
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/evidence.py` at line 26, The hash calculation in the
sha_text() function call is concatenating source_path and summary without a
delimiter, which can cause hash collisions when the same characters are split
differently between the two strings. Add a separator character (such as a pipe
"|" or colon ":") between the source_path and summary strings in the
concatenation within the sha_text() call to ensure that different combinations
of source_path and summary values produce different hashes.
MASTER_PROJECT_INDEX.md-15-15 (1)

15-15: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Naming inconsistency: "AIOS_habbit" (double-b) vs "AIOS_habit" (single-b).

Line 15 references "Repository m盻・c tiテェu c盻ァa n盻] t蘯」ng memory" for "AIOS_habbit", but other files and the Python package use "AIOS_habit". Confirm the correct project name after resolving the encoding issue.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@MASTER_PROJECT_INDEX.md` at line 15, In MASTER_PROJECT_INDEX.md line 15, fix
the spelling inconsistency by changing "AIOS_habbit" (double-b) to "AIOS_habit"
(single-b) to match the naming convention used in other files and the Python
package. Additionally, address the encoding issue affecting the description text
for this entry by re-encoding it to resolve the garbled characters (currently
showing as "Repository m盻・c tiテェu c盻ァa n盻] t蘯」ng memory") and ensure the text
displays correctly.
02_sources/README.md-1-13 (1)

1-13: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove UTF-8 BOM character from file start.

Same BOM issue as noted in previous files.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@02_sources/README.md` around lines 1 - 13, The file 02_sources/README.md has
a UTF-8 BOM (Byte Order Mark) character at the start of the file before the "#
Sources" heading. Remove this BOM character from the beginning of the file using
your editor's encoding settings or a text processing tool to save the file as
UTF-8 without BOM.
02_sources/README.md-3-13 (1)

3-13: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Clarify language choice: consider consistent English for public repository.

This file mixes Vietnamese and English inconsistently. Headers are in English, but the primary content descriptions are in Vietnamese. For a public MVP release, this could create accessibility barriers for non-Vietnamese-speaking contributors and users.

Recommendation: Either translate consistently to English (preferred for public repo) or provide both languages side-by-side with clear language labels, depending on team intent. If Vietnamese is intentional for a specific audience, document that design choice elsewhere.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@02_sources/README.md` around lines 3 - 13, The README.md file in the
02_sources directory has inconsistent language mixing, with English headers
("Important", "Files") but Vietnamese descriptions in the main content. For a
public repository, translate all the Vietnamese text to English consistently
throughout the file, including the introduction, Important section content, and
the file descriptions to ensure accessibility for non-Vietnamese-speaking
contributors and users.
docs/PRIVACY_MODEL.md-1-4 (1)

1-4: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove UTF-8 BOM character from file start.

Same BOM issue as noted in previous files.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/PRIVACY_MODEL.md` around lines 1 - 4, Remove the UTF-8 BOM (Byte Order
Mark) character that appears at the very beginning of the PRIVACY_MODEL.md file
before the "# Privacy Model" heading. The BOM character is invisible but present
and should be deleted to ensure the file starts cleanly with the heading text.
Use a text editor that can handle BOM removal or manually delete the invisible
character at the file's start.
docs/OPERATOR_RUNBOOK.md-1-7 (1)

1-7: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove UTF-8 BOM character from file start.

Same BOM issue as noted in INSTALL.md.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/OPERATOR_RUNBOOK.md` around lines 1 - 7, The OPERATOR_RUNBOOK.md file
has a UTF-8 BOM (Byte Order Mark) character at the very beginning of the file,
before the "# Operator Runbook" heading. Remove this invisible BOM character
from the start of the file to ensure proper file encoding without the BOM
prefix. You can do this by opening the file in a text editor that supports BOM
removal or by re-saving the file with UTF-8 encoding without BOM.
docs/PHASE_GATE_PROCESS.md-1-2 (1)

1-2: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove UTF-8 BOM character from file start.

Same BOM issue as noted in previous files.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/PHASE_GATE_PROCESS.md` around lines 1 - 2, Remove the UTF-8 BOM (Byte
Order Mark) character that appears at the start of the PHASE_GATE_PROCESS.md
file, before the "#" symbol in the heading "# Phase Gate Process". The file
should start directly with the "#" character without any invisible BOM prefix.
This can typically be done by saving the file with UTF-8 encoding (without BOM)
in your text editor.
docs/RECOVERY_GUIDE.md-1-2 (1)

1-2: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove UTF-8 BOM character from file start.

Same BOM issue as noted in previous files.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/RECOVERY_GUIDE.md` around lines 1 - 2, Remove the UTF-8 BOM (Byte Order
Mark) character from the beginning of the RECOVERY_GUIDE.md file. The BOM
appears as an invisible character before the "# Recovery Guide" heading and
should be deleted. Open the file in your editor, position the cursor at the very
beginning before the "#" character, and delete any invisible characters at the
start of the file. Save the file after removing the BOM to ensure the file
starts directly with the "# Recovery Guide" text without any byte order mark
prefix.
docs/OPERATOR_RUNBOOK.md-4-5 (1)

4-5: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Clarify status terms for unsupported claims.

Line 5 references UNKNOWN or needs_evidence as status values but doesn't define them. These should be explained in context or linked to documentation that defines the memory/evidence state model.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/OPERATOR_RUNBOOK.md` around lines 4 - 5, The OPERATOR_RUNBOOK.md file
references status values `UNKNOWN` and `needs_evidence` for unsupported claims
but does not define what these terms mean. Add clear definitions or explanations
for these status values either inline in the document where they are mentioned
or by including a reference/link to documentation that explains the memory and
evidence state model. Ensure readers understand the distinction between these
states and when each should be used.
docs/INSTALL.md-1-7 (1)

1-7: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Remove UTF-8 BOM character from file start.

The file begins with a UTF-8 byte-order mark (BOM) which can cause parsing issues with some tools, CI systems, and editor configurations. This should be removed.

🔧 Proposed fix

Remove the BOM character from line 1. The file should start directly with # Install without any prefix.

-# Install
+# Install

(Note: In your editor, ensure you're saving as UTF-8 without BOM, or use a tool like sed or dos2unix to strip it.)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/INSTALL.md` around lines 1 - 7, The file starts with a UTF-8 byte-order
mark (BOM) character before the "# Install" heading which can cause parsing
issues with tools and CI systems. Remove the BOM character from the beginning of
the file so that it starts directly with the "# Install" heading with no prefix.
You can do this by saving the file as UTF-8 without BOM in your editor, or use
command-line tools like sed or dos2unix to strip the BOM character.
09_handover/README.md-3-3 (1)

3-3: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Mixed language in documentation.

Line 3 contains Vietnamese text while the rest of the documentation is in English. For consistency and accessibility, consider translating this line to English or establishing a clear i18n strategy for the documentation.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@09_handover/README.md` at line 3, The line containing "Lưu handover theo
phase hoặc theo mốc quan trọng." is written in Vietnamese while the rest of the
README.md documentation is in English, creating inconsistency. Translate this
Vietnamese text to English to maintain language consistency throughout the
documentation. Ensure the translated version maintains the original meaning
about saving handover information according to phases or important milestones.
🧹 Nitpick comments (13)
src/aios_habit/cli.py (7)

204-218: ⚡ Quick win

Add explicit return code for consistency.

This function does not return an explicit exit code. While the implicit None return converts to 0 in main(), adding an explicit return statement improves consistency with other command handlers.

♻️ Proposed fix
     print_json({"status": "PASS", "dry_run": args.dry_run, "verified_exportable": len(memories)})
+    return 0
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 204 - 218, The cmd_profile function does
not have an explicit return statement, which reduces consistency with other
command handlers. Add an explicit return statement at the end of the cmd_profile
function after the print_json call to ensure it returns an explicit exit code
(typically 0 for success) rather than implicitly returning None.

275-281: ⚡ Quick win

Add explicit return code for consistency.

This function does not return an explicit exit code. While the implicit None return converts to 0 in main(), adding an explicit return statement improves consistency with other command handlers.

♻️ Proposed fix
     print_json({"status": "PASS", "dry_run": args.dry_run, "path": "09_handover/final_handover.md"})
+    return 0
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 275 - 281, The cmd_handover function
lacks an explicit return statement at the end, which reduces consistency with
other command handlers in the module. Add an explicit return statement (return 0
or simply return) after the print_json call to make the exit code explicit and
improve consistency across all command handler functions.

144-154: ⚡ Quick win

Add explicit return code for consistency.

This function does not return an explicit exit code. While the implicit None return converts to 0 in main(), adding an explicit return statement improves consistency with other command handlers.

♻️ Proposed fix
     print_json({"status": "PASS", "dry_run": args.dry_run, "candidate": asdict(candidate)})
+    return 0
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 144 - 154, The cmd_extract function does
not include an explicit return statement at the end, which is inconsistent with
other command handlers. Add an explicit return statement at the end of the
cmd_extract function after the print_json call to provide a clear and consistent
exit code for the command handler.

49-60: ⚡ Quick win

Add explicit return code for consistency.

This function does not return an explicit exit code, unlike cmd_evidence, cmd_memory, cmd_export, cmd_audit, and cmd_phase. While Python's implicit None return converts to 0 in main(), explicit return statements improve consistency and clarity.

♻️ Proposed fix
     print_json({"status": "PASS", "dry_run": args.dry_run, "count": len(cards), "projects": data})
+    return 0
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 49 - 60, The cmd_discover function lacks
an explicit return statement for exit code consistency, while other command
functions like cmd_evidence, cmd_memory, cmd_export, cmd_audit, and cmd_phase
all explicitly return 0. Add an explicit return 0 statement at the end of the
cmd_discover function to match the pattern used by the other command functions
and improve code consistency.

63-96: ⚡ Quick win

Standardize validation failure exit codes.

The add subcommand returns exit code 2 on validation failure (line 81), while the validate subcommand returns 1 (line 96). This inconsistency may confuse scripts or automation that check exit codes.

♻️ Proposed fix to standardize on exit code 1
         if errors:
             print_json({"status": "FAIL", "errors": errors})
-            return 2
+            return 1
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 63 - 96, In the cmd_evidence function,
the "add" subcommand returns exit code 2 when validation fails (when errors is
truthy after calling record.validate()), while the "validate" subcommand returns
exit code 1 for validation failures. Change the return statement in the "add"
subcommand's validation failure block from `return 2` to `return 1` to
standardize both subcommands to use the same exit code for validation failures.

99-141: ⚡ Quick win

Standardize validation failure exit codes.

Similar to cmd_evidence, the add subcommand returns exit code 2 on validation failure (line 116), while the validate subcommand returns 1 (line 141). For consistency across all commands, standardize on a single exit code.

♻️ Proposed fix to standardize on exit code 1
         if errors:
             print_json({"status": "FAIL", "errors": errors})
-            return 2
+            return 1
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 99 - 141, The cmd_memory function uses
inconsistent exit codes for validation failures across different subcommands. In
the "add" subcommand branch (args.mem_cmd == "add"), when validation errors are
found, the function returns exit code 2, but the validate subcommand at the end
of the function returns exit code 1 for similar validation failures. Change the
return statement in the "add" branch from return 2 to return 1 to standardize on
exit code 1 for all validation failures throughout the cmd_memory function.

157-201: ⚡ Quick win

Standardize validation failure exit codes.

The add subcommand returns exit code 2 on validation failure (line 187), while the validate subcommand returns 1 (line 201). For consistency with the proposed standardization across all commands, use exit code 1.

♻️ Proposed fix to standardize on exit code 1
         if errors:
             print_json({"status": "FAIL", "errors": errors})
-            return 2
+            return 1
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` around lines 157 - 201, In the handle_generic
function, standardize the validation failure exit code by changing the return
value from 2 to 1 when validation errors occur during the add subcommand. Locate
the line where errors are validated after creating either a WorkflowCard or
DecisionPattern object, and change the return statement from return 2 to return
1 to match the exit code used by the validate subcommand at the end of the
function.
src/aios_habit/storage.py (2)

11-14: ⚖️ Poor tradeoff

Concurrent append risk: append_jsonl lacks atomic write protection.

Multiple concurrent calls to append_jsonl on the same file (e.g., from parallel CLI invocations) can interleave writes and corrupt the JSONL structure. While less likely in a local-first single-user context, the PR objectives emphasize reliability for public MVP.

Consider adding file locking (e.g., fcntl.flock on Unix, msvcrt.locking on Windows) if concurrent writes are expected, or document that the CLI is not safe for concurrent execution.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/storage.py` around lines 11 - 14, The append_jsonl function
lacks file locking protection, which can cause data corruption when multiple
concurrent processes write to the same JSONL file. Add platform-specific file
locking to the function: use fcntl.flock on Unix-based systems and
msvcrt.locking on Windows. Acquire the lock before opening the file in append
mode, perform the write operation, and release the lock afterward.
Alternatively, if concurrent writes are not expected in the design, add clear
documentation stating that the CLI is not safe for concurrent execution.

8-8: ⚡ Quick win

Memory inefficiency: read_text().splitlines() loads entire file.

Using read_text().splitlines() loads the entire JSONL file into memory before parsing. For large evidence or memory registries, this could consume significant memory unnecessarily.

♻️ Recommended refactor: stream line-by-line
 def read_jsonl(path: Path) -> list[dict]:
     if not path.exists():
         return []
-    return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
+    records = []
+    with path.open("r", encoding="utf-8") as f:
+        for line in f:
+            if line.strip():
+                try:
+                    records.append(json.loads(line))
+                except json.JSONDecodeError:
+                    continue
+    return records
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/storage.py` at line 8, The current implementation loads the
entire JSONL file into memory using read_text().splitlines() before parsing each
line, which is inefficient for large files. Replace this approach with streaming
the file line-by-line using a file context manager and the open() function to
iterate through the file, parsing each non-empty line with json.loads() as it's
read. This ensures only one line is held in memory at a time rather than loading
the entire file content at once.
src/aios_habit/paths.py (1)

5-5: ⚖️ Poor tradeoff

REPO_ROOT fragility: Path.cwd() changes with working directory.

Path.cwd() returns the current working directory at module import time, which can vary depending on where the CLI is invoked. This creates unpredictable behavior if the user runs aios-habit from different directories or if code changes the working directory at runtime.

🔧 Recommended fix: anchor to package location or CLI-provided root

Consider one of these alternatives:

  1. Anchor to the package installation directory:
-REPO_ROOT = Path.cwd()
+REPO_ROOT = Path(__file__).parent.parent.parent.resolve()
  1. Or accept repo root as a runtime parameter passed from the CLI (preferred for flexibility).
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/paths.py` at line 5, The REPO_ROOT variable in paths.py uses
Path.cwd() which is fragile because it depends on the current working directory
at import time. Replace the REPO_ROOT assignment with one of two approaches:
either anchor it to the package installation directory by deriving it from the
__file__ path and traversing up the parent directories, or accept repo_root as a
runtime parameter that can be passed from the CLI entry point. The runtime
parameter approach is preferred for flexibility, so consider modifying the paths
module to accept an optional root parameter during initialization that defaults
to deriving from __file__ if not provided.
src/aios_habit/audit.py (1)

27-29: ⚡ Quick win

Inefficient directory skip check.

The check any(part in SKIP_DIRS for part in path.parts) at line 28 iterates through all path parts for every file, which could be slow when scanning large repositories.

♻️ Optional optimization: use set intersection
-        if not path.is_file() or any(part in SKIP_DIRS for part in path.parts):
+        if not path.is_file() or SKIP_DIRS & set(path.parts):
             continue
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/audit.py` around lines 27 - 29, The directory skip check using
`any(part in SKIP_DIRS for part in path.parts)` iterates through all path parts
for every file found by rglob, resulting in poor performance with large
repositories. To fix this, ensure SKIP_DIRS is defined as a set (not a list) and
replace the any() check with a set intersection operation between the path parts
converted to a set and the SKIP_DIRS set. This will change the lookup complexity
from linear to constant time and significantly improve the scan performance for
large directories.
README.md (1)

11-18: ⚡ Quick win

Address adverb repetition in bullet list.

Both the line ending with "memory only." (line 17) and the next line starting with "Exports...only" (line 18) use "only" as a sentence-ender adverb, creating repetitive phrasing. Restructure one sentence to vary the syntax.

✏️ Proposed fix for adverb repetition
 - Builds profiles from verified/export-allowed memory only.
-- Exports AI packs only after redaction/audit checks.
+- Exports AI packs after redaction and audit checks.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@README.md` around lines 11 - 18, The bullet points in the What AIOS Habit
Does section contain repetitive use of the adverb "only" at the end of
consecutive lines - specifically in the line "Builds profiles from
verified/export-allowed memory only." and "Exports AI packs only after
redaction/audit checks." Restructure one of these sentences to vary the syntax
and eliminate the repetition, either by relocating the word "only" to a
different position in the sentence, replacing it with a synonym, or rewording
the clause entirely to maintain clarity while improving readability.
08_audit/phase_0_report.md (1)

1-5: 💤 Low value

Clarify intent of minimal phase report stubs.

Phase_0_report.md contains only a title and status line; no validation details, error descriptions, or summary. Comparing to cli.py lines 20-40 (record_phase_report function), reports are dynamically generated with error details appended.

This stub is either:

  1. A pre-populated audit record (should contain actual check results for user reference)
  2. A placeholder that will be overwritten at runtime (should be marked as auto-generated, or removed)

For clarity, either populate it with a real audit result summary, or explicitly mark it as "auto-generated—do not edit."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@08_audit/phase_0_report.md` around lines 1 - 5, The phase_0_report.md file is
currently a stub with only a title and status, lacking clarity about whether
it's a pre-populated reference document or a runtime-generated placeholder.
Since the record_phase_report function in cli.py (lines 20-40) dynamically
generates reports with error details appended at runtime, clarify the intent by
adding a header comment at the top of phase_0_report.md explicitly marking it as
"auto-generated—do not edit" to indicate this file will be overwritten during
execution, or alternatively populate it with actual validation check results and
error summaries to serve as a reference document for users.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@00_governance/PHASE_0_EXIT_CHECKLIST.md`:
- Around line 1-28: The Phase 0 status is inconsistent between
PHASE_0_EXIT_CHECKLIST.md which shows status PASS with all 18 checks completed,
and PHASE_GATE_LOG.md which shows status OPENED with note "Waiting for user
review" for the same date (2026-06-20). Determine the actual current state of
Phase 0: if it is truly complete and closed, update the PHASE_GATE_LOG.md entry
to reflect the closed status with appropriate closure evidence; if Phase 0 is
still under review and not yet closed, revert the checks in
PHASE_0_EXIT_CHECKLIST.md to BLOCKED or PENDING status to match the
PHASE_GATE_LOG state. Ensure both documents reflect the same phase status for
the same date.

In `@00_governance/VALIDATION_RULES.md`:
- Around line 11-22: The MemoryUnit class in src/aios_habit/core.py and the
validation rules documented in VALIDATION_RULES.md are misaligned. First, update
the MemoryUnit class to match the schema definition in memory_unit.schema.json
by renaming the "category" field to "memory_type" and adding the missing
required fields: "scope", "confidence", and "rollback". Then, expand the
validate() method in MemoryUnit to enforce all 8 documented validation criteria:
unique memory_id, valid memory_type, concise statement (not raw quotes), at
least one evidence record, presence of confidence, presence of boundary/scope,
absence of unlabeled inference, and existence of rollback/deprecation path.
Ensure the updated validation method raises appropriate exceptions or returns
validation status for each criterion to match the governance requirements.

In `@10_schemas/decision_pattern.schema.json`:
- Around line 5-42: The JSON schema field names in decision_pattern.schema.json
do not match the field names in the Python DecisionPattern dataclass, causing
validation and deserialization failures. Update the DecisionPattern dataclass in
src/aios_habit/core.py to rename: the field currently named title to name, the
field currently named criteria to rule, and the field currently named
evidence_ids to evidence. Additionally, add a new required boundary field to the
DecisionPattern dataclass to match the schema requirement. Ensure all field
names and types are consistent between the schema and the Python model.

In `@10_schemas/evidence_record.schema.json`:
- Around line 6-87: The EvidenceRecord dataclass in src/aios_habit/core.py
(lines 59-72) has fields that do not match the evidence_record.schema.json
definition, causing serialized records to be non-compliant. Update the
EvidenceRecord dataclass to match the schema by: removing fields not in the
schema (title, source_path, source_pointer, captured_at, classification, hash,
risk_level, allowed_for_export), adding missing required fields (source_id,
permission_status, retention_policy), and renaming mismatched fields to align
with schema names (source_reference instead of source_path, created_at instead
of captured_at, permission_status instead of classification, artifact_hash
instead of hash). Ensure all required fields from the schema are present and the
field types match the schema definitions.

In `@10_schemas/memory_unit.schema.json`:
- Around line 18-99: Add the missing export_allowed field to the properties
section of the memory_unit schema. Define export_allowed as a boolean type
property with appropriate description documenting that this field can only be
set to true when the memory status is verified. This will align the schema
definition with the Python MemoryUnit model that already uses and validates this
field, ensuring consistency between the data model and the schema validation.

In `@10_schemas/workflow_card.schema.json`:
- Around line 47-51: The schema file defines field names and types that do not
match the WorkflowCard dataclass model. In the workflow_card.schema.json file,
rename the field `outputs` to `output` and change its type from array to string
to match the WorkflowCard model's output field. Additionally, rename the field
`evidence` to `evidence_ids` to match the WorkflowCard model's evidence_ids
field which is a list of strings. These changes must be made in the schema
definitions (around lines 47 and 82) to ensure the schema validation aligns with
the actual WorkflowCard dataclass fields defined in src/aios_habit/core.py.

In `@11_templates/workflow_card.md`:
- Line 13: Replace the Vietnamese placeholder text in the workflow card template
with English equivalents to maintain consistency and professionalism for the
public MVP release. The Vietnamese phrase "Khi nào workflow này được kích hoạt?"
should be replaced with "When is this workflow triggered?" and the Vietnamese
phrase "Nếu workflow sai hoặc gây lỗi, quay lại thế nào?" should be replaced
with an appropriate English equivalent such as "How to rollback if the workflow
fails?" or similar. This ensures the entire template document uses consistent
English language throughout.
- Line 6: The workflow_card.md template uses field names that do not match the
WorkflowCard dataclass contract. Change the field label from "Name:" to "Title:"
to align with the WorkflowCard's title field, and change "Outputs" to "Output:"
to align with the WorkflowCard's singular output field rather than plural
outputs. These corrections ensure the template matches the actual dataclass
fields used during CLI and runtime validation.

In `@src/aios_habit/audit.py`:
- Around line 40-42: The MemoryUnit constructor call in the audit loop lacks
error handling for malformed JSONL records, causing the entire audit to crash
when encountering invalid data. Wrap the MemoryUnit(**record) instantiation in a
try-except block to catch TypeError and ValueError exceptions, and when caught,
append an appropriate error message (such as "memory_id_unknown: Invalid record
format - error_details") to the errors list so the audit continues processing
remaining records and reports all issues comprehensively.
- Line 33: The read_text() method call is using errors="ignore" which silently
suppresses encoding errors and could hide file corruption issues during an audit
operation. Replace the errors="ignore" parameter with errors="replace" to
preserve undecodable characters while still allowing the file to be read, or
alternatively add explicit error handling that logs a warning when encoding
issues are encountered. This will ensure the audit function doesn't miss
potentially corrupted files.

In `@src/aios_habit/cli.py`:
- Line 18: The REPO variable assignment assumes the CLI is always invoked from
the repository root without any validation, which will cause failures if run
from a subdirectory. Add a helper function that checks for the presence of a
repository marker file (such as a .git directory, setup.py, or other identifying
file specific to the project) to validate that the current working directory is
actually the repository root. Call this validation helper before or during the
REPO assignment and raise an appropriate error if the repository root cannot be
found, ensuring users get clear feedback rather than silent failures in file
operations.

In `@src/aios_habit/core.py`:
- Around line 55-56: The load_json() function lacks error handling for two
common failure cases: missing files (FileNotFoundError) and malformed JSON
(JSONDecodeError). Wrap the json.loads() call and path.read_text() operation in
a try-except block to catch both FileNotFoundError and JSONDecodeError
exceptions, and provide appropriate fallback behavior such as returning an empty
dictionary or None when either error occurs, similar to how read_jsonl() handles
missing files by returning an empty list.
- Around line 38-41: The read_jsonl function does not handle JSONDecodeError
that can occur when json.loads(line) encounters malformed JSON content. Wrap the
json.loads(line) call within a try-except block to catch JSONDecodeError
exceptions, and either skip malformed lines silently or log a warning message
before continuing to the next line. This will prevent the function from crashing
when processing JSONL files with corrupted or invalid JSON entries, ensuring all
callers including audit, evidence validation, and CLI commands remain stable.

In `@src/aios_habit/memory.py`:
- Around line 11-13: Wrap the `MemoryUnit(**record)` instantiation in a
try-except block to handle TypeError exceptions that occur when JSONL records
have missing required fields or unexpected keys. When a construction error is
caught, append an error message to the errors list that identifies which record
failed and includes the error details, then continue processing the next record
instead of crashing. This ensures the validation loop completes even when
encountering malformed records in hand-edited JSONL files.

In `@src/aios_habit/storage.py`:
- Around line 5-8: The read_jsonl function lacks error handling for malformed
JSON lines which will cause a JSONDecodeError and crash the entire operation.
Wrap the json.loads(line) call in a try-except block within the list
comprehension to catch JSONDecodeError, and skip malformed lines gracefully
(optionally logging a warning) so that corrupted or hand-edited JSONL files can
still be processed without failing completely.

---

Outside diff comments:
In `@07_ai_export_packs/README.md`:
- Around line 1-23: The README.md file contains Vietnamese text on the line
following the "AI Export Packs" heading that is not translated to English. To
maintain consistency with the English-primary repository for the public MVP
release, either provide a complete English translation of the Vietnamese phrase
"Thư mục này chứa bản chuyển đổi profile/memory cho từng AI." or remove the
Vietnamese line entirely. Choose whichever approach best fits the documentation
style and ensures the README is fully comprehensible to English-speaking users.

In `@08_audit/open_issues.md`:
- Around line 1-6: ISS-0001 in the open_issues.md table is marked as OPEN with
Pending review status, but the Phase 0 validation has passed, creating ambiguity
for public release. Update the ISS-0001 row to resolve this: either complete the
user review and update Status to reflect closure, or change the Status to
explicitly defer ISS-0001 to a future phase (e.g., Phase 1) with a target date
in the Resolution column, or re-assess and lower the Severity if this is truly
non-blocking for the MVP release. Choose the appropriate action based on product
requirements and ensure the table clearly communicates the decision for
stakeholders.

In `@08_audit/README.md`:
- Around line 1-6: The README.md file in the 08_audit directory contains
untranslated Vietnamese text on lines 3 and 5 that should be translated to
English for consistency in a public-facing repository. Replace the Vietnamese
text "Lưu issue, validation log và rollback log." and "Audit là bắt buộc trước
fix trong các workflow kỹ thuật." with their English equivalents to ensure the
documentation is clear and maintainable for all users. Alternatively, remove
these lines entirely if they are not essential to understanding the Audit
feature.

In `@MASTER_PROJECT_INDEX.md`:
- Around line 5-70: The MASTER_PROJECT_INDEX.md file contains character encoding
corruption affecting Vietnamese text throughout (visible as mojibake like "Lニーu
danh m盻・c" instead of proper Vietnamese characters). Open the file in a text
editor, change the file encoding to UTF-8 (without BOM), and re-save it to
ensure all Vietnamese text renders correctly and the document becomes readable.
Verify after saving that the text displays properly (for example, "Lưu danh mục
project được biết" should appear instead of the corrupted version).

---

Minor comments:
In `@02_sources/README.md`:
- Around line 1-13: The file 02_sources/README.md has a UTF-8 BOM (Byte Order
Mark) character at the start of the file before the "# Sources" heading. Remove
this BOM character from the beginning of the file using your editor's encoding
settings or a text processing tool to save the file as UTF-8 without BOM.
- Around line 3-13: The README.md file in the 02_sources directory has
inconsistent language mixing, with English headers ("Important", "Files") but
Vietnamese descriptions in the main content. For a public repository, translate
all the Vietnamese text to English consistently throughout the file, including
the introduction, Important section content, and the file descriptions to ensure
accessibility for non-Vietnamese-speaking contributors and users.

In `@09_handover/README.md`:
- Line 3: The line containing "Lưu handover theo phase hoặc theo mốc quan
trọng." is written in Vietnamese while the rest of the README.md documentation
is in English, creating inconsistency. Translate this Vietnamese text to English
to maintain language consistency throughout the documentation. Ensure the
translated version maintains the original meaning about saving handover
information according to phases or important milestones.

In `@docs/INSTALL.md`:
- Around line 1-7: The file starts with a UTF-8 byte-order mark (BOM) character
before the "# Install" heading which can cause parsing issues with tools and CI
systems. Remove the BOM character from the beginning of the file so that it
starts directly with the "# Install" heading with no prefix. You can do this by
saving the file as UTF-8 without BOM in your editor, or use command-line tools
like sed or dos2unix to strip the BOM character.

In `@docs/OPERATOR_RUNBOOK.md`:
- Around line 1-7: The OPERATOR_RUNBOOK.md file has a UTF-8 BOM (Byte Order
Mark) character at the very beginning of the file, before the "# Operator
Runbook" heading. Remove this invisible BOM character from the start of the file
to ensure proper file encoding without the BOM prefix. You can do this by
opening the file in a text editor that supports BOM removal or by re-saving the
file with UTF-8 encoding without BOM.
- Around line 4-5: The OPERATOR_RUNBOOK.md file references status values
`UNKNOWN` and `needs_evidence` for unsupported claims but does not define what
these terms mean. Add clear definitions or explanations for these status values
either inline in the document where they are mentioned or by including a
reference/link to documentation that explains the memory and evidence state
model. Ensure readers understand the distinction between these states and when
each should be used.

In `@docs/PHASE_GATE_PROCESS.md`:
- Around line 1-2: Remove the UTF-8 BOM (Byte Order Mark) character that appears
at the start of the PHASE_GATE_PROCESS.md file, before the "#" symbol in the
heading "# Phase Gate Process". The file should start directly with the "#"
character without any invisible BOM prefix. This can typically be done by saving
the file with UTF-8 encoding (without BOM) in your text editor.

In `@docs/PRIVACY_MODEL.md`:
- Around line 1-4: Remove the UTF-8 BOM (Byte Order Mark) character that appears
at the very beginning of the PRIVACY_MODEL.md file before the "# Privacy Model"
heading. The BOM character is invisible but present and should be deleted to
ensure the file starts cleanly with the heading text. Use a text editor that can
handle BOM removal or manually delete the invisible character at the file's
start.

In `@docs/RECOVERY_GUIDE.md`:
- Around line 1-2: Remove the UTF-8 BOM (Byte Order Mark) character from the
beginning of the RECOVERY_GUIDE.md file. The BOM appears as an invisible
character before the "# Recovery Guide" heading and should be deleted. Open the
file in your editor, position the cursor at the very beginning before the "#"
character, and delete any invisible characters at the start of the file. Save
the file after removing the BOM to ensure the file starts directly with the "#
Recovery Guide" text without any byte order mark prefix.

In `@MASTER_PROJECT_INDEX.md`:
- Line 15: In MASTER_PROJECT_INDEX.md line 15, fix the spelling inconsistency by
changing "AIOS_habbit" (double-b) to "AIOS_habit" (single-b) to match the naming
convention used in other files and the Python package. Additionally, address the
encoding issue affecting the description text for this entry by re-encoding it
to resolve the garbled characters (currently showing as "Repository m盻・c tiテェu
c盻ァa n盻] t蘯」ng memory") and ensure the text displays correctly.

In `@src/aios_habit/evidence.py`:
- Line 26: The hash calculation in the sha_text() function call is concatenating
source_path and summary without a delimiter, which can cause hash collisions
when the same characters are split differently between the two strings. Add a
separator character (such as a pipe "|" or colon ":") between the source_path
and summary strings in the concatenation within the sha_text() call to ensure
that different combinations of source_path and summary values produce different
hashes.

In `@src/aios_habit/paths.py`:
- Around line 7-19: Remove the unused IGNORE_DIRS constant definition from
paths.py (lines 7-19) since it is never imported or used anywhere in the
codebase and an identical definition already exists and is actively used in
discovery.py. Simply delete the entire IGNORE_DIRS set definition block along
with any associated whitespace to eliminate the dead code.

In `@src/aios_habit/profiles.py`:
- Around line 8-10: The memory dictionary access in the loop over items is not
defensive. Replace the direct dictionary access `memory['statement']` with the
safe `.get()` method using a default value (similar to how
`memory.get("evidence_ids", [])` is already done on the previous line) to
prevent KeyError when the 'statement' key is missing from the memory dictionary.
This ensures the utility function remains defensive even though the CLI may
pre-filter the data.

---

Nitpick comments:
In `@08_audit/phase_0_report.md`:
- Around line 1-5: The phase_0_report.md file is currently a stub with only a
title and status, lacking clarity about whether it's a pre-populated reference
document or a runtime-generated placeholder. Since the record_phase_report
function in cli.py (lines 20-40) dynamically generates reports with error
details appended at runtime, clarify the intent by adding a header comment at
the top of phase_0_report.md explicitly marking it as "auto-generated—do not
edit" to indicate this file will be overwritten during execution, or
alternatively populate it with actual validation check results and error
summaries to serve as a reference document for users.

In `@README.md`:
- Around line 11-18: The bullet points in the What AIOS Habit Does section
contain repetitive use of the adverb "only" at the end of consecutive lines -
specifically in the line "Builds profiles from verified/export-allowed memory
only." and "Exports AI packs only after redaction/audit checks." Restructure one
of these sentences to vary the syntax and eliminate the repetition, either by
relocating the word "only" to a different position in the sentence, replacing it
with a synonym, or rewording the clause entirely to maintain clarity while
improving readability.

In `@src/aios_habit/audit.py`:
- Around line 27-29: The directory skip check using `any(part in SKIP_DIRS for
part in path.parts)` iterates through all path parts for every file found by
rglob, resulting in poor performance with large repositories. To fix this,
ensure SKIP_DIRS is defined as a set (not a list) and replace the any() check
with a set intersection operation between the path parts converted to a set and
the SKIP_DIRS set. This will change the lookup complexity from linear to
constant time and significantly improve the scan performance for large
directories.

In `@src/aios_habit/cli.py`:
- Around line 204-218: The cmd_profile function does not have an explicit return
statement, which reduces consistency with other command handlers. Add an
explicit return statement at the end of the cmd_profile function after the
print_json call to ensure it returns an explicit exit code (typically 0 for
success) rather than implicitly returning None.
- Around line 275-281: The cmd_handover function lacks an explicit return
statement at the end, which reduces consistency with other command handlers in
the module. Add an explicit return statement (return 0 or simply return) after
the print_json call to make the exit code explicit and improve consistency
across all command handler functions.
- Around line 144-154: The cmd_extract function does not include an explicit
return statement at the end, which is inconsistent with other command handlers.
Add an explicit return statement at the end of the cmd_extract function after
the print_json call to provide a clear and consistent exit code for the command
handler.
- Around line 49-60: The cmd_discover function lacks an explicit return
statement for exit code consistency, while other command functions like
cmd_evidence, cmd_memory, cmd_export, cmd_audit, and cmd_phase all explicitly
return 0. Add an explicit return 0 statement at the end of the cmd_discover
function to match the pattern used by the other command functions and improve
code consistency.
- Around line 63-96: In the cmd_evidence function, the "add" subcommand returns
exit code 2 when validation fails (when errors is truthy after calling
record.validate()), while the "validate" subcommand returns exit code 1 for
validation failures. Change the return statement in the "add" subcommand's
validation failure block from `return 2` to `return 1` to standardize both
subcommands to use the same exit code for validation failures.
- Around line 99-141: The cmd_memory function uses inconsistent exit codes for
validation failures across different subcommands. In the "add" subcommand branch
(args.mem_cmd == "add"), when validation errors are found, the function returns
exit code 2, but the validate subcommand at the end of the function returns exit
code 1 for similar validation failures. Change the return statement in the "add"
branch from return 2 to return 1 to standardize on exit code 1 for all
validation failures throughout the cmd_memory function.
- Around line 157-201: In the handle_generic function, standardize the
validation failure exit code by changing the return value from 2 to 1 when
validation errors occur during the add subcommand. Locate the line where errors
are validated after creating either a WorkflowCard or DecisionPattern object,
and change the return statement from return 2 to return 1 to match the exit code
used by the validate subcommand at the end of the function.

In `@src/aios_habit/paths.py`:
- Line 5: The REPO_ROOT variable in paths.py uses Path.cwd() which is fragile
because it depends on the current working directory at import time. Replace the
REPO_ROOT assignment with one of two approaches: either anchor it to the package
installation directory by deriving it from the __file__ path and traversing up
the parent directories, or accept repo_root as a runtime parameter that can be
passed from the CLI entry point. The runtime parameter approach is preferred for
flexibility, so consider modifying the paths module to accept an optional root
parameter during initialization that defaults to deriving from __file__ if not
provided.

In `@src/aios_habit/storage.py`:
- Around line 11-14: The append_jsonl function lacks file locking protection,
which can cause data corruption when multiple concurrent processes write to the
same JSONL file. Add platform-specific file locking to the function: use
fcntl.flock on Unix-based systems and msvcrt.locking on Windows. Acquire the
lock before opening the file in append mode, perform the write operation, and
release the lock afterward. Alternatively, if concurrent writes are not expected
in the design, add clear documentation stating that the CLI is not safe for
concurrent execution.
- Line 8: The current implementation loads the entire JSONL file into memory
using read_text().splitlines() before parsing each line, which is inefficient
for large files. Replace this approach with streaming the file line-by-line
using a file context manager and the open() function to iterate through the
file, parsing each non-empty line with json.loads() as it's read. This ensures
only one line is held in memory at a time rather than loading the entire file
content at once.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

Comment on lines +1 to +28
# Phase 0 Exit Checklist

Phase 0 closed only after concrete validation. The explicit `/goal` execution request is recorded as user approval evidence for this run.

| ID | Check | Status | Evidence / File | Notes |
|---|---|---|---|---|
| P0-01 | Constitution exists and reflects project philosophy | PASS | `CONSTITUTION.md` | Read and aligned |
| P0-02 | Roadmap has required phase fields | PASS | `ROADMAP.md` | Updated to Phase 0-9 |
| P0-03 | Architecture defines layered system and data flow | PASS | `ARCHITECTURE.md` | Existing architecture retained |
| P0-04 | Repository folder structure defined | PASS | `ARCHITECTURE.md` | Existing numbered structure preserved plus compatibility folders |
| P0-05 | Master identity profile exists | PASS | `MASTER_IDENTITY.md` | Candidate claims remain non-verified |
| P0-06 | Master behavior profile exists | PASS | `MASTER_BEHAVIOR_PROFILE.md` | Candidate claims remain non-verified |
| P0-07 | Master language profile exists | PASS | `MASTER_LANGUAGE_PROFILE.md` | Candidate claims remain non-verified |
| P0-08 | Master project index exists | PASS | `MASTER_PROJECT_INDEX.md` | Seed list remains non-exhaustive |
| P0-09 | Master workflow profile exists | PASS | `MASTER_WORKFLOW_PROFILE.md` | Candidate workflows remain non-verified |
| P0-10 | Memory schema exists and parses | PASS | `10_schemas/memory_unit.schema.json` | Parsed with `py -3` |
| P0-11 | Evidence schema exists and parses | PASS | `10_schemas/evidence_record.schema.json` | Parsed with `py -3` |
| P0-12 | Source policy blocks raw chat as direct memory | PASS | `00_governance/SOURCE_POLICY.md` | Confirmed |
| P0-13 | Data policy states local-first principle | PASS | `00_governance/DATA_POLICY.md` | Confirmed |
| P0-14 | Changelog initialized | PASS | `CHANGELOG.md` | Updated |
| P0-15 | Handover initialized | PASS | `PROJECT_HANDOVER.md` | Updated by generator |
| P0-16 | Rollback path exists | PASS | `08_audit/rollback_log.md` | Existing |
| P0-17 | `.gitignore` protects raw/local/secrets | PASS | `.gitignore` | Existing |
| P0-18 | User reviewed and approved execution | PASS | `/goal` request | Treated as explicit approval to execute all phases |

## Current Phase 0 Result

Status: `PASS`

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚖️ Poor tradeoff

Reconcile phase status inconsistency between PHASE_0_EXIT_CHECKLIST and PHASE_GATE_LOG.

This document marks Phase 0 as PASS with all 18 checks completed (line 28), but PHASE_GATE_LOG.md records the same date (2026-06-20) with status OPENED and the note "Waiting for user review" (line 11). These statuses contradict each other.

Clarify whether Phase 0 is actually PASS (in which case update the PHASE_GATE_LOG entry to reflect closure) or still OPENED (in which case revert checks to BLOCKED or PENDING until gate closure conditions are met).

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@00_governance/PHASE_0_EXIT_CHECKLIST.md` around lines 1 - 28, The Phase 0
status is inconsistent between PHASE_0_EXIT_CHECKLIST.md which shows status PASS
with all 18 checks completed, and PHASE_GATE_LOG.md which shows status OPENED
with note "Waiting for user review" for the same date (2026-06-20). Determine
the actual current state of Phase 0: if it is truly complete and closed, update
the PHASE_GATE_LOG.md entry to reflect the closed status with appropriate
closure evidence; if Phase 0 is still under review and not yet closed, revert
the checks in PHASE_0_EXIT_CHECKLIST.md to BLOCKED or PENDING status to match
the PHASE_GATE_LOG state. Ensure both documents reflect the same phase status
for the same date.

Comment on lines +11 to +22
## Memory Validation

Memory đạt PASS khi:

1. Có `memory_id` duy nhất.
2. Có `memory_type` hợp lệ.
3. Có statement ngắn gọn, không phải raw quote dài.
4. Có ít nhất một evidence record.
5. Có confidence.
6. Có boundary/scope.
7. Không chứa suy đoán chưa gắn nhãn.
8. Có rollback/deprecation path.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify MemoryUnit.validate() implementation against documented criteria
ast-grep --pattern $'class MemoryUnit {
  $$$
  def validate($_) {
    $$$
  }
}'

Repository: Nakazasen/AIOS_habbit

Length of output: 47


🏁 Script executed:

# Find core.py and related files
find . -name "core.py" -o -name "*memory*.py" | head -20

Repository: Nakazasen/AIOS_habbit

Length of output: 117


🏁 Script executed:

# Also check the repo structure
git ls-files | grep -E "(core|memory)" | head -20

Repository: Nakazasen/AIOS_habbit

Length of output: 515


🏁 Script executed:

# Check if VALIDATION_RULES.md exists
git ls-files | grep -i validation

Repository: Nakazasen/AIOS_habbit

Length of output: 126


🏁 Script executed:

# Read VALIDATION_RULES.md to see the documented criteria
cat -n 00_governance/VALIDATION_RULES.md

Repository: Nakazasen/AIOS_habbit

Length of output: 1584


🏁 Script executed:

# Check memory.py for MemoryUnit class and validate method
cat -n src/aios_habit/memory.py

Repository: Nakazasen/AIOS_habbit

Length of output: 829


🏁 Script executed:

# Check core.py as well
cat -n src/aios_habit/core.py

Repository: Nakazasen/AIOS_habbit

Length of output: 6937


🏁 Script executed:

# Check the memory_unit.schema.json to see if additional validation is defined there
cat 10_schemas/memory_unit.schema.json

Repository: Nakazasen/AIOS_habbit

Length of output: 2990


Critical gap between documented validation criteria, schema definition, and code implementation.

This document lists 8 criteria for memory PASS status (lines 11-22): unique ID, valid type, concise statement, evidence, confidence, boundary/scope, no unlabeled inference, and rollback path.

However, the MemoryUnit.validate() method in src/aios_habit/core.py only enforces 4 checks: valid category, valid status, evidence requirement for "verified" status, and export-allowed restriction. It does not validate statement conciseness, confidence presence, scope, rollback path, or unlabeled inference.

Additionally, 10_schemas/memory_unit.schema.json defines a schema with required fields (memory_type, scope, confidence, evidence, rollback) that don't match the Python MemoryUnit class fields (which uses "category" instead of "memory_type" and lacks "scope" and "rollback" fields). This schema-code misalignment means neither the schema nor the validation rules are properly enforced.

Memory can pass code validation without meeting documented criteria, creating false confidence. Align the Python implementation with the schema definition, then implement full validation for all 8 documented criteria to support governance robustness.

🧰 Tools
🪛 LanguageTool

[style] ~17-~17: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...uy nhất. 2. Có memory_type hợp lệ. 3. Có statement ngắn gọn, không phải raw quot...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)


[style] ~18-~18: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ... ngắn gọn, không phải raw quote dài. 4. Có ít nhất một evidence record. 5. Có conf...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)


[style] ~19-~19: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: .... 4. Có ít nhất một evidence record. 5. Có confidence. 6. Có boundary/scope. 7. Kh...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)


[style] ~20-~20: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...t evidence record. 5. Có confidence. 6. Có boundary/scope. 7. Không chứa suy đoán ...

(ENGLISH_WORD_REPEAT_BEGINNING_RULE)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@00_governance/VALIDATION_RULES.md` around lines 11 - 22, The MemoryUnit class
in src/aios_habit/core.py and the validation rules documented in
VALIDATION_RULES.md are misaligned. First, update the MemoryUnit class to match
the schema definition in memory_unit.schema.json by renaming the "category"
field to "memory_type" and adding the missing required fields: "scope",
"confidence", and "rollback". Then, expand the validate() method in MemoryUnit
to enforce all 8 documented validation criteria: unique memory_id, valid
memory_type, concise statement (not raw quotes), at least one evidence record,
presence of confidence, presence of boundary/scope, absence of unlabeled
inference, and existence of rollback/deprecation path. Ensure the updated
validation method raises appropriate exceptions or returns validation status for
each criterion to match the governance requirements.

Comment on lines +5 to +42
"required": [
"decision_id",
"name",
"context",
"rule",
"evidence",
"boundary",
"status"
],
"properties": {
"decision_id": {
"type": "string",
"pattern": "^DEC-[0-9]{4,}$"
},
"name": {
"type": "string"
},
"context": {
"type": "string"
},
"rule": {
"type": "string"
},
"tradeoffs": {
"type": "array",
"items": {
"type": "string"
}
},
"boundary": {
"type": "string"
},
"evidence": {
"type": "array",
"items": {
"type": "string"
}
},

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify DecisionPattern dataclass field names

# Find DecisionPattern class definition
ast-grep --pattern $'class DecisionPattern:
  $$$'

# Find DecisionPattern instantiation patterns
rg -nP '\bDecisionPattern\s*\(' -A 3

Repository: Nakazasen/AIOS_habbit

Length of output: 1891


🏁 Script executed:

cat 10_schemas/decision_pattern.schema.json

Repository: Nakazasen/AIOS_habbit

Length of output: 960


Multiple schema-code field name mismatches and missing field.

The schema defines required fields that don't align with the Python DecisionPattern dataclass in src/aios_habit/core.py:160:

  • Schema name ↔ Code title
  • Schema rule ↔ Code criteria
  • Schema evidence ↔ Code evidence_ids
  • Schema requires boundary ↔ Code has no equivalent field

This will cause JSON validation failures and deserialization errors. Align all field names and add the missing boundary field to the Python model.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@10_schemas/decision_pattern.schema.json` around lines 5 - 42, The JSON schema
field names in decision_pattern.schema.json do not match the field names in the
Python DecisionPattern dataclass, causing validation and deserialization
failures. Update the DecisionPattern dataclass in src/aios_habit/core.py to
rename: the field currently named title to name, the field currently named
criteria to rule, and the field currently named evidence_ids to evidence.
Additionally, add a new required boundary field to the DecisionPattern dataclass
to match the schema requirement. Ensure all field names and types are consistent
between the schema and the Python model.

Comment on lines +6 to +87
"required": [
"evidence_id",
"source_id",
"source_type",
"source_reference",
"summary",
"permission_status",
"retention_policy",
"created_at"
],
"properties": {
"evidence_id": {
"type": "string",
"pattern": "^EVD-[0-9]{4,}$"
},
"source_id": {
"type": "string",
"pattern": "^SRC-[0-9]{4,}$"
},
"source_type": {
"type": "string",
"enum": [
"chat_transcript",
"markdown",
"audit_report",
"commit_history",
"roadmap",
"specification",
"project_folder",
"prompt_library",
"user_interview",
"other"
]
},
"source_reference": {
"type": "string",
"description": "Local path, commit hash, artifact ID, or controlled reference."
},
"artifact_hash": {
"type": "string"
},
"summary": {
"type": "string"
},
"content_boundary": {
"type": "string",
"description": "What part of source this evidence covers."
},
"permission_status": {
"type": "string",
"enum": [
"user_provided",
"user_approved",
"local_only",
"excluded",
"unknown"
]
},
"retention_policy": {
"type": "string",
"enum": [
"metadata_only",
"summary_only",
"local_raw_allowed",
"do_not_store_raw",
"delete_after_extraction"
]
},
"linked_memory_ids": {
"type": "array",
"items": {
"type": "string"
}
},
"created_at": {
"type": "string",
"format": "date"
},
"notes": {
"type": "string"
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify EvidenceRecord dataclass field definitions

# Find EvidenceRecord class definition
ast-grep --pattern $'`@dataclass`
class EvidenceRecord:
  $$$'

# Find all EvidenceRecord instantiations
rg -nP '\bEvidenceRecord\s*\(' -A 12 -B 1

Repository: Nakazasen/AIOS_habbit

Length of output: 8424


🏁 Script executed:

# Check how the schema is used in the codebase
rg -l "evidence_record.*schema" -g "*.py" -g "*.json"

# Check if there's any schema validation being performed
rg -n "jsonschema|schema.*validat|validat.*schema" -g "*.py"

# Look at how records are serialized to JSON
rg -n "__dict__|to_dict|json\.dumps" -g "*.py" src/aios_habit/

Repository: Nakazasen/AIOS_habbit

Length of output: 675


🏁 Script executed:

# Check the schema file location and its full content
find . -name "evidence_record.schema.json" -type f

# Verify if schema validation is happening anywhere
rg -n "evidence\.jsonl|EVIDENCE_PATH" -A 5 -B 2 -g "*.py"

Repository: Nakazasen/AIOS_habbit

Length of output: 5622


🏁 Script executed:

# Check the test file to see if schema validation is tested
cat tests/test_aios_habit.py | head -100

Repository: Nakazasen/AIOS_habbit

Length of output: 3160


🏁 Script executed:

# Check if schema files are used for validation anywhere
rg -n "jsonschema|\.schema\.json|validate.*schema" -g "*.py"

# Check phase_gate to see if it validates against schemas
cat src/aios_habit/phase_gate.py | head -150

Repository: Nakazasen/AIOS_habbit

Length of output: 5196


🏁 Script executed:

# Check if there are any imports of jsonschema
rg -n "from jsonschema|import jsonschema" -g "*.py"

# Check requirements or dependencies
cat pyproject.toml 2>/dev/null || cat requirements.txt 2>/dev/null || cat setup.py 2>/dev/null | head -50

Repository: Nakazasen/AIOS_habbit

Length of output: 484


🏁 Script executed:

# Check if schema is referenced in any documentation or comments
rg -n "evidence_record.*schema|schema.*evidence" --include="*.md" --include="*.py"

# Check what the actual dataclass __dict__ produces
python3 << 'PY'
from dataclasses import asdict, dataclass, field
from datetime import datetime

def now():
    return datetime.now().isoformat()

`@dataclass`
class EvidenceRecord:
    evidence_id: str
    title: str
    source_type: str
    source_path: str
    source_pointer: str = ""
    captured_at: str = field(default_factory=now)
    classification: str = "metadata_only"
    summary: str = ""
    hash: str = ""
    risk_level: str = "low"
    allowed_for_export: bool = True
    notes: str = ""

record = EvidenceRecord(
    evidence_id="EVD-1",
    title="Test",
    source_type="markdown",
    source_path="README.md"
)

# Show what gets serialized
import json
serialized = asdict(record)
print("Serialized fields:")
for key in sorted(serialized.keys()):
    print(f"  {key}")
PY

Repository: Nakazasen/AIOS_habbit

Length of output: 701


🏁 Script executed:

# Check if there's any actual JSON schema validation using jsonschema library
rg -n "validate|Draft|validator" src/aios_habit/ -g "*.py" | grep -i schema

# Verify no jsonschema dependency
grep -i "jsonschema" pyproject.toml

Repository: Nakazasen/AIOS_habbit

Length of output: 47


🏁 Script executed:

# Check how records are actually read back from evidence.jsonl
rg -B2 -A5 "EvidenceRecord\(\*\*record\)" src/aios_habit/

Repository: Nakazasen/AIOS_habbit

Length of output: 927


🏁 Script executed:

# Check if core.py or models.py has the definition
head -5 src/aios_habit/core.py src/aios_habit/models.py 2>&1 | grep -A 5 "^==>"

# Check what's being imported from models/core in evidence.py
head -20 src/aios_habit/evidence.py

# Confirm the actual location of EvidenceRecord
grep -n "^from.*EvidenceRecord\|^import.*EvidenceRecord" src/aios_habit/

Repository: Nakazasen/AIOS_habbit

Length of output: 898


🏁 Script executed:

# Verify what happens when deserializing with mismatched fields
python3 << 'PY'
from dataclasses import asdict, dataclass, field

`@dataclass`
class EvidenceRecord:
    evidence_id: str
    title: str
    source_type: str
    source_path: str
    source_pointer: str = ""
    captured_at: str = ""
    classification: str = "metadata_only"
    summary: str = ""
    hash: str = ""
    risk_level: str = "low"
    allowed_for_export: bool = True
    notes: str = ""

# Test 1: Deserialize from JSON that matches code fields
from_code = {
    "evidence_id": "EVD-1",
    "title": "Test",
    "source_type": "markdown",
    "source_path": "README.md",
    "summary": "Summary",
    "classification": "metadata_only",
    "captured_at": "2024-01-01"
}

try:
    r1 = EvidenceRecord(**from_code)
    print("✓ Code-generated fields deserialize successfully")
except Exception as e:
    print(f"✗ Code fields error: {e}")

# Test 2: Deserialize from JSON that matches schema fields
from_schema = {
    "evidence_id": "EVD-1",
    "source_id": "SRC-1234",
    "source_type": "markdown",
    "source_reference": "README.md",
    "summary": "Summary",
    "permission_status": "user_provided",
    "retention_policy": "metadata_only",
    "created_at": "2024-01-01"
}

try:
    r2 = EvidenceRecord(**from_schema)
    print("✓ Schema fields deserialize successfully")
except Exception as e:
    print(f"✗ Schema fields error: {e}")
PY

Repository: Nakazasen/AIOS_habbit

Length of output: 210


🏁 Script executed:

# Check if the schema file is actually referenced anywhere in the codebase
rg -n "evidence_record.schema" -g "*.py" -g "*.md"

# Verify the actual location referenced in the review comment
sed -n '59,96p' src/aios_habit/cli.py

Repository: Nakazasen/AIOS_habbit

Length of output: 1875


🏁 Script executed:

# Check if the review location is accurate - look at cli.py around line 63-96
cat -n src/aios_habit/cli.py | sed -n '59,97p'

Repository: Nakazasen/AIOS_habbit

Length of output: 1874


Schema and code field definitions are misaligned.

The evidence_record.schema.json and EvidenceRecord dataclass (src/aios_habit/core.py:59-72) define incompatible field sets. Records serialized by the code use asdict() and will not conform to the schema structure.

Code provides but schema doesn't define:

  • title
  • source_path (schema expects source_reference)
  • source_pointer
  • captured_at (schema expects created_at)
  • classification (schema expects permission_status)
  • hash (schema has artifact_hash)
  • risk_level
  • allowed_for_export

Schema requires but code doesn't provide:

  • source_id
  • permission_status (mapped from classification)
  • retention_policy

While JSON schema validation is not currently enforced in the codebase (no jsonschema dependency), this contract break indicates the schema is either aspirational or disconnected from the implementation. If the schema represents the intended data model, the EvidenceRecord class must be updated to match.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@10_schemas/evidence_record.schema.json` around lines 6 - 87, The
EvidenceRecord dataclass in src/aios_habit/core.py (lines 59-72) has fields that
do not match the evidence_record.schema.json definition, causing serialized
records to be non-compliant. Update the EvidenceRecord dataclass to match the
schema by: removing fields not in the schema (title, source_path,
source_pointer, captured_at, classification, hash, risk_level,
allowed_for_export), adding missing required fields (source_id,
permission_status, retention_policy), and renaming mismatched fields to align
with schema names (source_reference instead of source_path, created_at instead
of captured_at, permission_status instead of classification, artifact_hash
instead of hash). Ensure all required fields from the schema are present and the
field types match the schema definitions.

Comment on lines +18 to +99
"properties": {
"memory_id": {
"type": "string",
"pattern": "^MEM-[A-Z]+-[0-9]{4,}$"
},
"memory_type": {
"type": "string",
"enum": [
"identity",
"behavior",
"language",
"workflow",
"project_knowledge",
"lessons_learned",
"decision_patterns"
]
},
"title": {
"type": "string",
"minLength": 3
},
"statement": {
"type": "string",
"description": "Pattern/knowledge statement, not raw transcript."
},
"status": {
"type": "string",
"enum": [
"candidate",
"validated",
"deprecated",
"conflicted",
"rejected"
]
},
"confidence": {
"type": "string",
"enum": [
"low",
"medium",
"high",
"verified"
]
},
"scope": {
"type": "string",
"description": "Where this memory applies and does not apply."
},
"evidence": {
"type": "array",
"minItems": 1,
"items": {
"$ref": "#/$defs/evidence_link"
}
},
"tags": {
"type": "array",
"items": {
"type": "string"
}
},
"linked_projects": {
"type": "array",
"items": {
"type": "string"
}
},
"validation": {
"$ref": "#/$defs/validation"
},
"rollback": {
"type": "string"
},
"created_at": {
"type": "string",
"format": "date"
},
"updated_at": {
"type": "string",
"format": "date"
}
},

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Missing export_allowed field definition.

The Python MemoryUnit model in src/aios_habit/phase_gate.py:95 uses and validates an export_allowed field with the constraint "only verified memory can be export_allowed", but this property is not defined in the schema.

This creates a schema-code contract gap where:

  • Python code can create records with export_allowed
  • Schema validation will reject those records as having unexpected properties (if additionalProperties: false is enforced)
  • The documented data model is incomplete

Add the export_allowed boolean property to the schema with appropriate documentation of the verified-status constraint.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@10_schemas/memory_unit.schema.json` around lines 18 - 99, Add the missing
export_allowed field to the properties section of the memory_unit schema. Define
export_allowed as a boolean type property with appropriate description
documenting that this field can only be set to true when the memory status is
verified. This will align the schema definition with the Python MemoryUnit model
that already uses and validates this field, ensuring consistency between the
data model and the schema validation.

Comment thread src/aios_habit/cli.py
from .profiles import build_profile_text
from .storage import append_jsonl, read_jsonl, write_json

REPO = Path.cwd()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Validate that the current directory is the repository root.

REPO = Path.cwd() assumes the CLI is always invoked from the repository root. If a user runs the CLI from a subdirectory, all file operations will fail or write to incorrect locations, degrading the user experience.

🛡️ Proposed fix to add repository validation

Add a helper to verify the presence of a repository marker file:

+def validate_repo(repo: Path) -> None:
+    if not (repo / "CONSTITUTION.md").exists():
+        print_json({"status": "FAIL", "error": "Not in AIOS Habit repository root"})
+        sys.exit(1)
+
 REPO = Path.cwd()
+validate_repo(REPO)
 EVIDENCE_PATH = REPO / "03_evidence_registry" / "records" / "evidence.jsonl"
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/cli.py` at line 18, The REPO variable assignment assumes the
CLI is always invoked from the repository root without any validation, which
will cause failures if run from a subdirectory. Add a helper function that
checks for the presence of a repository marker file (such as a .git directory,
setup.py, or other identifying file specific to the project) to validate that
the current working directory is actually the repository root. Call this
validation helper before or during the REPO assignment and raise an appropriate
error if the repository root cannot be found, ensuring users get clear feedback
rather than silent failures in file operations.

Comment thread src/aios_habit/core.py
Comment on lines +38 to +41
def read_jsonl(path: Path) -> list[dict]:
if not path.exists():
return []
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Add error handling for malformed JSON lines.

If a JSONL file contains an invalid JSON line (e.g., due to manual editing or corruption), json.loads(line) will raise JSONDecodeError and crash. This affects all callers including audit, evidence validation, and CLI commands.

🛡️ Proposed fix with error handling
 def read_jsonl(path: Path) -> list[dict]:
     if not path.exists():
         return []
-    return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
+    results = []
+    for line in path.read_text(encoding="utf-8").splitlines():
+        if line.strip():
+            try:
+                results.append(json.loads(line))
+            except json.JSONDecodeError as e:
+                # Log or skip malformed lines to prevent crashes
+                continue
+    return results
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/core.py` around lines 38 - 41, The read_jsonl function does
not handle JSONDecodeError that can occur when json.loads(line) encounters
malformed JSON content. Wrap the json.loads(line) call within a try-except block
to catch JSONDecodeError exceptions, and either skip malformed lines silently or
log a warning message before continuing to the next line. This will prevent the
function from crashing when processing JSONL files with corrupted or invalid
JSON entries, ensuring all callers including audit, evidence validation, and CLI
commands remain stable.

Comment thread src/aios_habit/core.py
Comment on lines +55 to +56
def load_json(path: Path):
return json.loads(path.read_text(encoding="utf-8"))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Add error handling for missing files and malformed JSON.

load_json() can crash with FileNotFoundError if the file is missing or JSONDecodeError if the JSON is invalid. Unlike read_jsonl(), which returns an empty list for missing files, this function provides no fallback.

🛡️ Proposed fix with error handling
 def load_json(path: Path):
+    if not path.exists():
+        return None  # or raise a more descriptive error
+    try:
-    return json.loads(path.read_text(encoding="utf-8"))
+        return json.loads(path.read_text(encoding="utf-8"))
+    except json.JSONDecodeError as e:
+        raise ValueError(f"Invalid JSON in {path}: {e}")
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
def load_json(path: Path):
return json.loads(path.read_text(encoding="utf-8"))
def load_json(path: Path):
if not path.exists():
return None # or raise a more descriptive error
try:
return json.loads(path.read_text(encoding="utf-8"))
except json.JSONDecodeError as e:
raise ValueError(f"Invalid JSON in {path}: {e}")
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/core.py` around lines 55 - 56, The load_json() function lacks
error handling for two common failure cases: missing files (FileNotFoundError)
and malformed JSON (JSONDecodeError). Wrap the json.loads() call and
path.read_text() operation in a try-except block to catch both FileNotFoundError
and JSONDecodeError exceptions, and provide appropriate fallback behavior such
as returning an empty dictionary or None when either error occurs, similar to
how read_jsonl() handles missing files by returning an empty list.

Comment thread src/aios_habit/memory.py
Comment on lines +11 to +13
for record in read_jsonl(path):
memory = MemoryUnit(**record)
errors.extend(f"{memory.memory_id}: {error}" for error in memory.validate())

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Missing error handling for malformed JSONL records.

MemoryUnit(**record) at line 12 will raise TypeError if the JSONL record is missing required fields or contains unexpected keys, causing the entire validation to crash. Since JSONL files may be hand-edited, this creates a runtime crash risk.

🛡️ Proposed fix: catch construction errors
     for record in read_jsonl(path):
-        memory = MemoryUnit(**record)
+        try:
+            memory = MemoryUnit(**record)
+        except TypeError as exc:
+            errors.append(f"malformed record: {exc}")
+            continue
         errors.extend(f"{memory.memory_id}: {error}" for error in memory.validate())
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/memory.py` around lines 11 - 13, Wrap the
`MemoryUnit(**record)` instantiation in a try-except block to handle TypeError
exceptions that occur when JSONL records have missing required fields or
unexpected keys. When a construction error is caught, append an error message to
the errors list that identifies which record failed and includes the error
details, then continue processing the next record instead of crashing. This
ensures the validation loop completes even when encountering malformed records
in hand-edited JSONL files.

Comment thread src/aios_habit/storage.py
Comment on lines +5 to +8
def read_jsonl(path: Path) -> list[dict]:
if not path.exists():
return []
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Missing error handling for malformed JSON in read_jsonl.

json.loads(line) at line 8 will raise JSONDecodeError if any line contains invalid JSON, causing the entire operation to fail. Since JSONL files may be hand-edited or corrupted, this creates a runtime crash risk.

🛡️ Proposed fix: handle malformed lines gracefully
 def read_jsonl(path: Path) -> list[dict]:
     if not path.exists():
         return []
-    return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
+    records = []
+    for line in path.read_text(encoding="utf-8").splitlines():
+        if not line.strip():
+            continue
+        try:
+            records.append(json.loads(line))
+        except json.JSONDecodeError:
+            # Skip malformed lines or log/raise as needed
+            continue
+    return records
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
def read_jsonl(path: Path) -> list[dict]:
if not path.exists():
return []
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
def read_jsonl(path: Path) -> list[dict]:
if not path.exists():
return []
records = []
for line in path.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError:
# Skip malformed lines or log/raise as needed
continue
return records
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/storage.py` around lines 5 - 8, The read_jsonl function lacks
error handling for malformed JSON lines which will cause a JSONDecodeError and
crash the entire operation. Wrap the json.loads(line) call in a try-except block
within the list comprehension to catch JSONDecodeError, and skip malformed lines
gracefully (optionally logging a warning) so that corrupted or hand-edited JSONL
files can still be processed without failing completely.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 10

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
RUN_AIOS_CASE_COCKPIT.bat (1)

1-14: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Use CRLF line endings for Windows batch reliability.

This .bat file is flagged as LF-only; converting to CRLF avoids known Windows batch parsing edge cases.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@RUN_AIOS_CASE_COCKPIT.bat` around lines 1 - 14, The RUN_AIOS_CASE_COCKPIT.bat
file currently uses LF (Unix-style) line endings, which can cause parsing issues
in Windows batch scripts. Convert all line endings in this file from LF to CRLF
(Windows-style) to ensure proper batch file execution on Windows systems. This
can be done through your text editor's line ending conversion feature or by
using Git configuration to normalize line endings for batch files.

Source: Linters/SAST tools

docs/CASE_COCKPIT.md (1)

17-20: ⚠️ Potential issue | 🟠 Major

Documented command requires package installation step that is omitted from docs.

Line 19 documents py -3 -m streamlit run src\aios_habit\case_cockpit.py, but this command will fail with relative import errors because the package is not installed. The .bat and .ps1 scripts correctly execute py -3 -m pip install -e . first to register the package; the docs omit this critical prerequisite. Either include the installation step in the documentation or reference the entry point defined in pyproject.toml: aios-case-cockpit (which requires package installation).

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/CASE_COCKPIT.md` around lines 17 - 20, The documentation for running the
case cockpit via the direct streamlit command is incomplete because it omits the
required package installation step that the .bat and .ps1 scripts correctly
include. Update the documentation to either include the py -3 -m pip install -e
. installation step before the streamlit run command, or replace it with the
entry point command aios-case-cockpit (which requires the package to be
installed first). Reference the pyproject.toml file where the aios-case-cockpit
entry point is defined to ensure consistency with the project setup.
🧹 Nitpick comments (2)
tests/test_case_cockpit.py (1)

17-23: ⚡ Quick win

Add regression tests for privacy and upload-path hardening paths.

Please add tests that assert filename sanitization in upload persistence and local_only evidence exclusion in prompt construction, so the hardening behavior is locked in.

Also applies to: 24-32

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/test_case_cockpit.py` around lines 17 - 23, The current
test_csv_ingest_reads_synthetic_file test only verifies basic CSV ingest
functionality but lacks regression tests for the security hardening features.
Add additional test functions that specifically assert filename sanitization
behavior in upload persistence (verify that special characters or path traversal
attempts in filenames are properly sanitized before being persisted) and test
that evidence marked with local_only flag is properly excluded from prompt
construction (verify that such evidence does not appear in the extracted text or
other prompt-related outputs). These regression tests will lock in the hardening
behavior and prevent future regressions of these security features.
src/aios_habit/case_models.py (1)

17-17: ⚡ Quick win

Use the shared UTC timestamp helper for persisted model defaults.

These defaults currently use local naive timestamps, while the project already exposes a canonical UTC formatter (src/aios_habit/core.py:26-27). Keeping one format avoids mixed timestamp semantics in JSONL records.

Suggested patch
 from dataclasses import dataclass, field
-from datetime import datetime
 from typing import List, Optional, Dict, Any
+from .core import now
@@
-    created_at: str = field(default_factory=lambda: datetime.now().isoformat())
+    created_at: str = field(default_factory=now)
@@
-    created_at: str = field(default_factory=lambda: datetime.now().isoformat())
-    updated_at: str = field(default_factory=lambda: datetime.now().isoformat())
+    created_at: str = field(default_factory=now)
+    updated_at: str = field(default_factory=now)

Also applies to: 25-26

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_models.py` at line 17, The `created_at` field (and other
timestamp fields on lines 25-26) are using local naive timestamps via
`datetime.now().isoformat()`, which creates inconsistent timestamp formats in
JSONL records. Replace the lambda default factories for all timestamp fields to
use the canonical UTC timestamp helper that is already exposed in
`src/aios_habit/core.py:26-27`, ensuring all persisted model defaults use
consistent UTC timestamp semantics across the codebase.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@RUN_AIOS_CASE_COCKPIT.bat`:
- Around line 5-12: The dependency check on line 5 only verifies third-party
packages (streamlit, pandas, openpyxl) but does not check if the local
aios_habit package is installed. If the third-party packages are already present
but the local package is missing, the errorlevel check passes and line 8 (pip
install -e .) is skipped, causing the streamlit run command on line 12 to fail
at import time. Modify the pip show command to also verify the aios_habit
package is installed, or add a separate verification step that attempts to
import the aios_habit module to ensure both the third-party dependencies and the
local package are available before running the streamlit application.

In `@RUN_AIOS_HABIT_STUDIO.bat`:
- Around line 1-13: The RUN_AIOS_HABIT_STUDIO.bat file currently uses Unix line
endings (LF) which can cause batch parser failures on Windows systems. Convert
the entire file to use Windows line endings (CRLF) by saving the file with CRLF
line ending format in your editor, or using a command-line tool like dos2unix
with the DOS option to perform the conversion. Ensure that all line breaks in
the batch file are properly formatted as CRLF before committing.

In `@scripts/run_case_cockpit.ps1`:
- Around line 5-11: The dependency check at lines 5-8 only verifies that the
streamlit command exists in PATH using Get-Command, but does not confirm that
the aios_habit package is properly installed and importable in the Python
interpreter context that will execute case_cockpit.py. Replace or supplement the
Get-Command check for streamlit with a Python-based verification that attempts
to import both the streamlit module and the aios_habit package using py -3 to
ensure both dependencies are available in the same interpreter instance before
proceeding to launch the application.

In `@scripts/run_studio.ps1`:
- Around line 6-12: The current check only verifies that the streamlit command
exists but does not verify that the aios_habit package is installed, which will
cause the script to fail at runtime when studio.py tries to import from the
aios_habit module. Modify the dependency verification logic to check if both
streamlit and aios_habit are importable using Python import checks, and only
skip the pip install step if both are successfully importable. This ensures that
running py -3 -m streamlit run src\aios_habit\studio.py will have all required
dependencies available.

In `@src/aios_habit/case_cockpit.py`:
- Around line 205-207: The privacy filter for local_only evidence in the loop
currently only excludes content when the target contains "NotebookLM", but
local_only content should be excluded for all external targets. Modify the
condition in the if statement that checks both "NotebookLM" in target and
e.privacy_level == "local_only" to only check e.privacy_level == "local_only" so
that local_only evidence is consistently filtered out regardless of whether the
target is NotebookLM, GPT, Gemini, or Copilot.
- Around line 181-183: The code in the "Add Action" button handler does not
validate that new_action is not empty before appending it to
active_case.next_actions and persisting it with save_case(). Add a validation
check to ensure new_action contains non-empty, non-whitespace content before
proceeding with the append and save operations. Only execute the append and
save_case call if the new_action passes this validation check.

In `@src/aios_habit/case_ingest.py`:
- Around line 64-66: The uploaded filename from uploaded_file.name is used
directly without sanitization, creating a security risk for path traversal
attacks and file overwrite collisions. Extract only the base filename from
uploaded_file.name using os.path.basename() or pathlib.Path.name to remove any
directory path components before constructing the dest_path variable and writing
the file to disk. This ensures that only safe filenames without path separators
are written to the assets_dir directory.

In `@src/aios_habit/case_store.py`:
- Around line 31-32: The bare except Exception: pass blocks at lines 31-32 and
44-45 are silently suppressing parse and schema validation errors when loading
JSONL records, causing data loss without any notification or logging. Replace
these silent exception blocks with proper error handling that logs the specific
error details, includes information about which record failed, and either
re-raises the exception or handles it in a way that makes the failure visible to
the user instead of silently excluding persisted records from the application.
- Around line 60-63: The current implementation of save_cases and the related
code at lines 76-78 write directly to CASES_FILE, which is not atomic and can
leave partial files on interruption or cause data loss when multiple sessions
write concurrently. Refactor the file write operation to first write all JSON
lines to a temporary file, then atomically rename that temporary file to
CASES_FILE to ensure the entire write either succeeds completely or fails
without corrupting the target file.

In `@src/aios_habit/studio.py`:
- Around line 218-226: The approval logic in the button handler checks only that
evidence_ids is non-empty but does not validate that the referenced evidence
actually exists, and it unconditionally appends to MEMORY_PATH without checking
for duplicates. To fix this, before setting status to verified and calling
append_jsonl, add validation to confirm that each evidence_id referenced in data
actually exists in the evidence store, and implement a check to prevent
duplicate entries in MEMORY_PATH by either verifying the record does not already
exist before appending or by replacing an existing record if it does. Ensure the
success message is only shown after successful validation and append.

---

Outside diff comments:
In `@docs/CASE_COCKPIT.md`:
- Around line 17-20: The documentation for running the case cockpit via the
direct streamlit command is incomplete because it omits the required package
installation step that the .bat and .ps1 scripts correctly include. Update the
documentation to either include the py -3 -m pip install -e . installation step
before the streamlit run command, or replace it with the entry point command
aios-case-cockpit (which requires the package to be installed first). Reference
the pyproject.toml file where the aios-case-cockpit entry point is defined to
ensure consistency with the project setup.

In `@RUN_AIOS_CASE_COCKPIT.bat`:
- Around line 1-14: The RUN_AIOS_CASE_COCKPIT.bat file currently uses LF
(Unix-style) line endings, which can cause parsing issues in Windows batch
scripts. Convert all line endings in this file from LF to CRLF (Windows-style)
to ensure proper batch file execution on Windows systems. This can be done
through your text editor's line ending conversion feature or by using Git
configuration to normalize line endings for batch files.

---

Nitpick comments:
In `@src/aios_habit/case_models.py`:
- Line 17: The `created_at` field (and other timestamp fields on lines 25-26)
are using local naive timestamps via `datetime.now().isoformat()`, which creates
inconsistent timestamp formats in JSONL records. Replace the lambda default
factories for all timestamp fields to use the canonical UTC timestamp helper
that is already exposed in `src/aios_habit/core.py:26-27`, ensuring all
persisted model defaults use consistent UTC timestamp semantics across the
codebase.

In `@tests/test_case_cockpit.py`:
- Around line 17-23: The current test_csv_ingest_reads_synthetic_file test only
verifies basic CSV ingest functionality but lacks regression tests for the
security hardening features. Add additional test functions that specifically
assert filename sanitization behavior in upload persistence (verify that special
characters or path traversal attempts in filenames are properly sanitized before
being persisted) and test that evidence marked with local_only flag is properly
excluded from prompt construction (verify that such evidence does not appear in
the extracted text or other prompt-related outputs). These regression tests will
lock in the hardening behavior and prevent future regressions of these security
features.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 5a5fe4f4-9a2f-4109-878c-380725637b82

📥 Commits

Reviewing files that changed from the base of the PR and between c03a932 and 477246d.

📒 Files selected for processing (22)
  • .gitignore
  • CHANGELOG.md
  • PROJECT_HANDOVER.md
  • README.md
  • RUN_AIOS_CASE_COCKPIT.bat
  • RUN_AIOS_HABIT_STUDIO.bat
  • docs/CASE_COCKPIT.md
  • docs/INSTALL.md
  • docs/OPERATOR_RUNBOOK.md
  • docs/STUDIO_UI.md
  • pyproject.toml
  • scripts/run_case_cockpit.ps1
  • scripts/run_studio.ps1
  • src/aios_habit/case_actions.py
  • src/aios_habit/case_cockpit.py
  • src/aios_habit/case_graph.py
  • src/aios_habit/case_ingest.py
  • src/aios_habit/case_models.py
  • src/aios_habit/case_store.py
  • src/aios_habit/studio.py
  • tests/test_aios_habit.py
  • tests/test_case_cockpit.py
✅ Files skipped from review due to trivial changes (5)
  • docs/STUDIO_UI.md
  • PROJECT_HANDOVER.md
  • CHANGELOG.md
  • .gitignore
  • README.md
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/test_aios_habit.py

Comment thread RUN_AIOS_CASE_COCKPIT.bat
Comment on lines +5 to +12
py -3 -m pip show streamlit pandas openpyxl >nul 2>&1
if %errorlevel% neq 0 (
echo Dependencies missing. Installing...
py -3 -m pip install -e .
)

echo Starting AIOS Case Cockpit...
py -3 -m streamlit run src\aios_habit\case_cockpit.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

cd /tmp && find . -maxdepth 3 -type f \( -name "setup.py" -o -name "pyproject.toml" -o -name "setup.cfg" \) 2>/dev/null | head -20

Repository: Nakazasen/AIOS_habbit

Length of output: 47


🏁 Script executed:

ls -la

Repository: Nakazasen/AIOS_habbit

Length of output: 3822


🏁 Script executed:

git ls-files | head -30

Repository: Nakazasen/AIOS_habbit

Length of output: 1065


🏁 Script executed:

cat pyproject.toml

Repository: Nakazasen/AIOS_habbit

Length of output: 650


🏁 Script executed:

ls -la src/

Repository: Nakazasen/AIOS_habbit

Length of output: 239


🏁 Script executed:

cat RUN_AIOS_CASE_COCKPIT.bat

Repository: Nakazasen/AIOS_habbit

Length of output: 388


🏁 Script executed:

find src -type f -name "*.py" | head -20

Repository: Nakazasen/AIOS_habbit

Length of output: 612


Dependency check misses local package installation.

Line 5 only verifies third-party packages (streamlit, pandas, openpyxl). If those are already installed but the local aios_habit package is not, the check passes and Line 8 is skipped, causing Line 12 to fail at import time.

Suggested fix
- py -3 -m pip show streamlit pandas openpyxl >nul 2>&1
+ py -3 -c "import aios_habit, streamlit, pandas, openpyxl" >nul 2>&1
  if %errorlevel% neq 0 (
      echo Dependencies missing. Installing...
      py -3 -m pip install -e .
  )
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@RUN_AIOS_CASE_COCKPIT.bat` around lines 5 - 12, The dependency check on line
5 only verifies third-party packages (streamlit, pandas, openpyxl) but does not
check if the local aios_habit package is installed. If the third-party packages
are already present but the local package is missing, the errorlevel check
passes and line 8 (pip install -e .) is skipped, causing the streamlit run
command on line 12 to fail at import time. Modify the pip show command to also
verify the aios_habit package is installed, or add a separate verification step
that attempts to import the aios_habit module to ensure both the third-party
dependencies and the local package are available before running the streamlit
application.

Comment thread RUN_AIOS_HABIT_STUDIO.bat
Comment on lines +1 to +13
@echo off
cd /d "%~dp0"

echo Checking dependencies...
py -3 -m pip show streamlit >nul 2>&1
if %errorlevel% neq 0 (
echo Streamlit not found. Installing dependencies...
py -3 -m pip install -e .
)

echo Starting AIOS Habit Studio...
py -3 -m streamlit run src\aios_habit\studio.py
pause

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical | ⚡ Quick win

Convert file to Windows line endings (CRLF).

The batch file uses Unix line endings (LF only) which can cause GOTO/CALL label parsing failures on Windows due to documented batch parser boundary bugs. Windows batch files must use CRLF line endings for reliable parsing.

This is not a cosmetic issue—it can cause script malfunction on some Windows systems.

Fix: Convert the file to CRLF using one of these methods:

  • In your editor: Save with "Windows (CRLF)" line ending
  • Via Git config: git config core.autocrlf true to auto-convert on checkout
  • Via command-line tools: dos2unix -c DOS RUN_AIOS_HABIT_STUDIO.bat
🧰 Tools
🪛 Blinter (1.0.112)

[error] 1-1: Unix line endings detected. Explanation: Batch file uses Unix line endings (LF-only) which can cause GOTO/CALL label parsing failures and script malfunction due to Windows batch parser 512-byte boundary bugs. Recommendation: Convert file to Windows line endings (CRLF). Use tools like dos2unix, notepad++, or configure git with 'git config core.autocrlf true'. Context: File uses Unix line endings (LF-only) - 13 LF sequences found

(E018)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@RUN_AIOS_HABIT_STUDIO.bat` around lines 1 - 13, The RUN_AIOS_HABIT_STUDIO.bat
file currently uses Unix line endings (LF) which can cause batch parser failures
on Windows systems. Convert the entire file to use Windows line endings (CRLF)
by saving the file with CRLF line ending format in your editor, or using a
command-line tool like dos2unix with the DOS option to perform the conversion.
Ensure that all line breaks in the batch file are properly formatted as CRLF
before committing.

Comment on lines +5 to +11
if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
Write-Host "Dependencies missing. Installing..."
py -3 -m pip install -e .
}

Write-Host "Starting AIOS Case Cockpit..."
py -3 -m streamlit run src\aios_habit\case_cockpit.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, find and examine the case_cockpit.py file
find . -name "case_cockpit.py" -type f

Repository: Nakazasen/AIOS_habbit

Length of output: 98


🏁 Script executed:

# Look for dependency files
find . -type f \( -name "setup.py" -o -name "setup.cfg" -o -name "pyproject.toml" -o -name "requirements.txt" \) | head -20

Repository: Nakazasen/AIOS_habbit

Length of output: 82


🏁 Script executed:

# Examine the PowerShell script structure
cat -n scripts/run_case_cockpit.ps1

Repository: Nakazasen/AIOS_habbit

Length of output: 530


🏁 Script executed:

# Examine case_cockpit.py imports
cat -n src/aios_habit/case_cockpit.py | head -50

Repository: Nakazasen/AIOS_habbit

Length of output: 2409


🏁 Script executed:

# Examine pyproject.toml for dependencies
cat -n pyproject.toml

Repository: Nakazasen/AIOS_habbit

Length of output: 832


🏁 Script executed:

# Check if aios_habit package structure exists
find src/aios_habit -type f -name "*.py" | head -20

Repository: Nakazasen/AIOS_habbit

Length of output: 612


Launcher should verify all dependencies are importable, not just streamlit command presence.

At lines 5–8, checking Get-Command "streamlit" is insufficient. The case_cockpit.py script imports from the local aios_habit package (lines 8–12), and the launcher can still fail at startup if the package itself isn't installed in that interpreter context, even if the streamlit command exists in PATH.

Suggested fix
-if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
+py -3 -c "import aios_habit, streamlit, pandas, openpyxl" 2>$null
+if ($LASTEXITCODE -ne 0) {
     Write-Host "Dependencies missing. Installing..."
     py -3 -m pip install -e .
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/run_case_cockpit.ps1` around lines 5 - 11, The dependency check at
lines 5-8 only verifies that the streamlit command exists in PATH using
Get-Command, but does not confirm that the aios_habit package is properly
installed and importable in the Python interpreter context that will execute
case_cockpit.py. Replace or supplement the Get-Command check for streamlit with
a Python-based verification that attempts to import both the streamlit module
and the aios_habit package using py -3 to ensure both dependencies are available
in the same interpreter instance before proceeding to launch the application.

Comment thread scripts/run_studio.ps1
Comment on lines +6 to +12
if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
Write-Host "Streamlit not found. Attempting to install dependencies..."
py -3 -m pip install -e .
}

Write-Host "Starting AIOS Habit Studio..."
py -3 -m streamlit run src\aios_habit\studio.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, let's examine the run_studio.ps1 file
cat -n scripts/run_studio.ps1

Repository: Nakazasen/AIOS_habbit

Length of output: 588


🏁 Script executed:

# Find where aios_habit is defined to understand if it's a real package
find . -type f -name "*.py" | head -20 | xargs grep -l "aios_habit" | head -5

Repository: Nakazasen/AIOS_habbit

Length of output: 96


🏁 Script executed:

# Check for setup.py or pyproject.toml to understand package structure
ls -la | grep -E "setup\.(py|cfg)|pyproject\.toml|requirements"

Repository: Nakazasen/AIOS_habbit

Length of output: 130


🏁 Script executed:

# Search for similar dependency checks in other scripts
fd -e "ps1" -o -e "sh" -o -e "bash" | xargs grep -l "pip install" 2>/dev/null | head -5

Repository: Nakazasen/AIOS_habbit

Length of output: 182


🏁 Script executed:

# Check the pyproject.toml to understand the package structure
cat pyproject.toml

Repository: Nakazasen/AIOS_habbit

Length of output: 650


🏁 Script executed:

# Look at the studio.py file to understand what it imports
head -20 src/aios_habit/studio.py

Repository: Nakazasen/AIOS_habbit

Length of output: 755


🏁 Script executed:

# Check if there are other scripts with similar patterns
find . -name "run_*.ps1" -o -name "run_*.sh" | xargs cat 2>/dev/null

Repository: Nakazasen/AIOS_habbit

Length of output: 892


🏁 Script executed:

# Verify the structure of aios_habit package
ls -la src/aios_habit/ | head -15

Repository: Nakazasen/AIOS_habbit

Length of output: 931


Studio launcher needs to verify both aios_habit and streamlit are importable before proceeding.

The current check at line 6 only verifies the streamlit command exists, which is insufficient. If streamlit is installed globally but aios-habit is not installed via pip install -e ., line 12 will fail when trying to import from aios_habit (see imports in studio.py: audit, discovery, export_pack, handover, models, phase_gate, profiles, storage).

Suggested fix
-# Check if streamlit is available
-if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
+py -3 -c "import aios_habit, streamlit" 2>$null
+if ($LASTEXITCODE -ne 0) {
     Write-Host "Streamlit not found. Attempting to install dependencies..."
     py -3 -m pip install -e .
 }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
Write-Host "Streamlit not found. Attempting to install dependencies..."
py -3 -m pip install -e .
}
Write-Host "Starting AIOS Habit Studio..."
py -3 -m streamlit run src\aios_habit\studio.py
py -3 -c "import aios_habit, streamlit" 2>$null
if ($LASTEXITCODE -ne 0) {
Write-Host "Streamlit not found. Attempting to install dependencies..."
py -3 -m pip install -e .
}
Write-Host "Starting AIOS Habit Studio..."
py -3 -m streamlit run src\aios_habit\studio.py
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/run_studio.ps1` around lines 6 - 12, The current check only verifies
that the streamlit command exists but does not verify that the aios_habit
package is installed, which will cause the script to fail at runtime when
studio.py tries to import from the aios_habit module. Modify the dependency
verification logic to check if both streamlit and aios_habit are importable
using Python import checks, and only skip the pip install step if both are
successfully importable. This ensures that running py -3 -m streamlit run
src\aios_habit\studio.py will have all required dependencies available.

Comment thread src/aios_habit/case_cockpit.py Outdated
Comment thread src/aios_habit/case_cockpit.py Outdated
Comment thread src/aios_habit/case_ingest.py Outdated
Comment on lines +31 to +32
except Exception:
pass

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Stop silently dropping invalid JSONL records.

except Exception: pass hides parse/schema failures and silently excludes persisted records from the app, which is a data-loss failure mode.

Suggested patch
+import logging
@@
-                except Exception:
-                    pass
+                except (json.JSONDecodeError, TypeError) as exc:
+                    logging.warning("Skipping invalid case record in %s: %s", CASES_FILE, exc)
@@
-                except Exception:
-                    pass
+                except (json.JSONDecodeError, TypeError) as exc:
+                    logging.warning("Skipping invalid evidence record in %s: %s", EVIDENCE_FILE, exc)

Also applies to: 44-45

🧰 Tools
🪛 Ruff (0.15.17)

[error] 31-32: try-except-pass detected, consider logging the exception

(S110)


[warning] 31-31: Do not catch blind exception: Exception

(BLE001)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_store.py` around lines 31 - 32, The bare except
Exception: pass blocks at lines 31-32 and 44-45 are silently suppressing parse
and schema validation errors when loading JSONL records, causing data loss
without any notification or logging. Replace these silent exception blocks with
proper error handling that logs the specific error details, includes information
about which record failed, and either re-raises the exception or handles it in a
way that makes the failure visible to the user instead of silently excluding
persisted records from the application.

Source: Linters/SAST tools

Comment on lines +60 to +63
with open(CASES_FILE, 'w', encoding='utf-8') as f:
for c in cases:
f.write(json.dumps(asdict(c), ensure_ascii=False) + '\n')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | 🏗️ Heavy lift

Avoid non-atomic whole-file rewrites for JSONL persistence.

Writing directly to the target file can leave partial files on interruption and can lose updates when multiple sessions write close together.

Suggested direction
+import tempfile
@@
+def _write_jsonl_atomic(path: Path, rows: list[dict]) -> None:
+    fd, tmp = tempfile.mkstemp(dir=path.parent, suffix=".tmp")
+    try:
+        with os.fdopen(fd, "w", encoding="utf-8") as f:
+            for row in rows:
+                f.write(json.dumps(row, ensure_ascii=False) + "\n")
+        os.replace(tmp, path)
+    finally:
+        if os.path.exists(tmp):
+            os.remove(tmp)
@@
-    with open(CASES_FILE, 'w', encoding='utf-8') as f:
-        for c in cases:
-            f.write(json.dumps(asdict(c), ensure_ascii=False) + '\n')
+    _write_jsonl_atomic(CASES_FILE, [asdict(c) for c in cases])
@@
-    with open(EVIDENCE_FILE, 'w', encoding='utf-8') as f:
-        for e in items:
-            f.write(json.dumps(asdict(e), ensure_ascii=False) + '\n')
+    _write_jsonl_atomic(EVIDENCE_FILE, [asdict(e) for e in items])

Also applies to: 76-78

🧰 Tools
🪛 ast-grep (0.43.0)

[info] 61-61: use jsonify instead of json.dumps for JSON output
Context: json.dumps(asdict(c), ensure_ascii=False)
Note: Security best practice.

(use-jsonify)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_store.py` around lines 60 - 63, The current
implementation of save_cases and the related code at lines 76-78 write directly
to CASES_FILE, which is not atomic and can leave partial files on interruption
or cause data loss when multiple sessions write concurrently. Refactor the file
write operation to first write all JSON lines to a temporary file, then
atomically rename that temporary file to CASES_FILE to ensure the entire write
either succeeds completely or fails without corrupting the target file.

Comment thread src/aios_habit/studio.py
Comment on lines +218 to +226
if col1.button("Approve as Verified", key=f"app_{c_file.name}"):
if not data.get("evidence_ids"):
st.error("Cannot approve: Missing evidence_ids")
else:
data["status"] = "verified"
data["export_allowed"] = True
append_jsonl(MEMORY_PATH, data)
c_file.write_text(json.dumps(data, indent=2, ensure_ascii=False), encoding="utf-8")
st.success("Approved and added to Memory Vault.")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Review Queue approval can persist invalid and duplicate verified memories.

At Line 219-Line 225, approval only checks that evidence_ids is non-empty, then appends directly. This can (1) mark a memory as verified even when referenced evidence does not exist, and (2) append the same candidate repeatedly on re-approval.

Suggested fix
             if col1.button("Approve as Verified", key=f"app_{c_file.name}"):
-                if not data.get("evidence_ids"):
+                if not data.get("evidence_ids"):
                     st.error("Cannot approve: Missing evidence_ids")
                 else:
-                    data["status"] = "verified"
-                    data["export_allowed"] = True
-                    append_jsonl(MEMORY_PATH, data)
-                    c_file.write_text(json.dumps(data, indent=2, ensure_ascii=False), encoding="utf-8")
-                    st.success("Approved and added to Memory Vault.")
+                    existing_evs = {e.get("evidence_id") for e in load_evidence()}
+                    missing = [e for e in data.get("evidence_ids", []) if e not in existing_evs]
+                    if missing:
+                        st.error(f"Cannot approve: Missing evidence records: {missing}")
+                    else:
+                        data["status"] = "verified"
+                        data["export_allowed"] = True
+                        existing_mem_ids = {m.get("memory_id") for m in load_memory()}
+                        if data.get("memory_id") in existing_mem_ids:
+                            st.warning("Memory already exists in vault; skipping duplicate append.")
+                        else:
+                            append_jsonl(MEMORY_PATH, data)
+                        c_file.write_text(json.dumps(data, indent=2, ensure_ascii=False), encoding="utf-8")
+                        st.success("Approved and added to Memory Vault.")
🧰 Tools
🪛 ast-grep (0.43.0)

[info] 224-224: use jsonify instead of json.dumps for JSON output
Context: json.dumps(data, indent=2, ensure_ascii=False)
Note: Security best practice.

(use-jsonify)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/studio.py` around lines 218 - 226, The approval logic in the
button handler checks only that evidence_ids is non-empty but does not validate
that the referenced evidence actually exists, and it unconditionally appends to
MEMORY_PATH without checking for duplicates. To fix this, before setting status
to verified and calling append_jsonl, add validation to confirm that each
evidence_id referenced in data actually exists in the evidence store, and
implement a check to prevent duplicate entries in MEMORY_PATH by either
verifying the record does not already exist before appending or by replacing an
existing record if it does. Ensure the success message is only shown after
successful validation and append.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/aios_habit/case_audit.py`:
- Around line 31-33: The audit function hardcodes the assets directory path
using Path.cwd() instead of using the store's configured assets root, which
causes inconsistency when running from different working directories. Replace
the hardcoded path construction for the assets_dir variable with a call to
retrieve the assets root from the case store configuration (similar to how
ingest/save operations obtain it), ensuring the audit uses the same asset
location as other store operations regardless of the current working directory.
- Around line 64-68: The leak detection logic in the conditional block checking
privacy_level and prompt_outputs is too simplistic - it only verifies if the
complete ev.extracted_text string appears within prompt_text as a full
substring, but this misses cases where the extracted text is truncated or
fragmented across the prompt. Instead of just checking `ev.extracted_text in
prompt_text`, implement a more robust detection mechanism that can identify
partial or fragmented versions of the sensitive text (such as checking if
substantial portions or key segments of ev.extracted_text appear within
prompt_text, or using fuzzy matching to catch truncated leaks), ensuring that
both complete and partial exposures of local_only evidence are properly caught
and flagged in the errors list.
- Around line 55-61: The path containment validation in the try block uses
string prefix matching with startswith() which is unreliable and can be bypassed
by similarly-prefixed sibling directories. Replace the string-based containment
check with proper path operations using Path methods like is_relative_to() or by
comparing the resolved paths using Path.relative_to() within a try-except block.
Additionally, replace the overly broad Exception catch with specific exception
types that Path operations can actually raise, such as ValueError or OSError, to
properly handle only the expected error conditions when resolving or validating
paths.

In `@src/aios_habit/case_graph.py`:
- Around line 31-35: The evidence node ID generation on line 34 uses raw
ev.evidence_id directly, which is inconsistent with the synthetic ID pattern
used for hypotheses, actions, and decisions elsewhere in the function. Replace
the raw ev.evidence_id in the ev_node assignment with a synthetic ID generated
via the nid function (similar to how hypotheses, actions, and decisions use
indexed synthetic IDs like nid("HYP"), nid("ACT"), nid("DEC")). This ensures all
node identifiers follow a consistent, safe alphanumeric format that aligns with
Mermaid best practices and prevents fragility if ID formats change in the
future.

In `@src/aios_habit/case_prompt.py`:
- Around line 19-27: The evidence items being processed in the loop are not
filtered by the active case's case_id, which can cause evidence from other cases
to leak into the prompt. Before iterating through evidence_items in the loop
that starts with "for e in evidence_items:", add a filter to only include items
where e.case_id matches the current case's case_id. This filtering should happen
before the loop begins so that only evidence belonging to the active case is
included in the prompt assembly.
- Around line 8-13: The cloud targets (gemini, gpt, copilot) in the conditional
statement currently allow bypassing local-only protections by setting
should_include_local_only equal to the include_local_only parameter. To fix this
security issue, change the assignment for cloud targets (gemini, gpt, copilot)
from should_include_local_only = include_local_only to should_include_local_only
= False, ensuring local-only protected content cannot be sent to cloud services
regardless of the include_local_only parameter value, consistent with the
privacy hardening already applied to notebooklm_safe.

In `@tests/test_case_cockpit_hardening.py`:
- Around line 154-155: The bat_content and ps1_content assignments on lines
154-155 use relative paths that depend on the current working directory, causing
failures when tests run from different locations. Make these paths absolute by
anchoring them to the test file location using __file__ and constructing paths
relative to the test file's directory, so the Path().read_text() calls work
regardless of the invocation directory.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: a80618ce-e3fc-4646-90ed-59ccfebea0ed

📥 Commits

Reviewing files that changed from the base of the PR and between 477246d and c59a363.

📒 Files selected for processing (23)
  • CASE_COCKPIT_ACCEPTANCE_CRITERIA.md
  • CHANGELOG.md
  • DEPRECATION_PAUSE_LIST.md
  • HARVEST_BACKLOG.md
  • INTEGRATION_DECISION_LOG.md
  • MIGRATION_POLICY.md
  • MONDAY_PILOT_CHECKLIST.md
  • PRODUCT_NORTH_STAR.md
  • PROJECT_HANDOVER.md
  • REPO_INHERITANCE_MAP.md
  • WORKLENS_ARCHITECTURE.md
  • WORKLENS_MASTER_ROADMAP.md
  • docs/CASE_COCKPIT.md
  • docs/inheritance_audit/ABW_NVIDIA_FUSION_CONTROL_AUDIT.md
  • docs/inheritance_audit/INHERITANCE_SUMMARY.md
  • docs/inheritance_audit/Nvidia_AUDIT.md
  • docs/inheritance_audit/skill_Anti_brain_wiki_note_AUDIT.md
  • src/aios_habit/case_audit.py
  • src/aios_habit/case_cockpit.py
  • src/aios_habit/case_graph.py
  • src/aios_habit/case_ingest.py
  • src/aios_habit/case_prompt.py
  • tests/test_case_cockpit_hardening.py
✅ Files skipped from review due to trivial changes (11)
  • DEPRECATION_PAUSE_LIST.md
  • MIGRATION_POLICY.md
  • INTEGRATION_DECISION_LOG.md
  • HARVEST_BACKLOG.md
  • docs/inheritance_audit/INHERITANCE_SUMMARY.md
  • WORKLENS_ARCHITECTURE.md
  • PRODUCT_NORTH_STAR.md
  • MONDAY_PILOT_CHECKLIST.md
  • REPO_INHERITANCE_MAP.md
  • PROJECT_HANDOVER.md
  • CHANGELOG.md
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/aios_habit/case_cockpit.py

Comment on lines +31 to +33
cwd = Path.cwd().resolve()
assets_dir = (cwd / "local_cases" / "assets").resolve()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Audit asset root should use store configuration, not current working directory.

Lines 31-33 hardcode Path.cwd()/local_cases/assets, but ingest/save uses the case store assets root. Running from a different CWD can produce false FAIL audits.

Suggested fix
 from pathlib import Path
 from typing import List, Dict, Optional, Any
 from .case_models import Case, EvidenceItem
+from .case_store import ASSETS_DIR
@@
-    cwd = Path.cwd().resolve()
-    assets_dir = (cwd / "local_cases" / "assets").resolve()
+    assets_dir = ASSETS_DIR.resolve()
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_audit.py` around lines 31 - 33, The audit function
hardcodes the assets directory path using Path.cwd() instead of using the
store's configured assets root, which causes inconsistency when running from
different working directories. Replace the hardcoded path construction for the
assets_dir variable with a call to retrieve the assets root from the case store
configuration (similar to how ingest/save operations obtain it), ensuring the
audit uses the same asset location as other store operations regardless of the
current working directory.

Comment on lines +55 to +61
if ev.source_path and ev.source_path not in ("clipboard", "manual"):
try:
p = Path(ev.source_path).resolve()
if not str(p).startswith(str(assets_dir)):
errors.append(f"Evidence item {ev.evidence_id} asset path '{ev.source_path}' is outside assets directory '{assets_dir}'.")
except Exception as e:
errors.append(f"Evidence item {ev.evidence_id} has invalid asset path: {e}")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

cd /tmp && find . -name "case_audit.py" -type f 2>/dev/null | head -5

Repository: Nakazasen/AIOS_habbit

Length of output: 47


🏁 Script executed:

git ls-files | grep -i case_audit

Repository: Nakazasen/AIOS_habbit

Length of output: 94


🏁 Script executed:

cat -n src/aios_habit/case_audit.py | sed -n '50,70p'

Repository: Nakazasen/AIOS_habbit

Length of output: 1588


🏁 Script executed:

cat -n src/aios_habit/case_audit.py | head -40

Repository: Nakazasen/AIOS_habbit

Length of output: 1737


🏁 Script executed:

cat -n src/aios_habit/case_audit.py | sed -n '25,65p'

Repository: Nakazasen/AIOS_habbit

Length of output: 2574


Use path-aware containment checks instead of string prefix matching.

Line 58 uses startswith() on path strings, which can be bypassed by prefix-like siblings (e.g., /local_cases/assets_backup/ would match if checking only string prefix). This is not robust path validation. Additionally, line 60 catches a blind Exception instead of the specific exceptions that should be handled.

Suggested fix
         if ev.source_path and ev.source_path not in ("clipboard", "manual"):
             try:
                 p = Path(ev.source_path).resolve()
-                if not str(p).startswith(str(assets_dir)):
-                    errors.append(f"Evidence item {ev.evidence_id} asset path '{ev.source_path}' is outside assets directory '{assets_dir}'.")
-            except Exception as e:
+                p.relative_to(assets_dir)
+            except ValueError:
+                errors.append(
+                    f"Evidence item {ev.evidence_id} asset path '{ev.source_path}' is outside assets directory '{assets_dir}'."
+                )
+            except OSError as e:
                 errors.append(f"Evidence item {ev.evidence_id} has invalid asset path: {e}")
🧰 Tools
🪛 Ruff (0.15.17)

[warning] 60-60: Do not catch blind exception: Exception

(BLE001)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_audit.py` around lines 55 - 61, The path containment
validation in the try block uses string prefix matching with startswith() which
is unreliable and can be bypassed by similarly-prefixed sibling directories.
Replace the string-based containment check with proper path operations using
Path methods like is_relative_to() or by comparing the resolved paths using
Path.relative_to() within a try-except block. Additionally, replace the overly
broad Exception catch with specific exception types that Path operations can
actually raise, such as ValueError or OSError, to properly handle only the
expected error conditions when resolving or validating paths.

Source: Linters/SAST tools

Comment on lines +64 to +68
if ev.privacy_level == "local_only" and prompt_outputs:
for target, prompt_text in prompt_outputs.items():
if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"):
if has_text and ev.extracted_text in prompt_text:
errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Leak detection misses partial local-only text exposure.

Line 67 checks only full-string inclusion (ev.extracted_text in prompt_text). If prompt generation includes a truncated snippet, this audit can pass despite leakage.

Suggested fix
                 if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"):
-                    if has_text and ev.extracted_text in prompt_text:
+                    leak_probe = (ev.extracted_text or "")[:200]
+                    if has_text and leak_probe and leak_probe in prompt_text:
                         errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.")
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if ev.privacy_level == "local_only" and prompt_outputs:
for target, prompt_text in prompt_outputs.items():
if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"):
if has_text and ev.extracted_text in prompt_text:
errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.")
if ev.privacy_level == "local_only" and prompt_outputs:
for target, prompt_text in prompt_outputs.items():
if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"):
leak_probe = (ev.extracted_text or "")[:200]
if has_text and leak_probe and leak_probe in prompt_text:
errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.")
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_audit.py` around lines 64 - 68, The leak detection logic
in the conditional block checking privacy_level and prompt_outputs is too
simplistic - it only verifies if the complete ev.extracted_text string appears
within prompt_text as a full substring, but this misses cases where the
extracted text is truncated or fragmented across the prompt. Instead of just
checking `ev.extracted_text in prompt_text`, implement a more robust detection
mechanism that can identify partial or fragmented versions of the sensitive text
(such as checking if substantial portions or key segments of ev.extracted_text
appear within prompt_text, or using fuzzy matching to catch truncated leaks),
ensuring that both complete and partial exposures of local_only evidence are
properly caught and flagged in the errors list.

Comment on lines +31 to +35
for ev in evidence_items:
if ev.case_id == case.case_id:
label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}"
ev_node = f'{ev.evidence_id}["{label}"]'
lines.append(f'C -->|contains| {ev_node}')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, let's examine the actual file and context around the mentioned lines
cat -n src/aios_habit/case_graph.py | head -60

Repository: Nakazasen/AIOS_habbit

Length of output: 2300


🏁 Script executed:

# Find the safe_mermaid_label function definition
rg "def safe_mermaid_label" -A 10 src/

Repository: Nakazasen/AIOS_habbit

Length of output: 915


🏁 Script executed:

# Check how evidence_id is defined/used
rg "evidence_id" src/aios_habit/ -B 2 -A 2

Repository: Nakazasen/AIOS_habbit

Length of output: 21358


🏁 Script executed:

# Look at the full generate_case_mermaid function to understand the complete context
rg "def generate_case_mermaid" -A 50 src/aios_habit/case_graph.py

Repository: Nakazasen/AIOS_habbit

Length of output: 1098


🏁 Script executed:

# Find the nid() function definition
rg "def nid" -A 5 src/

Repository: Nakazasen/AIOS_habbit

Length of output: 381


🌐 Web query:

Mermaid node ID syntax requirements restrictions characters allowed

💡 Result:

In Mermaid flowchart syntax, node IDs serve as the internal identifiers for nodes and are distinct from their displayed labels [1]. While Mermaid does not strictly define a single character set for IDs, the parser is sensitive to certain keywords and patterns that can cause rendering errors or conflicts [2][3][1][4]. Core Requirements and Restrictions: Reserved Words: You should avoid using Mermaid reserved words as bare node IDs, as these can conflict with the parser and cause the diagram to break [4][5]. Reserved words include: end, graph, subgraph, click, style, linkStyle, class, classDef, default, call, href, and interpolate [4][5]. If you must use these as labels, define the node with a safe ID and provide the reserved word as a quoted label (e.g., safeId["end"]) [4][5]. Starting Characters: Node IDs that begin with specific characters like 'o' or 'x' can sometimes be misinterpreted by the parser as edge types (e.g., leading to unintended edge styling or shapes) [5]. It is best practice to use descriptive alphanumeric IDs or camelCase [5][6]. Special Characters: While some special characters (like '.' or '/') may be permitted in certain contexts, they can lead to ambiguous or broken syntax [3][7]. To avoid issues with spaces or special characters in labels, always use double quotes [8][1][4]. Best Practices for Node IDs: Use Descriptive, Alphanumeric IDs: Adopt naming conventions like camelCase (e.g., orderService) rather than generic labels like s1 or single letters [5][6]. This improves readability and reduces the risk of collisions with parser keywords [7][6]. Separate IDs from Labels: Always explicitly define nodes if they require complex text or reserved words, using the format ID["Display Text"] [1][5]. This ensures the ID remains safe for internal referencing while the label displays correctly [1]. If a diagram fails to render, check for node IDs that match reserved keywords or contain unexpected special characters, and verify that all labels with spaces or symbols are enclosed in double quotes [1][4].

Citations:


Use synthetic node IDs for evidence items to maintain consistency and safety.

Line 34 uses raw ev.evidence_id directly as the Mermaid node identifier, while hypotheses (line 38), actions (line 42), and decisions (line 46) use indexed synthetic IDs. Evidence node IDs generated via nid("EVD") produce format "EVD-{hex}" with hyphens. While this may render, Mermaid best practices favor pure alphanumeric IDs, and the current approach creates an inconsistent pattern and fragility if ID format changes.

Suggested fix
-    for ev in evidence_items:
+    ev_index = 0
+    for ev in evidence_items:
         if ev.case_id == case.case_id:
             label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}"
-            ev_node = f'{ev.evidence_id}["{label}"]'
-            lines.append(f'C -->|contains| {ev_node}')
+            ev_node_id = f"EV{ev_index}"
+            ev_index += 1
+            lines.append(f'{ev_node_id}["{label}"]')
+            lines.append(f'C -->|contains| {ev_node_id}')
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
for ev in evidence_items:
if ev.case_id == case.case_id:
label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}"
ev_node = f'{ev.evidence_id}["{label}"]'
lines.append(f'C -->|contains| {ev_node}')
ev_index = 0
for ev in evidence_items:
if ev.case_id == case.case_id:
label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}"
ev_node_id = f"EV{ev_index}"
ev_index += 1
lines.append(f'{ev_node_id}["{label}"]')
lines.append(f'C -->|contains| {ev_node_id}')
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_graph.py` around lines 31 - 35, The evidence node ID
generation on line 34 uses raw ev.evidence_id directly, which is inconsistent
with the synthetic ID pattern used for hypotheses, actions, and decisions
elsewhere in the function. Replace the raw ev.evidence_id in the ev_node
assignment with a synthetic ID generated via the nid function (similar to how
hypotheses, actions, and decisions use indexed synthetic IDs like nid("HYP"),
nid("ACT"), nid("DEC")). This ensures all node identifiers follow a consistent,
safe alphanumeric format that aligns with Mermaid best practices and prevents
fragility if ID formats change in the future.

Comment on lines +8 to +13
if target_lower == "notebooklm_safe":
should_include_local_only = False
elif target_lower in ("gemini", "gpt", "copilot"):
should_include_local_only = include_local_only
elif target_lower == "local_ai":
should_include_local_only = include_local_only

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Cloud targets can bypass local_only protections via include_local_only=True.

For gemini/gpt/copilot, should_include_local_only = include_local_only permits explicit leakage to cloud prompts. That conflicts with public-safe privacy hardening.

Suggested fix
-    elif target_lower in ("gemini", "gpt", "copilot"):
-        should_include_local_only = include_local_only
+    elif target_lower in ("gemini", "gpt", "copilot"):
+        should_include_local_only = False
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_prompt.py` around lines 8 - 13, The cloud targets
(gemini, gpt, copilot) in the conditional statement currently allow bypassing
local-only protections by setting should_include_local_only equal to the
include_local_only parameter. To fix this security issue, change the assignment
for cloud targets (gemini, gpt, copilot) from should_include_local_only =
include_local_only to should_include_local_only = False, ensuring local-only
protected content cannot be sent to cloud services regardless of the
include_local_only parameter value, consistent with the privacy hardening
already applied to notebooklm_safe.

Comment on lines +19 to +27
for e in evidence_items:
if e.privacy_level == "local_only":
if not should_include_local_only:
excluded_any = True
lines.append(f"- [{e.source_type}] {e.title}: [EXCLUDED FOR PRIVACY - local_only]")
continue

snippet = e.extracted_text[:200] if e.extracted_text else ""
lines.append(f"- [{e.source_type}] {e.title}: {snippet}...")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Scope evidence to the active case before prompt assembly.

The current flow summarizes every item passed in, without filtering by case.case_id. That can leak another case’s evidence into a prompt pack.

Suggested fix
-def summarize_evidence_for_prompt(evidence_items: List[EvidenceItem], target: str, include_local_only: bool = False) -> str:
+def summarize_evidence_for_prompt(
+    evidence_items: List[EvidenceItem],
+    target: str,
+    include_local_only: bool = False,
+    case_id: str | None = None,
+) -> str:
@@
-    for e in evidence_items:
+    for e in evidence_items:
+        if case_id is not None and e.case_id != case_id:
+            continue
         if e.privacy_level == "local_only":
@@
-    prompt += summarize_evidence_for_prompt(evidence_items, target, include_local_only)
+    prompt += summarize_evidence_for_prompt(
+        evidence_items, target, include_local_only, case_id=case.case_id
+    )

Also applies to: 36-40

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/aios_habit/case_prompt.py` around lines 19 - 27, The evidence items being
processed in the loop are not filtered by the active case's case_id, which can
cause evidence from other cases to leak into the prompt. Before iterating
through evidence_items in the loop that starts with "for e in evidence_items:",
add a filter to only include items where e.case_id matches the current case's
case_id. This filtering should happen before the loop begins so that only
evidence belonging to the active case is included in the prompt assembly.

Comment on lines +154 to +155
bat_content = Path("RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
ps1_content = Path("scripts/run_case_cockpit.ps1").read_text(encoding="utf-8")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Make launcher path assertions independent of working directory.

Lines 154-155 depend on current working directory. This can fail in IDE/CI invocations that run tests from another path.

Suggested fix
 def test_launcher_contents_check():
-    bat_content = Path("RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
-    ps1_content = Path("scripts/run_case_cockpit.ps1").read_text(encoding="utf-8")
+    repo_root = Path(__file__).resolve().parents[1]
+    bat_content = (repo_root / "RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
+    ps1_content = (repo_root / "scripts" / "run_case_cockpit.ps1").read_text(encoding="utf-8")
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
bat_content = Path("RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
ps1_content = Path("scripts/run_case_cockpit.ps1").read_text(encoding="utf-8")
def test_launcher_contents_check():
repo_root = Path(__file__).resolve().parents[1]
bat_content = (repo_root / "RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
ps1_content = (repo_root / "scripts" / "run_case_cockpit.ps1").read_text(encoding="utf-8")
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/test_case_cockpit_hardening.py` around lines 154 - 155, The bat_content
and ps1_content assignments on lines 154-155 use relative paths that depend on
the current working directory, causing failures when tests run from different
locations. Make these paths absolute by anchoring them to the test file location
using __file__ and constructing paths relative to the test file's directory, so
the Path().read_text() calls work regardless of the invocation directory.

@Nakazasen
Nakazasen merged commit e1476d1 into main Jun 20, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant