Public-safe AIOS Habit MVP hardening - #1
Conversation
📝 WalkthroughWalkthroughThis PR introduces the complete AIOS Habit platform—a local-first, evidence-based personal memory system. It establishes governance and architectural foundations, repository structure with discovery and audit logs, JSON schemas and markdown templates, a core Python package with domain services and validation logic, two Streamlit applications (Studio for memory management and Case Cockpit for case-driven evidence capture), a multi-command CLI, comprehensive test coverage, and GitHub Actions CI. ChangesAIOS Habit Complete Platform Launch
Phase Gate Source->>Studio: Add Evidence Vault->>Studio: Build Profiles CLI->>Vault: Memory Commands |
There was a problem hiding this comment.
Actionable comments posted: 15
Note
Due to the large number of review comments, Critical, Major severity comments were prioritized as inline comments.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (4)
MASTER_PROJECT_INDEX.md (1)
5-70:⚠️ Potential issue | 🔴 Critical | ⚡ Quick winCRITICAL: Character encoding corruption in Vietnamese text renders document unreadable.
Lines 5, 9, 15-21, 24-25, 39-40 and throughout contain mojibake (corrupted character encoding). Vietnamese text like "Lニーu danh m盻・c" should be readable Vietnamese. The file appears to have been saved or processed with incorrect UTF-8 encoding (likely interpreted as Latin-1 and re-saved).
This is a data integrity violation per CONSTITUTION.md §6 (raw data policy) and violates source-of-truth requirements. The document is unreadable for Vietnamese speakers and will break any validation expecting valid UTF-8.
🔧 Recommended fix: Re-save file in proper UTF-8 encoding
- Open
MASTER_PROJECT_INDEX.mdin a text editor (e.g., VS Code, Notepad++, Sublime)- Ensure encoding is set to UTF-8 (without BOM) or UTF-8 with BOM (consistent with other files)
- Re-save the file
- Verify: All Vietnamese text should render correctly (e.g., "Lưu danh mục project được biết" instead of "Lニーu danh m盻・c")
Alternatively, if the source Vietnamese text is available, restore it to this file.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@MASTER_PROJECT_INDEX.md` around lines 5 - 70, The MASTER_PROJECT_INDEX.md file contains character encoding corruption affecting Vietnamese text throughout (visible as mojibake like "Lニーu danh m盻・c" instead of proper Vietnamese characters). Open the file in a text editor, change the file encoding to UTF-8 (without BOM), and re-save it to ensure all Vietnamese text renders correctly and the document becomes readable. Verify after saving that the text displays properly (for example, "Lưu danh mục project được biết" should appear instead of the corrupted version).08_audit/README.md (1)
1-6:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winTranslate or remove Vietnamese text for public release consistency.
Lines 3 and 5 contain untranslated Vietnamese:
- Line 3: "Lưu issue, validation log và rollback log."
- Line 5: "Audit là bắt buộc trước fix trong các workflow kỹ thuật."
For a public English-primary repository, this harms clarity and maintainability. Translate to English or remove entirely.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@08_audit/README.md` around lines 1 - 6, The README.md file in the 08_audit directory contains untranslated Vietnamese text on lines 3 and 5 that should be translated to English for consistency in a public-facing repository. Replace the Vietnamese text "Lưu issue, validation log và rollback log." and "Audit là bắt buộc trước fix trong các workflow kỹ thuật." with their English equivalents to ensure the documentation is clear and maintainable for all users. Alternatively, remove these lines entirely if they are not essential to understanding the Audit feature.07_ai_export_packs/README.md (1)
1-23:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove Vietnamese text for consistency in public release.
Line 3 contains untranslated Vietnamese: "Thư mục này chứa bản chuyển đổi profile/memory cho từng AI." The repository is English-primary; mixing languages without translation harms discoverability and maintenance for the public MVP release.
Translate to English or remove the Vietnamese line entirely.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@07_ai_export_packs/README.md` around lines 1 - 23, The README.md file contains Vietnamese text on the line following the "AI Export Packs" heading that is not translated to English. To maintain consistency with the English-primary repository for the public MVP release, either provide a complete English translation of the Vietnamese phrase "Thư mục này chứa bản chuyển đổi profile/memory cho từng AI." or remove the Vietnamese line entirely. Choose whichever approach best fits the documentation style and ensures the README is fully comprehensible to English-speaking users.08_audit/open_issues.md (1)
1-6:⚠️ Potential issue | 🟡 MinorResolve or re-prioritize ISS-0001 before public release.
ISS-0001 is marked
OPENwith statusPending review, while phase_0_report.md showsPASS. The technical Phase 0 validation (required files and schemas) passes, but the user review gate remains open. For a public MVP, clarify the status: either complete the user review, defer ISS-0001 to a later phase with explicit target, or re-assess severity if it's truly non-blocking.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@08_audit/open_issues.md` around lines 1 - 6, ISS-0001 in the open_issues.md table is marked as OPEN with Pending review status, but the Phase 0 validation has passed, creating ambiguity for public release. Update the ISS-0001 row to resolve this: either complete the user review and update Status to reflect closure, or change the Status to explicitly defer ISS-0001 to a future phase (e.g., Phase 1) with a target date in the Resolution column, or re-assess and lower the Severity if this is truly non-blocking for the MVP release. Choose the appropriate action based on product requirements and ensure the table clearly communicates the decision for stakeholders.
🟡 Minor comments (13)
src/aios_habit/profiles.py-8-10 (1)
8-10:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winKeyError risk: missing guard for 'statement' key.
Line 10 accesses
memory['statement']directly without a guard, which will raiseKeyErrorif the dictionary is missing the 'statement' key. While the CLI pre-filters memory items, this function is a reusable utility that should be defensive.🛡️ Proposed fix: add defensive guard
for memory in items: evidence = ", ".join(memory.get("evidence_ids", [])) - body += f"- {memory['statement']} Evidence: {evidence}\n" + body += f"- {memory.get('statement', 'UNKNOWN')} Evidence: {evidence}\n"🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/profiles.py` around lines 8 - 10, The memory dictionary access in the loop over items is not defensive. Replace the direct dictionary access `memory['statement']` with the safe `.get()` method using a default value (similar to how `memory.get("evidence_ids", [])` is already done on the previous line) to prevent KeyError when the 'statement' key is missing from the memory dictionary. This ensures the utility function remains defensive even though the CLI may pre-filter the data.src/aios_habit/paths.py-7-19 (1)
7-19:⚠️ Potential issue | 🟡 MinorRemove unused
IGNORE_DIRSfrompaths.py.
IGNORE_DIRSis defined in bothpaths.py(lines 7-19) anddiscovery.py(lines 5-18) with identical content. However,paths.IGNORE_DIRSis never imported or used anywhere in the codebase, whilediscovery.IGNORE_DIRSis actively used on lines 39 and 60 ofdiscovery.py. Remove the unused definition frompaths.pyto eliminate dead code and reduce maintenance burden.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/paths.py` around lines 7 - 19, Remove the unused IGNORE_DIRS constant definition from paths.py (lines 7-19) since it is never imported or used anywhere in the codebase and an identical definition already exists and is actively used in discovery.py. Simply delete the entire IGNORE_DIRS set definition block along with any associated whitespace to eliminate the dead code.src/aios_habit/evidence.py-26-26 (1)
26-26:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winAdd separator when concatenating strings for hashing.
Concatenating
source_pathandsummarywithout a delimiter can cause hash collisions when the same characters are split differently. For example,path="abc"+summary="def"produces the same hash aspath="ab"+summary="cdef".🔧 Proposed fix with delimiter
- hash=sha_text((source_path or "") + (summary or "")), + hash=sha_text((source_path or "") + "\n" + (summary or "")),🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/evidence.py` at line 26, The hash calculation in the sha_text() function call is concatenating source_path and summary without a delimiter, which can cause hash collisions when the same characters are split differently between the two strings. Add a separator character (such as a pipe "|" or colon ":") between the source_path and summary strings in the concatenation within the sha_text() call to ensure that different combinations of source_path and summary values produce different hashes.MASTER_PROJECT_INDEX.md-15-15 (1)
15-15:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winNaming inconsistency: "AIOS_habbit" (double-b) vs "AIOS_habit" (single-b).
Line 15 references "Repository m盻・c tiテェu c盻ァa n盻] t蘯」ng memory" for "AIOS_habbit", but other files and the Python package use "AIOS_habit". Confirm the correct project name after resolving the encoding issue.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@MASTER_PROJECT_INDEX.md` at line 15, In MASTER_PROJECT_INDEX.md line 15, fix the spelling inconsistency by changing "AIOS_habbit" (double-b) to "AIOS_habit" (single-b) to match the naming convention used in other files and the Python package. Additionally, address the encoding issue affecting the description text for this entry by re-encoding it to resolve the garbled characters (currently showing as "Repository m盻・c tiテェu c盻ァa n盻] t蘯」ng memory") and ensure the text displays correctly.02_sources/README.md-1-13 (1)
1-13:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove UTF-8 BOM character from file start.
Same BOM issue as noted in previous files.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@02_sources/README.md` around lines 1 - 13, The file 02_sources/README.md has a UTF-8 BOM (Byte Order Mark) character at the start of the file before the "# Sources" heading. Remove this BOM character from the beginning of the file using your editor's encoding settings or a text processing tool to save the file as UTF-8 without BOM.02_sources/README.md-3-13 (1)
3-13:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winClarify language choice: consider consistent English for public repository.
This file mixes Vietnamese and English inconsistently. Headers are in English, but the primary content descriptions are in Vietnamese. For a public MVP release, this could create accessibility barriers for non-Vietnamese-speaking contributors and users.
Recommendation: Either translate consistently to English (preferred for public repo) or provide both languages side-by-side with clear language labels, depending on team intent. If Vietnamese is intentional for a specific audience, document that design choice elsewhere.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@02_sources/README.md` around lines 3 - 13, The README.md file in the 02_sources directory has inconsistent language mixing, with English headers ("Important", "Files") but Vietnamese descriptions in the main content. For a public repository, translate all the Vietnamese text to English consistently throughout the file, including the introduction, Important section content, and the file descriptions to ensure accessibility for non-Vietnamese-speaking contributors and users.docs/PRIVACY_MODEL.md-1-4 (1)
1-4:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove UTF-8 BOM character from file start.
Same BOM issue as noted in previous files.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/PRIVACY_MODEL.md` around lines 1 - 4, Remove the UTF-8 BOM (Byte Order Mark) character that appears at the very beginning of the PRIVACY_MODEL.md file before the "# Privacy Model" heading. The BOM character is invisible but present and should be deleted to ensure the file starts cleanly with the heading text. Use a text editor that can handle BOM removal or manually delete the invisible character at the file's start.docs/OPERATOR_RUNBOOK.md-1-7 (1)
1-7:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove UTF-8 BOM character from file start.
Same BOM issue as noted in INSTALL.md.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/OPERATOR_RUNBOOK.md` around lines 1 - 7, The OPERATOR_RUNBOOK.md file has a UTF-8 BOM (Byte Order Mark) character at the very beginning of the file, before the "# Operator Runbook" heading. Remove this invisible BOM character from the start of the file to ensure proper file encoding without the BOM prefix. You can do this by opening the file in a text editor that supports BOM removal or by re-saving the file with UTF-8 encoding without BOM.docs/PHASE_GATE_PROCESS.md-1-2 (1)
1-2:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove UTF-8 BOM character from file start.
Same BOM issue as noted in previous files.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/PHASE_GATE_PROCESS.md` around lines 1 - 2, Remove the UTF-8 BOM (Byte Order Mark) character that appears at the start of the PHASE_GATE_PROCESS.md file, before the "#" symbol in the heading "# Phase Gate Process". The file should start directly with the "#" character without any invisible BOM prefix. This can typically be done by saving the file with UTF-8 encoding (without BOM) in your text editor.docs/RECOVERY_GUIDE.md-1-2 (1)
1-2:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove UTF-8 BOM character from file start.
Same BOM issue as noted in previous files.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/RECOVERY_GUIDE.md` around lines 1 - 2, Remove the UTF-8 BOM (Byte Order Mark) character from the beginning of the RECOVERY_GUIDE.md file. The BOM appears as an invisible character before the "# Recovery Guide" heading and should be deleted. Open the file in your editor, position the cursor at the very beginning before the "#" character, and delete any invisible characters at the start of the file. Save the file after removing the BOM to ensure the file starts directly with the "# Recovery Guide" text without any byte order mark prefix.docs/OPERATOR_RUNBOOK.md-4-5 (1)
4-5:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winClarify status terms for unsupported claims.
Line 5 references
UNKNOWNorneeds_evidenceas status values but doesn't define them. These should be explained in context or linked to documentation that defines the memory/evidence state model.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/OPERATOR_RUNBOOK.md` around lines 4 - 5, The OPERATOR_RUNBOOK.md file references status values `UNKNOWN` and `needs_evidence` for unsupported claims but does not define what these terms mean. Add clear definitions or explanations for these status values either inline in the document where they are mentioned or by including a reference/link to documentation that explains the memory and evidence state model. Ensure readers understand the distinction between these states and when each should be used.docs/INSTALL.md-1-7 (1)
1-7:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winRemove UTF-8 BOM character from file start.
The file begins with a UTF-8 byte-order mark (BOM) which can cause parsing issues with some tools, CI systems, and editor configurations. This should be removed.
🔧 Proposed fix
Remove the BOM character from line 1. The file should start directly with
# Installwithout any prefix.-# Install +# Install(Note: In your editor, ensure you're saving as UTF-8 without BOM, or use a tool like
sedordos2unixto strip it.)🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/INSTALL.md` around lines 1 - 7, The file starts with a UTF-8 byte-order mark (BOM) character before the "# Install" heading which can cause parsing issues with tools and CI systems. Remove the BOM character from the beginning of the file so that it starts directly with the "# Install" heading with no prefix. You can do this by saving the file as UTF-8 without BOM in your editor, or use command-line tools like sed or dos2unix to strip the BOM character.09_handover/README.md-3-3 (1)
3-3:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winMixed language in documentation.
Line 3 contains Vietnamese text while the rest of the documentation is in English. For consistency and accessibility, consider translating this line to English or establishing a clear i18n strategy for the documentation.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@09_handover/README.md` at line 3, The line containing "Lưu handover theo phase hoặc theo mốc quan trọng." is written in Vietnamese while the rest of the README.md documentation is in English, creating inconsistency. Translate this Vietnamese text to English to maintain language consistency throughout the documentation. Ensure the translated version maintains the original meaning about saving handover information according to phases or important milestones.
🧹 Nitpick comments (13)
src/aios_habit/cli.py (7)
204-218: ⚡ Quick winAdd explicit return code for consistency.
This function does not return an explicit exit code. While the implicit
Nonereturn converts to0inmain(), adding an explicit return statement improves consistency with other command handlers.♻️ Proposed fix
print_json({"status": "PASS", "dry_run": args.dry_run, "verified_exportable": len(memories)}) + return 0🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 204 - 218, The cmd_profile function does not have an explicit return statement, which reduces consistency with other command handlers. Add an explicit return statement at the end of the cmd_profile function after the print_json call to ensure it returns an explicit exit code (typically 0 for success) rather than implicitly returning None.
275-281: ⚡ Quick winAdd explicit return code for consistency.
This function does not return an explicit exit code. While the implicit
Nonereturn converts to0inmain(), adding an explicit return statement improves consistency with other command handlers.♻️ Proposed fix
print_json({"status": "PASS", "dry_run": args.dry_run, "path": "09_handover/final_handover.md"}) + return 0🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 275 - 281, The cmd_handover function lacks an explicit return statement at the end, which reduces consistency with other command handlers in the module. Add an explicit return statement (return 0 or simply return) after the print_json call to make the exit code explicit and improve consistency across all command handler functions.
144-154: ⚡ Quick winAdd explicit return code for consistency.
This function does not return an explicit exit code. While the implicit
Nonereturn converts to0inmain(), adding an explicit return statement improves consistency with other command handlers.♻️ Proposed fix
print_json({"status": "PASS", "dry_run": args.dry_run, "candidate": asdict(candidate)}) + return 0🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 144 - 154, The cmd_extract function does not include an explicit return statement at the end, which is inconsistent with other command handlers. Add an explicit return statement at the end of the cmd_extract function after the print_json call to provide a clear and consistent exit code for the command handler.
49-60: ⚡ Quick winAdd explicit return code for consistency.
This function does not return an explicit exit code, unlike
cmd_evidence,cmd_memory,cmd_export,cmd_audit, andcmd_phase. While Python's implicitNonereturn converts to0inmain(), explicit return statements improve consistency and clarity.♻️ Proposed fix
print_json({"status": "PASS", "dry_run": args.dry_run, "count": len(cards), "projects": data}) + return 0🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 49 - 60, The cmd_discover function lacks an explicit return statement for exit code consistency, while other command functions like cmd_evidence, cmd_memory, cmd_export, cmd_audit, and cmd_phase all explicitly return 0. Add an explicit return 0 statement at the end of the cmd_discover function to match the pattern used by the other command functions and improve code consistency.
63-96: ⚡ Quick winStandardize validation failure exit codes.
The
addsubcommand returns exit code2on validation failure (line 81), while thevalidatesubcommand returns1(line 96). This inconsistency may confuse scripts or automation that check exit codes.♻️ Proposed fix to standardize on exit code 1
if errors: print_json({"status": "FAIL", "errors": errors}) - return 2 + return 1🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 63 - 96, In the cmd_evidence function, the "add" subcommand returns exit code 2 when validation fails (when errors is truthy after calling record.validate()), while the "validate" subcommand returns exit code 1 for validation failures. Change the return statement in the "add" subcommand's validation failure block from `return 2` to `return 1` to standardize both subcommands to use the same exit code for validation failures.
99-141: ⚡ Quick winStandardize validation failure exit codes.
Similar to
cmd_evidence, theaddsubcommand returns exit code2on validation failure (line 116), while thevalidatesubcommand returns1(line 141). For consistency across all commands, standardize on a single exit code.♻️ Proposed fix to standardize on exit code 1
if errors: print_json({"status": "FAIL", "errors": errors}) - return 2 + return 1🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 99 - 141, The cmd_memory function uses inconsistent exit codes for validation failures across different subcommands. In the "add" subcommand branch (args.mem_cmd == "add"), when validation errors are found, the function returns exit code 2, but the validate subcommand at the end of the function returns exit code 1 for similar validation failures. Change the return statement in the "add" branch from return 2 to return 1 to standardize on exit code 1 for all validation failures throughout the cmd_memory function.
157-201: ⚡ Quick winStandardize validation failure exit codes.
The
addsubcommand returns exit code2on validation failure (line 187), while thevalidatesubcommand returns1(line 201). For consistency with the proposed standardization across all commands, use exit code1.♻️ Proposed fix to standardize on exit code 1
if errors: print_json({"status": "FAIL", "errors": errors}) - return 2 + return 1🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/cli.py` around lines 157 - 201, In the handle_generic function, standardize the validation failure exit code by changing the return value from 2 to 1 when validation errors occur during the add subcommand. Locate the line where errors are validated after creating either a WorkflowCard or DecisionPattern object, and change the return statement from return 2 to return 1 to match the exit code used by the validate subcommand at the end of the function.src/aios_habit/storage.py (2)
11-14: ⚖️ Poor tradeoffConcurrent append risk: append_jsonl lacks atomic write protection.
Multiple concurrent calls to
append_jsonlon the same file (e.g., from parallel CLI invocations) can interleave writes and corrupt the JSONL structure. While less likely in a local-first single-user context, the PR objectives emphasize reliability for public MVP.Consider adding file locking (e.g.,
fcntl.flockon Unix,msvcrt.lockingon Windows) if concurrent writes are expected, or document that the CLI is not safe for concurrent execution.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/storage.py` around lines 11 - 14, The append_jsonl function lacks file locking protection, which can cause data corruption when multiple concurrent processes write to the same JSONL file. Add platform-specific file locking to the function: use fcntl.flock on Unix-based systems and msvcrt.locking on Windows. Acquire the lock before opening the file in append mode, perform the write operation, and release the lock afterward. Alternatively, if concurrent writes are not expected in the design, add clear documentation stating that the CLI is not safe for concurrent execution.
8-8: ⚡ Quick winMemory inefficiency: read_text().splitlines() loads entire file.
Using
read_text().splitlines()loads the entire JSONL file into memory before parsing. For large evidence or memory registries, this could consume significant memory unnecessarily.♻️ Recommended refactor: stream line-by-line
def read_jsonl(path: Path) -> list[dict]: if not path.exists(): return [] - return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()] + records = [] + with path.open("r", encoding="utf-8") as f: + for line in f: + if line.strip(): + try: + records.append(json.loads(line)) + except json.JSONDecodeError: + continue + return records🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/storage.py` at line 8, The current implementation loads the entire JSONL file into memory using read_text().splitlines() before parsing each line, which is inefficient for large files. Replace this approach with streaming the file line-by-line using a file context manager and the open() function to iterate through the file, parsing each non-empty line with json.loads() as it's read. This ensures only one line is held in memory at a time rather than loading the entire file content at once.src/aios_habit/paths.py (1)
5-5: ⚖️ Poor tradeoffREPO_ROOT fragility: Path.cwd() changes with working directory.
Path.cwd()returns the current working directory at module import time, which can vary depending on where the CLI is invoked. This creates unpredictable behavior if the user runsaios-habitfrom different directories or if code changes the working directory at runtime.🔧 Recommended fix: anchor to package location or CLI-provided root
Consider one of these alternatives:
- Anchor to the package installation directory:
-REPO_ROOT = Path.cwd() +REPO_ROOT = Path(__file__).parent.parent.parent.resolve()
- Or accept repo root as a runtime parameter passed from the CLI (preferred for flexibility).
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/paths.py` at line 5, The REPO_ROOT variable in paths.py uses Path.cwd() which is fragile because it depends on the current working directory at import time. Replace the REPO_ROOT assignment with one of two approaches: either anchor it to the package installation directory by deriving it from the __file__ path and traversing up the parent directories, or accept repo_root as a runtime parameter that can be passed from the CLI entry point. The runtime parameter approach is preferred for flexibility, so consider modifying the paths module to accept an optional root parameter during initialization that defaults to deriving from __file__ if not provided.src/aios_habit/audit.py (1)
27-29: ⚡ Quick winInefficient directory skip check.
The check
any(part in SKIP_DIRS for part in path.parts)at line 28 iterates through all path parts for every file, which could be slow when scanning large repositories.♻️ Optional optimization: use set intersection
- if not path.is_file() or any(part in SKIP_DIRS for part in path.parts): + if not path.is_file() or SKIP_DIRS & set(path.parts): continue🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/audit.py` around lines 27 - 29, The directory skip check using `any(part in SKIP_DIRS for part in path.parts)` iterates through all path parts for every file found by rglob, resulting in poor performance with large repositories. To fix this, ensure SKIP_DIRS is defined as a set (not a list) and replace the any() check with a set intersection operation between the path parts converted to a set and the SKIP_DIRS set. This will change the lookup complexity from linear to constant time and significantly improve the scan performance for large directories.README.md (1)
11-18: ⚡ Quick winAddress adverb repetition in bullet list.
Both the line ending with "memory only." (line 17) and the next line starting with "Exports...only" (line 18) use "only" as a sentence-ender adverb, creating repetitive phrasing. Restructure one sentence to vary the syntax.
✏️ Proposed fix for adverb repetition
- Builds profiles from verified/export-allowed memory only. -- Exports AI packs only after redaction/audit checks. +- Exports AI packs after redaction and audit checks.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@README.md` around lines 11 - 18, The bullet points in the What AIOS Habit Does section contain repetitive use of the adverb "only" at the end of consecutive lines - specifically in the line "Builds profiles from verified/export-allowed memory only." and "Exports AI packs only after redaction/audit checks." Restructure one of these sentences to vary the syntax and eliminate the repetition, either by relocating the word "only" to a different position in the sentence, replacing it with a synonym, or rewording the clause entirely to maintain clarity while improving readability.08_audit/phase_0_report.md (1)
1-5: 💤 Low valueClarify intent of minimal phase report stubs.
Phase_0_report.md contains only a title and status line; no validation details, error descriptions, or summary. Comparing to cli.py lines 20-40 (record_phase_report function), reports are dynamically generated with error details appended.
This stub is either:
- A pre-populated audit record (should contain actual check results for user reference)
- A placeholder that will be overwritten at runtime (should be marked as auto-generated, or removed)
For clarity, either populate it with a real audit result summary, or explicitly mark it as "auto-generated—do not edit."
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@08_audit/phase_0_report.md` around lines 1 - 5, The phase_0_report.md file is currently a stub with only a title and status, lacking clarity about whether it's a pre-populated reference document or a runtime-generated placeholder. Since the record_phase_report function in cli.py (lines 20-40) dynamically generates reports with error details appended at runtime, clarify the intent by adding a header comment at the top of phase_0_report.md explicitly marking it as "auto-generated—do not edit" to indicate this file will be overwritten during execution, or alternatively populate it with actual validation check results and error summaries to serve as a reference document for users.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@00_governance/PHASE_0_EXIT_CHECKLIST.md`:
- Around line 1-28: The Phase 0 status is inconsistent between
PHASE_0_EXIT_CHECKLIST.md which shows status PASS with all 18 checks completed,
and PHASE_GATE_LOG.md which shows status OPENED with note "Waiting for user
review" for the same date (2026-06-20). Determine the actual current state of
Phase 0: if it is truly complete and closed, update the PHASE_GATE_LOG.md entry
to reflect the closed status with appropriate closure evidence; if Phase 0 is
still under review and not yet closed, revert the checks in
PHASE_0_EXIT_CHECKLIST.md to BLOCKED or PENDING status to match the
PHASE_GATE_LOG state. Ensure both documents reflect the same phase status for
the same date.
In `@00_governance/VALIDATION_RULES.md`:
- Around line 11-22: The MemoryUnit class in src/aios_habit/core.py and the
validation rules documented in VALIDATION_RULES.md are misaligned. First, update
the MemoryUnit class to match the schema definition in memory_unit.schema.json
by renaming the "category" field to "memory_type" and adding the missing
required fields: "scope", "confidence", and "rollback". Then, expand the
validate() method in MemoryUnit to enforce all 8 documented validation criteria:
unique memory_id, valid memory_type, concise statement (not raw quotes), at
least one evidence record, presence of confidence, presence of boundary/scope,
absence of unlabeled inference, and existence of rollback/deprecation path.
Ensure the updated validation method raises appropriate exceptions or returns
validation status for each criterion to match the governance requirements.
In `@10_schemas/decision_pattern.schema.json`:
- Around line 5-42: The JSON schema field names in decision_pattern.schema.json
do not match the field names in the Python DecisionPattern dataclass, causing
validation and deserialization failures. Update the DecisionPattern dataclass in
src/aios_habit/core.py to rename: the field currently named title to name, the
field currently named criteria to rule, and the field currently named
evidence_ids to evidence. Additionally, add a new required boundary field to the
DecisionPattern dataclass to match the schema requirement. Ensure all field
names and types are consistent between the schema and the Python model.
In `@10_schemas/evidence_record.schema.json`:
- Around line 6-87: The EvidenceRecord dataclass in src/aios_habit/core.py
(lines 59-72) has fields that do not match the evidence_record.schema.json
definition, causing serialized records to be non-compliant. Update the
EvidenceRecord dataclass to match the schema by: removing fields not in the
schema (title, source_path, source_pointer, captured_at, classification, hash,
risk_level, allowed_for_export), adding missing required fields (source_id,
permission_status, retention_policy), and renaming mismatched fields to align
with schema names (source_reference instead of source_path, created_at instead
of captured_at, permission_status instead of classification, artifact_hash
instead of hash). Ensure all required fields from the schema are present and the
field types match the schema definitions.
In `@10_schemas/memory_unit.schema.json`:
- Around line 18-99: Add the missing export_allowed field to the properties
section of the memory_unit schema. Define export_allowed as a boolean type
property with appropriate description documenting that this field can only be
set to true when the memory status is verified. This will align the schema
definition with the Python MemoryUnit model that already uses and validates this
field, ensuring consistency between the data model and the schema validation.
In `@10_schemas/workflow_card.schema.json`:
- Around line 47-51: The schema file defines field names and types that do not
match the WorkflowCard dataclass model. In the workflow_card.schema.json file,
rename the field `outputs` to `output` and change its type from array to string
to match the WorkflowCard model's output field. Additionally, rename the field
`evidence` to `evidence_ids` to match the WorkflowCard model's evidence_ids
field which is a list of strings. These changes must be made in the schema
definitions (around lines 47 and 82) to ensure the schema validation aligns with
the actual WorkflowCard dataclass fields defined in src/aios_habit/core.py.
In `@11_templates/workflow_card.md`:
- Line 13: Replace the Vietnamese placeholder text in the workflow card template
with English equivalents to maintain consistency and professionalism for the
public MVP release. The Vietnamese phrase "Khi nào workflow này được kích hoạt?"
should be replaced with "When is this workflow triggered?" and the Vietnamese
phrase "Nếu workflow sai hoặc gây lỗi, quay lại thế nào?" should be replaced
with an appropriate English equivalent such as "How to rollback if the workflow
fails?" or similar. This ensures the entire template document uses consistent
English language throughout.
- Line 6: The workflow_card.md template uses field names that do not match the
WorkflowCard dataclass contract. Change the field label from "Name:" to "Title:"
to align with the WorkflowCard's title field, and change "Outputs" to "Output:"
to align with the WorkflowCard's singular output field rather than plural
outputs. These corrections ensure the template matches the actual dataclass
fields used during CLI and runtime validation.
In `@src/aios_habit/audit.py`:
- Around line 40-42: The MemoryUnit constructor call in the audit loop lacks
error handling for malformed JSONL records, causing the entire audit to crash
when encountering invalid data. Wrap the MemoryUnit(**record) instantiation in a
try-except block to catch TypeError and ValueError exceptions, and when caught,
append an appropriate error message (such as "memory_id_unknown: Invalid record
format - error_details") to the errors list so the audit continues processing
remaining records and reports all issues comprehensively.
- Line 33: The read_text() method call is using errors="ignore" which silently
suppresses encoding errors and could hide file corruption issues during an audit
operation. Replace the errors="ignore" parameter with errors="replace" to
preserve undecodable characters while still allowing the file to be read, or
alternatively add explicit error handling that logs a warning when encoding
issues are encountered. This will ensure the audit function doesn't miss
potentially corrupted files.
In `@src/aios_habit/cli.py`:
- Line 18: The REPO variable assignment assumes the CLI is always invoked from
the repository root without any validation, which will cause failures if run
from a subdirectory. Add a helper function that checks for the presence of a
repository marker file (such as a .git directory, setup.py, or other identifying
file specific to the project) to validate that the current working directory is
actually the repository root. Call this validation helper before or during the
REPO assignment and raise an appropriate error if the repository root cannot be
found, ensuring users get clear feedback rather than silent failures in file
operations.
In `@src/aios_habit/core.py`:
- Around line 55-56: The load_json() function lacks error handling for two
common failure cases: missing files (FileNotFoundError) and malformed JSON
(JSONDecodeError). Wrap the json.loads() call and path.read_text() operation in
a try-except block to catch both FileNotFoundError and JSONDecodeError
exceptions, and provide appropriate fallback behavior such as returning an empty
dictionary or None when either error occurs, similar to how read_jsonl() handles
missing files by returning an empty list.
- Around line 38-41: The read_jsonl function does not handle JSONDecodeError
that can occur when json.loads(line) encounters malformed JSON content. Wrap the
json.loads(line) call within a try-except block to catch JSONDecodeError
exceptions, and either skip malformed lines silently or log a warning message
before continuing to the next line. This will prevent the function from crashing
when processing JSONL files with corrupted or invalid JSON entries, ensuring all
callers including audit, evidence validation, and CLI commands remain stable.
In `@src/aios_habit/memory.py`:
- Around line 11-13: Wrap the `MemoryUnit(**record)` instantiation in a
try-except block to handle TypeError exceptions that occur when JSONL records
have missing required fields or unexpected keys. When a construction error is
caught, append an error message to the errors list that identifies which record
failed and includes the error details, then continue processing the next record
instead of crashing. This ensures the validation loop completes even when
encountering malformed records in hand-edited JSONL files.
In `@src/aios_habit/storage.py`:
- Around line 5-8: The read_jsonl function lacks error handling for malformed
JSON lines which will cause a JSONDecodeError and crash the entire operation.
Wrap the json.loads(line) call in a try-except block within the list
comprehension to catch JSONDecodeError, and skip malformed lines gracefully
(optionally logging a warning) so that corrupted or hand-edited JSONL files can
still be processed without failing completely.
---
Outside diff comments:
In `@07_ai_export_packs/README.md`:
- Around line 1-23: The README.md file contains Vietnamese text on the line
following the "AI Export Packs" heading that is not translated to English. To
maintain consistency with the English-primary repository for the public MVP
release, either provide a complete English translation of the Vietnamese phrase
"Thư mục này chứa bản chuyển đổi profile/memory cho từng AI." or remove the
Vietnamese line entirely. Choose whichever approach best fits the documentation
style and ensures the README is fully comprehensible to English-speaking users.
In `@08_audit/open_issues.md`:
- Around line 1-6: ISS-0001 in the open_issues.md table is marked as OPEN with
Pending review status, but the Phase 0 validation has passed, creating ambiguity
for public release. Update the ISS-0001 row to resolve this: either complete the
user review and update Status to reflect closure, or change the Status to
explicitly defer ISS-0001 to a future phase (e.g., Phase 1) with a target date
in the Resolution column, or re-assess and lower the Severity if this is truly
non-blocking for the MVP release. Choose the appropriate action based on product
requirements and ensure the table clearly communicates the decision for
stakeholders.
In `@08_audit/README.md`:
- Around line 1-6: The README.md file in the 08_audit directory contains
untranslated Vietnamese text on lines 3 and 5 that should be translated to
English for consistency in a public-facing repository. Replace the Vietnamese
text "Lưu issue, validation log và rollback log." and "Audit là bắt buộc trước
fix trong các workflow kỹ thuật." with their English equivalents to ensure the
documentation is clear and maintainable for all users. Alternatively, remove
these lines entirely if they are not essential to understanding the Audit
feature.
In `@MASTER_PROJECT_INDEX.md`:
- Around line 5-70: The MASTER_PROJECT_INDEX.md file contains character encoding
corruption affecting Vietnamese text throughout (visible as mojibake like "Lニーu
danh m盻・c" instead of proper Vietnamese characters). Open the file in a text
editor, change the file encoding to UTF-8 (without BOM), and re-save it to
ensure all Vietnamese text renders correctly and the document becomes readable.
Verify after saving that the text displays properly (for example, "Lưu danh mục
project được biết" should appear instead of the corrupted version).
---
Minor comments:
In `@02_sources/README.md`:
- Around line 1-13: The file 02_sources/README.md has a UTF-8 BOM (Byte Order
Mark) character at the start of the file before the "# Sources" heading. Remove
this BOM character from the beginning of the file using your editor's encoding
settings or a text processing tool to save the file as UTF-8 without BOM.
- Around line 3-13: The README.md file in the 02_sources directory has
inconsistent language mixing, with English headers ("Important", "Files") but
Vietnamese descriptions in the main content. For a public repository, translate
all the Vietnamese text to English consistently throughout the file, including
the introduction, Important section content, and the file descriptions to ensure
accessibility for non-Vietnamese-speaking contributors and users.
In `@09_handover/README.md`:
- Line 3: The line containing "Lưu handover theo phase hoặc theo mốc quan
trọng." is written in Vietnamese while the rest of the README.md documentation
is in English, creating inconsistency. Translate this Vietnamese text to English
to maintain language consistency throughout the documentation. Ensure the
translated version maintains the original meaning about saving handover
information according to phases or important milestones.
In `@docs/INSTALL.md`:
- Around line 1-7: The file starts with a UTF-8 byte-order mark (BOM) character
before the "# Install" heading which can cause parsing issues with tools and CI
systems. Remove the BOM character from the beginning of the file so that it
starts directly with the "# Install" heading with no prefix. You can do this by
saving the file as UTF-8 without BOM in your editor, or use command-line tools
like sed or dos2unix to strip the BOM character.
In `@docs/OPERATOR_RUNBOOK.md`:
- Around line 1-7: The OPERATOR_RUNBOOK.md file has a UTF-8 BOM (Byte Order
Mark) character at the very beginning of the file, before the "# Operator
Runbook" heading. Remove this invisible BOM character from the start of the file
to ensure proper file encoding without the BOM prefix. You can do this by
opening the file in a text editor that supports BOM removal or by re-saving the
file with UTF-8 encoding without BOM.
- Around line 4-5: The OPERATOR_RUNBOOK.md file references status values
`UNKNOWN` and `needs_evidence` for unsupported claims but does not define what
these terms mean. Add clear definitions or explanations for these status values
either inline in the document where they are mentioned or by including a
reference/link to documentation that explains the memory and evidence state
model. Ensure readers understand the distinction between these states and when
each should be used.
In `@docs/PHASE_GATE_PROCESS.md`:
- Around line 1-2: Remove the UTF-8 BOM (Byte Order Mark) character that appears
at the start of the PHASE_GATE_PROCESS.md file, before the "#" symbol in the
heading "# Phase Gate Process". The file should start directly with the "#"
character without any invisible BOM prefix. This can typically be done by saving
the file with UTF-8 encoding (without BOM) in your text editor.
In `@docs/PRIVACY_MODEL.md`:
- Around line 1-4: Remove the UTF-8 BOM (Byte Order Mark) character that appears
at the very beginning of the PRIVACY_MODEL.md file before the "# Privacy Model"
heading. The BOM character is invisible but present and should be deleted to
ensure the file starts cleanly with the heading text. Use a text editor that can
handle BOM removal or manually delete the invisible character at the file's
start.
In `@docs/RECOVERY_GUIDE.md`:
- Around line 1-2: Remove the UTF-8 BOM (Byte Order Mark) character from the
beginning of the RECOVERY_GUIDE.md file. The BOM appears as an invisible
character before the "# Recovery Guide" heading and should be deleted. Open the
file in your editor, position the cursor at the very beginning before the "#"
character, and delete any invisible characters at the start of the file. Save
the file after removing the BOM to ensure the file starts directly with the "#
Recovery Guide" text without any byte order mark prefix.
In `@MASTER_PROJECT_INDEX.md`:
- Line 15: In MASTER_PROJECT_INDEX.md line 15, fix the spelling inconsistency by
changing "AIOS_habbit" (double-b) to "AIOS_habit" (single-b) to match the naming
convention used in other files and the Python package. Additionally, address the
encoding issue affecting the description text for this entry by re-encoding it
to resolve the garbled characters (currently showing as "Repository m盻・c tiテェu
c盻ァa n盻] t蘯」ng memory") and ensure the text displays correctly.
In `@src/aios_habit/evidence.py`:
- Line 26: The hash calculation in the sha_text() function call is concatenating
source_path and summary without a delimiter, which can cause hash collisions
when the same characters are split differently between the two strings. Add a
separator character (such as a pipe "|" or colon ":") between the source_path
and summary strings in the concatenation within the sha_text() call to ensure
that different combinations of source_path and summary values produce different
hashes.
In `@src/aios_habit/paths.py`:
- Around line 7-19: Remove the unused IGNORE_DIRS constant definition from
paths.py (lines 7-19) since it is never imported or used anywhere in the
codebase and an identical definition already exists and is actively used in
discovery.py. Simply delete the entire IGNORE_DIRS set definition block along
with any associated whitespace to eliminate the dead code.
In `@src/aios_habit/profiles.py`:
- Around line 8-10: The memory dictionary access in the loop over items is not
defensive. Replace the direct dictionary access `memory['statement']` with the
safe `.get()` method using a default value (similar to how
`memory.get("evidence_ids", [])` is already done on the previous line) to
prevent KeyError when the 'statement' key is missing from the memory dictionary.
This ensures the utility function remains defensive even though the CLI may
pre-filter the data.
---
Nitpick comments:
In `@08_audit/phase_0_report.md`:
- Around line 1-5: The phase_0_report.md file is currently a stub with only a
title and status, lacking clarity about whether it's a pre-populated reference
document or a runtime-generated placeholder. Since the record_phase_report
function in cli.py (lines 20-40) dynamically generates reports with error
details appended at runtime, clarify the intent by adding a header comment at
the top of phase_0_report.md explicitly marking it as "auto-generated—do not
edit" to indicate this file will be overwritten during execution, or
alternatively populate it with actual validation check results and error
summaries to serve as a reference document for users.
In `@README.md`:
- Around line 11-18: The bullet points in the What AIOS Habit Does section
contain repetitive use of the adverb "only" at the end of consecutive lines -
specifically in the line "Builds profiles from verified/export-allowed memory
only." and "Exports AI packs only after redaction/audit checks." Restructure one
of these sentences to vary the syntax and eliminate the repetition, either by
relocating the word "only" to a different position in the sentence, replacing it
with a synonym, or rewording the clause entirely to maintain clarity while
improving readability.
In `@src/aios_habit/audit.py`:
- Around line 27-29: The directory skip check using `any(part in SKIP_DIRS for
part in path.parts)` iterates through all path parts for every file found by
rglob, resulting in poor performance with large repositories. To fix this,
ensure SKIP_DIRS is defined as a set (not a list) and replace the any() check
with a set intersection operation between the path parts converted to a set and
the SKIP_DIRS set. This will change the lookup complexity from linear to
constant time and significantly improve the scan performance for large
directories.
In `@src/aios_habit/cli.py`:
- Around line 204-218: The cmd_profile function does not have an explicit return
statement, which reduces consistency with other command handlers. Add an
explicit return statement at the end of the cmd_profile function after the
print_json call to ensure it returns an explicit exit code (typically 0 for
success) rather than implicitly returning None.
- Around line 275-281: The cmd_handover function lacks an explicit return
statement at the end, which reduces consistency with other command handlers in
the module. Add an explicit return statement (return 0 or simply return) after
the print_json call to make the exit code explicit and improve consistency
across all command handler functions.
- Around line 144-154: The cmd_extract function does not include an explicit
return statement at the end, which is inconsistent with other command handlers.
Add an explicit return statement at the end of the cmd_extract function after
the print_json call to provide a clear and consistent exit code for the command
handler.
- Around line 49-60: The cmd_discover function lacks an explicit return
statement for exit code consistency, while other command functions like
cmd_evidence, cmd_memory, cmd_export, cmd_audit, and cmd_phase all explicitly
return 0. Add an explicit return 0 statement at the end of the cmd_discover
function to match the pattern used by the other command functions and improve
code consistency.
- Around line 63-96: In the cmd_evidence function, the "add" subcommand returns
exit code 2 when validation fails (when errors is truthy after calling
record.validate()), while the "validate" subcommand returns exit code 1 for
validation failures. Change the return statement in the "add" subcommand's
validation failure block from `return 2` to `return 1` to standardize both
subcommands to use the same exit code for validation failures.
- Around line 99-141: The cmd_memory function uses inconsistent exit codes for
validation failures across different subcommands. In the "add" subcommand branch
(args.mem_cmd == "add"), when validation errors are found, the function returns
exit code 2, but the validate subcommand at the end of the function returns exit
code 1 for similar validation failures. Change the return statement in the "add"
branch from return 2 to return 1 to standardize on exit code 1 for all
validation failures throughout the cmd_memory function.
- Around line 157-201: In the handle_generic function, standardize the
validation failure exit code by changing the return value from 2 to 1 when
validation errors occur during the add subcommand. Locate the line where errors
are validated after creating either a WorkflowCard or DecisionPattern object,
and change the return statement from return 2 to return 1 to match the exit code
used by the validate subcommand at the end of the function.
In `@src/aios_habit/paths.py`:
- Line 5: The REPO_ROOT variable in paths.py uses Path.cwd() which is fragile
because it depends on the current working directory at import time. Replace the
REPO_ROOT assignment with one of two approaches: either anchor it to the package
installation directory by deriving it from the __file__ path and traversing up
the parent directories, or accept repo_root as a runtime parameter that can be
passed from the CLI entry point. The runtime parameter approach is preferred for
flexibility, so consider modifying the paths module to accept an optional root
parameter during initialization that defaults to deriving from __file__ if not
provided.
In `@src/aios_habit/storage.py`:
- Around line 11-14: The append_jsonl function lacks file locking protection,
which can cause data corruption when multiple concurrent processes write to the
same JSONL file. Add platform-specific file locking to the function: use
fcntl.flock on Unix-based systems and msvcrt.locking on Windows. Acquire the
lock before opening the file in append mode, perform the write operation, and
release the lock afterward. Alternatively, if concurrent writes are not expected
in the design, add clear documentation stating that the CLI is not safe for
concurrent execution.
- Line 8: The current implementation loads the entire JSONL file into memory
using read_text().splitlines() before parsing each line, which is inefficient
for large files. Replace this approach with streaming the file line-by-line
using a file context manager and the open() function to iterate through the
file, parsing each non-empty line with json.loads() as it's read. This ensures
only one line is held in memory at a time rather than loading the entire file
content at once.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
| # Phase 0 Exit Checklist | ||
|
|
||
| Phase 0 closed only after concrete validation. The explicit `/goal` execution request is recorded as user approval evidence for this run. | ||
|
|
||
| | ID | Check | Status | Evidence / File | Notes | | ||
| |---|---|---|---|---| | ||
| | P0-01 | Constitution exists and reflects project philosophy | PASS | `CONSTITUTION.md` | Read and aligned | | ||
| | P0-02 | Roadmap has required phase fields | PASS | `ROADMAP.md` | Updated to Phase 0-9 | | ||
| | P0-03 | Architecture defines layered system and data flow | PASS | `ARCHITECTURE.md` | Existing architecture retained | | ||
| | P0-04 | Repository folder structure defined | PASS | `ARCHITECTURE.md` | Existing numbered structure preserved plus compatibility folders | | ||
| | P0-05 | Master identity profile exists | PASS | `MASTER_IDENTITY.md` | Candidate claims remain non-verified | | ||
| | P0-06 | Master behavior profile exists | PASS | `MASTER_BEHAVIOR_PROFILE.md` | Candidate claims remain non-verified | | ||
| | P0-07 | Master language profile exists | PASS | `MASTER_LANGUAGE_PROFILE.md` | Candidate claims remain non-verified | | ||
| | P0-08 | Master project index exists | PASS | `MASTER_PROJECT_INDEX.md` | Seed list remains non-exhaustive | | ||
| | P0-09 | Master workflow profile exists | PASS | `MASTER_WORKFLOW_PROFILE.md` | Candidate workflows remain non-verified | | ||
| | P0-10 | Memory schema exists and parses | PASS | `10_schemas/memory_unit.schema.json` | Parsed with `py -3` | | ||
| | P0-11 | Evidence schema exists and parses | PASS | `10_schemas/evidence_record.schema.json` | Parsed with `py -3` | | ||
| | P0-12 | Source policy blocks raw chat as direct memory | PASS | `00_governance/SOURCE_POLICY.md` | Confirmed | | ||
| | P0-13 | Data policy states local-first principle | PASS | `00_governance/DATA_POLICY.md` | Confirmed | | ||
| | P0-14 | Changelog initialized | PASS | `CHANGELOG.md` | Updated | | ||
| | P0-15 | Handover initialized | PASS | `PROJECT_HANDOVER.md` | Updated by generator | | ||
| | P0-16 | Rollback path exists | PASS | `08_audit/rollback_log.md` | Existing | | ||
| | P0-17 | `.gitignore` protects raw/local/secrets | PASS | `.gitignore` | Existing | | ||
| | P0-18 | User reviewed and approved execution | PASS | `/goal` request | Treated as explicit approval to execute all phases | | ||
|
|
||
| ## Current Phase 0 Result | ||
|
|
||
| Status: `PASS` |
There was a problem hiding this comment.
Reconcile phase status inconsistency between PHASE_0_EXIT_CHECKLIST and PHASE_GATE_LOG.
This document marks Phase 0 as PASS with all 18 checks completed (line 28), but PHASE_GATE_LOG.md records the same date (2026-06-20) with status OPENED and the note "Waiting for user review" (line 11). These statuses contradict each other.
Clarify whether Phase 0 is actually PASS (in which case update the PHASE_GATE_LOG entry to reflect closure) or still OPENED (in which case revert checks to BLOCKED or PENDING until gate closure conditions are met).
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@00_governance/PHASE_0_EXIT_CHECKLIST.md` around lines 1 - 28, The Phase 0
status is inconsistent between PHASE_0_EXIT_CHECKLIST.md which shows status PASS
with all 18 checks completed, and PHASE_GATE_LOG.md which shows status OPENED
with note "Waiting for user review" for the same date (2026-06-20). Determine
the actual current state of Phase 0: if it is truly complete and closed, update
the PHASE_GATE_LOG.md entry to reflect the closed status with appropriate
closure evidence; if Phase 0 is still under review and not yet closed, revert
the checks in PHASE_0_EXIT_CHECKLIST.md to BLOCKED or PENDING status to match
the PHASE_GATE_LOG state. Ensure both documents reflect the same phase status
for the same date.
| ## Memory Validation | ||
|
|
||
| Memory đạt PASS khi: | ||
|
|
||
| 1. Có `memory_id` duy nhất. | ||
| 2. Có `memory_type` hợp lệ. | ||
| 3. Có statement ngắn gọn, không phải raw quote dài. | ||
| 4. Có ít nhất một evidence record. | ||
| 5. Có confidence. | ||
| 6. Có boundary/scope. | ||
| 7. Không chứa suy đoán chưa gắn nhãn. | ||
| 8. Có rollback/deprecation path. |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Verify MemoryUnit.validate() implementation against documented criteria
ast-grep --pattern $'class MemoryUnit {
$$$
def validate($_) {
$$$
}
}'Repository: Nakazasen/AIOS_habbit
Length of output: 47
🏁 Script executed:
# Find core.py and related files
find . -name "core.py" -o -name "*memory*.py" | head -20Repository: Nakazasen/AIOS_habbit
Length of output: 117
🏁 Script executed:
# Also check the repo structure
git ls-files | grep -E "(core|memory)" | head -20Repository: Nakazasen/AIOS_habbit
Length of output: 515
🏁 Script executed:
# Check if VALIDATION_RULES.md exists
git ls-files | grep -i validationRepository: Nakazasen/AIOS_habbit
Length of output: 126
🏁 Script executed:
# Read VALIDATION_RULES.md to see the documented criteria
cat -n 00_governance/VALIDATION_RULES.mdRepository: Nakazasen/AIOS_habbit
Length of output: 1584
🏁 Script executed:
# Check memory.py for MemoryUnit class and validate method
cat -n src/aios_habit/memory.pyRepository: Nakazasen/AIOS_habbit
Length of output: 829
🏁 Script executed:
# Check core.py as well
cat -n src/aios_habit/core.pyRepository: Nakazasen/AIOS_habbit
Length of output: 6937
🏁 Script executed:
# Check the memory_unit.schema.json to see if additional validation is defined there
cat 10_schemas/memory_unit.schema.jsonRepository: Nakazasen/AIOS_habbit
Length of output: 2990
Critical gap between documented validation criteria, schema definition, and code implementation.
This document lists 8 criteria for memory PASS status (lines 11-22): unique ID, valid type, concise statement, evidence, confidence, boundary/scope, no unlabeled inference, and rollback path.
However, the MemoryUnit.validate() method in src/aios_habit/core.py only enforces 4 checks: valid category, valid status, evidence requirement for "verified" status, and export-allowed restriction. It does not validate statement conciseness, confidence presence, scope, rollback path, or unlabeled inference.
Additionally, 10_schemas/memory_unit.schema.json defines a schema with required fields (memory_type, scope, confidence, evidence, rollback) that don't match the Python MemoryUnit class fields (which uses "category" instead of "memory_type" and lacks "scope" and "rollback" fields). This schema-code misalignment means neither the schema nor the validation rules are properly enforced.
Memory can pass code validation without meeting documented criteria, creating false confidence. Align the Python implementation with the schema definition, then implement full validation for all 8 documented criteria to support governance robustness.
🧰 Tools
🪛 LanguageTool
[style] ~17-~17: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...uy nhất. 2. Có memory_type hợp lệ. 3. Có statement ngắn gọn, không phải raw quot...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
[style] ~18-~18: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ... ngắn gọn, không phải raw quote dài. 4. Có ít nhất một evidence record. 5. Có conf...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
[style] ~19-~19: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: .... 4. Có ít nhất một evidence record. 5. Có confidence. 6. Có boundary/scope. 7. Kh...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
[style] ~20-~20: Three successive sentences begin with the same word. Consider rewording the sentence or use a thesaurus to find a synonym.
Context: ...t evidence record. 5. Có confidence. 6. Có boundary/scope. 7. Không chứa suy đoán ...
(ENGLISH_WORD_REPEAT_BEGINNING_RULE)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@00_governance/VALIDATION_RULES.md` around lines 11 - 22, The MemoryUnit class
in src/aios_habit/core.py and the validation rules documented in
VALIDATION_RULES.md are misaligned. First, update the MemoryUnit class to match
the schema definition in memory_unit.schema.json by renaming the "category"
field to "memory_type" and adding the missing required fields: "scope",
"confidence", and "rollback". Then, expand the validate() method in MemoryUnit
to enforce all 8 documented validation criteria: unique memory_id, valid
memory_type, concise statement (not raw quotes), at least one evidence record,
presence of confidence, presence of boundary/scope, absence of unlabeled
inference, and existence of rollback/deprecation path. Ensure the updated
validation method raises appropriate exceptions or returns validation status for
each criterion to match the governance requirements.
| "required": [ | ||
| "decision_id", | ||
| "name", | ||
| "context", | ||
| "rule", | ||
| "evidence", | ||
| "boundary", | ||
| "status" | ||
| ], | ||
| "properties": { | ||
| "decision_id": { | ||
| "type": "string", | ||
| "pattern": "^DEC-[0-9]{4,}$" | ||
| }, | ||
| "name": { | ||
| "type": "string" | ||
| }, | ||
| "context": { | ||
| "type": "string" | ||
| }, | ||
| "rule": { | ||
| "type": "string" | ||
| }, | ||
| "tradeoffs": { | ||
| "type": "array", | ||
| "items": { | ||
| "type": "string" | ||
| } | ||
| }, | ||
| "boundary": { | ||
| "type": "string" | ||
| }, | ||
| "evidence": { | ||
| "type": "array", | ||
| "items": { | ||
| "type": "string" | ||
| } | ||
| }, |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Verify DecisionPattern dataclass field names
# Find DecisionPattern class definition
ast-grep --pattern $'class DecisionPattern:
$$$'
# Find DecisionPattern instantiation patterns
rg -nP '\bDecisionPattern\s*\(' -A 3Repository: Nakazasen/AIOS_habbit
Length of output: 1891
🏁 Script executed:
cat 10_schemas/decision_pattern.schema.jsonRepository: Nakazasen/AIOS_habbit
Length of output: 960
Multiple schema-code field name mismatches and missing field.
The schema defines required fields that don't align with the Python DecisionPattern dataclass in src/aios_habit/core.py:160:
- Schema
name↔ Codetitle - Schema
rule↔ Codecriteria - Schema
evidence↔ Codeevidence_ids - Schema requires
boundary↔ Code has no equivalent field
This will cause JSON validation failures and deserialization errors. Align all field names and add the missing boundary field to the Python model.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@10_schemas/decision_pattern.schema.json` around lines 5 - 42, The JSON schema
field names in decision_pattern.schema.json do not match the field names in the
Python DecisionPattern dataclass, causing validation and deserialization
failures. Update the DecisionPattern dataclass in src/aios_habit/core.py to
rename: the field currently named title to name, the field currently named
criteria to rule, and the field currently named evidence_ids to evidence.
Additionally, add a new required boundary field to the DecisionPattern dataclass
to match the schema requirement. Ensure all field names and types are consistent
between the schema and the Python model.
| "required": [ | ||
| "evidence_id", | ||
| "source_id", | ||
| "source_type", | ||
| "source_reference", | ||
| "summary", | ||
| "permission_status", | ||
| "retention_policy", | ||
| "created_at" | ||
| ], | ||
| "properties": { | ||
| "evidence_id": { | ||
| "type": "string", | ||
| "pattern": "^EVD-[0-9]{4,}$" | ||
| }, | ||
| "source_id": { | ||
| "type": "string", | ||
| "pattern": "^SRC-[0-9]{4,}$" | ||
| }, | ||
| "source_type": { | ||
| "type": "string", | ||
| "enum": [ | ||
| "chat_transcript", | ||
| "markdown", | ||
| "audit_report", | ||
| "commit_history", | ||
| "roadmap", | ||
| "specification", | ||
| "project_folder", | ||
| "prompt_library", | ||
| "user_interview", | ||
| "other" | ||
| ] | ||
| }, | ||
| "source_reference": { | ||
| "type": "string", | ||
| "description": "Local path, commit hash, artifact ID, or controlled reference." | ||
| }, | ||
| "artifact_hash": { | ||
| "type": "string" | ||
| }, | ||
| "summary": { | ||
| "type": "string" | ||
| }, | ||
| "content_boundary": { | ||
| "type": "string", | ||
| "description": "What part of source this evidence covers." | ||
| }, | ||
| "permission_status": { | ||
| "type": "string", | ||
| "enum": [ | ||
| "user_provided", | ||
| "user_approved", | ||
| "local_only", | ||
| "excluded", | ||
| "unknown" | ||
| ] | ||
| }, | ||
| "retention_policy": { | ||
| "type": "string", | ||
| "enum": [ | ||
| "metadata_only", | ||
| "summary_only", | ||
| "local_raw_allowed", | ||
| "do_not_store_raw", | ||
| "delete_after_extraction" | ||
| ] | ||
| }, | ||
| "linked_memory_ids": { | ||
| "type": "array", | ||
| "items": { | ||
| "type": "string" | ||
| } | ||
| }, | ||
| "created_at": { | ||
| "type": "string", | ||
| "format": "date" | ||
| }, | ||
| "notes": { | ||
| "type": "string" | ||
| } | ||
| } |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Verify EvidenceRecord dataclass field definitions
# Find EvidenceRecord class definition
ast-grep --pattern $'`@dataclass`
class EvidenceRecord:
$$$'
# Find all EvidenceRecord instantiations
rg -nP '\bEvidenceRecord\s*\(' -A 12 -B 1Repository: Nakazasen/AIOS_habbit
Length of output: 8424
🏁 Script executed:
# Check how the schema is used in the codebase
rg -l "evidence_record.*schema" -g "*.py" -g "*.json"
# Check if there's any schema validation being performed
rg -n "jsonschema|schema.*validat|validat.*schema" -g "*.py"
# Look at how records are serialized to JSON
rg -n "__dict__|to_dict|json\.dumps" -g "*.py" src/aios_habit/Repository: Nakazasen/AIOS_habbit
Length of output: 675
🏁 Script executed:
# Check the schema file location and its full content
find . -name "evidence_record.schema.json" -type f
# Verify if schema validation is happening anywhere
rg -n "evidence\.jsonl|EVIDENCE_PATH" -A 5 -B 2 -g "*.py"Repository: Nakazasen/AIOS_habbit
Length of output: 5622
🏁 Script executed:
# Check the test file to see if schema validation is tested
cat tests/test_aios_habit.py | head -100Repository: Nakazasen/AIOS_habbit
Length of output: 3160
🏁 Script executed:
# Check if schema files are used for validation anywhere
rg -n "jsonschema|\.schema\.json|validate.*schema" -g "*.py"
# Check phase_gate to see if it validates against schemas
cat src/aios_habit/phase_gate.py | head -150Repository: Nakazasen/AIOS_habbit
Length of output: 5196
🏁 Script executed:
# Check if there are any imports of jsonschema
rg -n "from jsonschema|import jsonschema" -g "*.py"
# Check requirements or dependencies
cat pyproject.toml 2>/dev/null || cat requirements.txt 2>/dev/null || cat setup.py 2>/dev/null | head -50Repository: Nakazasen/AIOS_habbit
Length of output: 484
🏁 Script executed:
# Check if schema is referenced in any documentation or comments
rg -n "evidence_record.*schema|schema.*evidence" --include="*.md" --include="*.py"
# Check what the actual dataclass __dict__ produces
python3 << 'PY'
from dataclasses import asdict, dataclass, field
from datetime import datetime
def now():
return datetime.now().isoformat()
`@dataclass`
class EvidenceRecord:
evidence_id: str
title: str
source_type: str
source_path: str
source_pointer: str = ""
captured_at: str = field(default_factory=now)
classification: str = "metadata_only"
summary: str = ""
hash: str = ""
risk_level: str = "low"
allowed_for_export: bool = True
notes: str = ""
record = EvidenceRecord(
evidence_id="EVD-1",
title="Test",
source_type="markdown",
source_path="README.md"
)
# Show what gets serialized
import json
serialized = asdict(record)
print("Serialized fields:")
for key in sorted(serialized.keys()):
print(f" {key}")
PYRepository: Nakazasen/AIOS_habbit
Length of output: 701
🏁 Script executed:
# Check if there's any actual JSON schema validation using jsonschema library
rg -n "validate|Draft|validator" src/aios_habit/ -g "*.py" | grep -i schema
# Verify no jsonschema dependency
grep -i "jsonschema" pyproject.tomlRepository: Nakazasen/AIOS_habbit
Length of output: 47
🏁 Script executed:
# Check how records are actually read back from evidence.jsonl
rg -B2 -A5 "EvidenceRecord\(\*\*record\)" src/aios_habit/Repository: Nakazasen/AIOS_habbit
Length of output: 927
🏁 Script executed:
# Check if core.py or models.py has the definition
head -5 src/aios_habit/core.py src/aios_habit/models.py 2>&1 | grep -A 5 "^==>"
# Check what's being imported from models/core in evidence.py
head -20 src/aios_habit/evidence.py
# Confirm the actual location of EvidenceRecord
grep -n "^from.*EvidenceRecord\|^import.*EvidenceRecord" src/aios_habit/Repository: Nakazasen/AIOS_habbit
Length of output: 898
🏁 Script executed:
# Verify what happens when deserializing with mismatched fields
python3 << 'PY'
from dataclasses import asdict, dataclass, field
`@dataclass`
class EvidenceRecord:
evidence_id: str
title: str
source_type: str
source_path: str
source_pointer: str = ""
captured_at: str = ""
classification: str = "metadata_only"
summary: str = ""
hash: str = ""
risk_level: str = "low"
allowed_for_export: bool = True
notes: str = ""
# Test 1: Deserialize from JSON that matches code fields
from_code = {
"evidence_id": "EVD-1",
"title": "Test",
"source_type": "markdown",
"source_path": "README.md",
"summary": "Summary",
"classification": "metadata_only",
"captured_at": "2024-01-01"
}
try:
r1 = EvidenceRecord(**from_code)
print("✓ Code-generated fields deserialize successfully")
except Exception as e:
print(f"✗ Code fields error: {e}")
# Test 2: Deserialize from JSON that matches schema fields
from_schema = {
"evidence_id": "EVD-1",
"source_id": "SRC-1234",
"source_type": "markdown",
"source_reference": "README.md",
"summary": "Summary",
"permission_status": "user_provided",
"retention_policy": "metadata_only",
"created_at": "2024-01-01"
}
try:
r2 = EvidenceRecord(**from_schema)
print("✓ Schema fields deserialize successfully")
except Exception as e:
print(f"✗ Schema fields error: {e}")
PYRepository: Nakazasen/AIOS_habbit
Length of output: 210
🏁 Script executed:
# Check if the schema file is actually referenced anywhere in the codebase
rg -n "evidence_record.schema" -g "*.py" -g "*.md"
# Verify the actual location referenced in the review comment
sed -n '59,96p' src/aios_habit/cli.pyRepository: Nakazasen/AIOS_habbit
Length of output: 1875
🏁 Script executed:
# Check if the review location is accurate - look at cli.py around line 63-96
cat -n src/aios_habit/cli.py | sed -n '59,97p'Repository: Nakazasen/AIOS_habbit
Length of output: 1874
Schema and code field definitions are misaligned.
The evidence_record.schema.json and EvidenceRecord dataclass (src/aios_habit/core.py:59-72) define incompatible field sets. Records serialized by the code use asdict() and will not conform to the schema structure.
Code provides but schema doesn't define:
titlesource_path(schema expectssource_reference)source_pointercaptured_at(schema expectscreated_at)classification(schema expectspermission_status)hash(schema hasartifact_hash)risk_levelallowed_for_export
Schema requires but code doesn't provide:
source_idpermission_status(mapped fromclassification)retention_policy
While JSON schema validation is not currently enforced in the codebase (no jsonschema dependency), this contract break indicates the schema is either aspirational or disconnected from the implementation. If the schema represents the intended data model, the EvidenceRecord class must be updated to match.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@10_schemas/evidence_record.schema.json` around lines 6 - 87, The
EvidenceRecord dataclass in src/aios_habit/core.py (lines 59-72) has fields that
do not match the evidence_record.schema.json definition, causing serialized
records to be non-compliant. Update the EvidenceRecord dataclass to match the
schema by: removing fields not in the schema (title, source_path,
source_pointer, captured_at, classification, hash, risk_level,
allowed_for_export), adding missing required fields (source_id,
permission_status, retention_policy), and renaming mismatched fields to align
with schema names (source_reference instead of source_path, created_at instead
of captured_at, permission_status instead of classification, artifact_hash
instead of hash). Ensure all required fields from the schema are present and the
field types match the schema definitions.
| "properties": { | ||
| "memory_id": { | ||
| "type": "string", | ||
| "pattern": "^MEM-[A-Z]+-[0-9]{4,}$" | ||
| }, | ||
| "memory_type": { | ||
| "type": "string", | ||
| "enum": [ | ||
| "identity", | ||
| "behavior", | ||
| "language", | ||
| "workflow", | ||
| "project_knowledge", | ||
| "lessons_learned", | ||
| "decision_patterns" | ||
| ] | ||
| }, | ||
| "title": { | ||
| "type": "string", | ||
| "minLength": 3 | ||
| }, | ||
| "statement": { | ||
| "type": "string", | ||
| "description": "Pattern/knowledge statement, not raw transcript." | ||
| }, | ||
| "status": { | ||
| "type": "string", | ||
| "enum": [ | ||
| "candidate", | ||
| "validated", | ||
| "deprecated", | ||
| "conflicted", | ||
| "rejected" | ||
| ] | ||
| }, | ||
| "confidence": { | ||
| "type": "string", | ||
| "enum": [ | ||
| "low", | ||
| "medium", | ||
| "high", | ||
| "verified" | ||
| ] | ||
| }, | ||
| "scope": { | ||
| "type": "string", | ||
| "description": "Where this memory applies and does not apply." | ||
| }, | ||
| "evidence": { | ||
| "type": "array", | ||
| "minItems": 1, | ||
| "items": { | ||
| "$ref": "#/$defs/evidence_link" | ||
| } | ||
| }, | ||
| "tags": { | ||
| "type": "array", | ||
| "items": { | ||
| "type": "string" | ||
| } | ||
| }, | ||
| "linked_projects": { | ||
| "type": "array", | ||
| "items": { | ||
| "type": "string" | ||
| } | ||
| }, | ||
| "validation": { | ||
| "$ref": "#/$defs/validation" | ||
| }, | ||
| "rollback": { | ||
| "type": "string" | ||
| }, | ||
| "created_at": { | ||
| "type": "string", | ||
| "format": "date" | ||
| }, | ||
| "updated_at": { | ||
| "type": "string", | ||
| "format": "date" | ||
| } | ||
| }, |
There was a problem hiding this comment.
Missing export_allowed field definition.
The Python MemoryUnit model in src/aios_habit/phase_gate.py:95 uses and validates an export_allowed field with the constraint "only verified memory can be export_allowed", but this property is not defined in the schema.
This creates a schema-code contract gap where:
- Python code can create records with
export_allowed - Schema validation will reject those records as having unexpected properties (if
additionalProperties: falseis enforced) - The documented data model is incomplete
Add the export_allowed boolean property to the schema with appropriate documentation of the verified-status constraint.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@10_schemas/memory_unit.schema.json` around lines 18 - 99, Add the missing
export_allowed field to the properties section of the memory_unit schema. Define
export_allowed as a boolean type property with appropriate description
documenting that this field can only be set to true when the memory status is
verified. This will align the schema definition with the Python MemoryUnit model
that already uses and validates this field, ensuring consistency between the
data model and the schema validation.
| from .profiles import build_profile_text | ||
| from .storage import append_jsonl, read_jsonl, write_json | ||
|
|
||
| REPO = Path.cwd() |
There was a problem hiding this comment.
Validate that the current directory is the repository root.
REPO = Path.cwd() assumes the CLI is always invoked from the repository root. If a user runs the CLI from a subdirectory, all file operations will fail or write to incorrect locations, degrading the user experience.
🛡️ Proposed fix to add repository validation
Add a helper to verify the presence of a repository marker file:
+def validate_repo(repo: Path) -> None:
+ if not (repo / "CONSTITUTION.md").exists():
+ print_json({"status": "FAIL", "error": "Not in AIOS Habit repository root"})
+ sys.exit(1)
+
REPO = Path.cwd()
+validate_repo(REPO)
EVIDENCE_PATH = REPO / "03_evidence_registry" / "records" / "evidence.jsonl"🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/cli.py` at line 18, The REPO variable assignment assumes the
CLI is always invoked from the repository root without any validation, which
will cause failures if run from a subdirectory. Add a helper function that
checks for the presence of a repository marker file (such as a .git directory,
setup.py, or other identifying file specific to the project) to validate that
the current working directory is actually the repository root. Call this
validation helper before or during the REPO assignment and raise an appropriate
error if the repository root cannot be found, ensuring users get clear feedback
rather than silent failures in file operations.
| def read_jsonl(path: Path) -> list[dict]: | ||
| if not path.exists(): | ||
| return [] | ||
| return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()] |
There was a problem hiding this comment.
Add error handling for malformed JSON lines.
If a JSONL file contains an invalid JSON line (e.g., due to manual editing or corruption), json.loads(line) will raise JSONDecodeError and crash. This affects all callers including audit, evidence validation, and CLI commands.
🛡️ Proposed fix with error handling
def read_jsonl(path: Path) -> list[dict]:
if not path.exists():
return []
- return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
+ results = []
+ for line in path.read_text(encoding="utf-8").splitlines():
+ if line.strip():
+ try:
+ results.append(json.loads(line))
+ except json.JSONDecodeError as e:
+ # Log or skip malformed lines to prevent crashes
+ continue
+ return results🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/core.py` around lines 38 - 41, The read_jsonl function does
not handle JSONDecodeError that can occur when json.loads(line) encounters
malformed JSON content. Wrap the json.loads(line) call within a try-except block
to catch JSONDecodeError exceptions, and either skip malformed lines silently or
log a warning message before continuing to the next line. This will prevent the
function from crashing when processing JSONL files with corrupted or invalid
JSON entries, ensuring all callers including audit, evidence validation, and CLI
commands remain stable.
| def load_json(path: Path): | ||
| return json.loads(path.read_text(encoding="utf-8")) |
There was a problem hiding this comment.
Add error handling for missing files and malformed JSON.
load_json() can crash with FileNotFoundError if the file is missing or JSONDecodeError if the JSON is invalid. Unlike read_jsonl(), which returns an empty list for missing files, this function provides no fallback.
🛡️ Proposed fix with error handling
def load_json(path: Path):
+ if not path.exists():
+ return None # or raise a more descriptive error
+ try:
- return json.loads(path.read_text(encoding="utf-8"))
+ return json.loads(path.read_text(encoding="utf-8"))
+ except json.JSONDecodeError as e:
+ raise ValueError(f"Invalid JSON in {path}: {e}")📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| def load_json(path: Path): | |
| return json.loads(path.read_text(encoding="utf-8")) | |
| def load_json(path: Path): | |
| if not path.exists(): | |
| return None # or raise a more descriptive error | |
| try: | |
| return json.loads(path.read_text(encoding="utf-8")) | |
| except json.JSONDecodeError as e: | |
| raise ValueError(f"Invalid JSON in {path}: {e}") |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/core.py` around lines 55 - 56, The load_json() function lacks
error handling for two common failure cases: missing files (FileNotFoundError)
and malformed JSON (JSONDecodeError). Wrap the json.loads() call and
path.read_text() operation in a try-except block to catch both FileNotFoundError
and JSONDecodeError exceptions, and provide appropriate fallback behavior such
as returning an empty dictionary or None when either error occurs, similar to
how read_jsonl() handles missing files by returning an empty list.
| for record in read_jsonl(path): | ||
| memory = MemoryUnit(**record) | ||
| errors.extend(f"{memory.memory_id}: {error}" for error in memory.validate()) |
There was a problem hiding this comment.
Missing error handling for malformed JSONL records.
MemoryUnit(**record) at line 12 will raise TypeError if the JSONL record is missing required fields or contains unexpected keys, causing the entire validation to crash. Since JSONL files may be hand-edited, this creates a runtime crash risk.
🛡️ Proposed fix: catch construction errors
for record in read_jsonl(path):
- memory = MemoryUnit(**record)
+ try:
+ memory = MemoryUnit(**record)
+ except TypeError as exc:
+ errors.append(f"malformed record: {exc}")
+ continue
errors.extend(f"{memory.memory_id}: {error}" for error in memory.validate())🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/memory.py` around lines 11 - 13, Wrap the
`MemoryUnit(**record)` instantiation in a try-except block to handle TypeError
exceptions that occur when JSONL records have missing required fields or
unexpected keys. When a construction error is caught, append an error message to
the errors list that identifies which record failed and includes the error
details, then continue processing the next record instead of crashing. This
ensures the validation loop completes even when encountering malformed records
in hand-edited JSONL files.
| def read_jsonl(path: Path) -> list[dict]: | ||
| if not path.exists(): | ||
| return [] | ||
| return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()] |
There was a problem hiding this comment.
Missing error handling for malformed JSON in read_jsonl.
json.loads(line) at line 8 will raise JSONDecodeError if any line contains invalid JSON, causing the entire operation to fail. Since JSONL files may be hand-edited or corrupted, this creates a runtime crash risk.
🛡️ Proposed fix: handle malformed lines gracefully
def read_jsonl(path: Path) -> list[dict]:
if not path.exists():
return []
- return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
+ records = []
+ for line in path.read_text(encoding="utf-8").splitlines():
+ if not line.strip():
+ continue
+ try:
+ records.append(json.loads(line))
+ except json.JSONDecodeError:
+ # Skip malformed lines or log/raise as needed
+ continue
+ return records📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| def read_jsonl(path: Path) -> list[dict]: | |
| if not path.exists(): | |
| return [] | |
| return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()] | |
| def read_jsonl(path: Path) -> list[dict]: | |
| if not path.exists(): | |
| return [] | |
| records = [] | |
| for line in path.read_text(encoding="utf-8").splitlines(): | |
| if not line.strip(): | |
| continue | |
| try: | |
| records.append(json.loads(line)) | |
| except json.JSONDecodeError: | |
| # Skip malformed lines or log/raise as needed | |
| continue | |
| return records |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/storage.py` around lines 5 - 8, The read_jsonl function lacks
error handling for malformed JSON lines which will cause a JSONDecodeError and
crash the entire operation. Wrap the json.loads(line) call in a try-except block
within the list comprehension to catch JSONDecodeError, and skip malformed lines
gracefully (optionally logging a warning) so that corrupted or hand-edited JSONL
files can still be processed without failing completely.
There was a problem hiding this comment.
Actionable comments posted: 10
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
RUN_AIOS_CASE_COCKPIT.bat (1)
1-14:⚠️ Potential issue | 🟡 Minor | ⚡ Quick winUse CRLF line endings for Windows batch reliability.
This
.batfile is flagged as LF-only; converting to CRLF avoids known Windows batch parsing edge cases.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@RUN_AIOS_CASE_COCKPIT.bat` around lines 1 - 14, The RUN_AIOS_CASE_COCKPIT.bat file currently uses LF (Unix-style) line endings, which can cause parsing issues in Windows batch scripts. Convert all line endings in this file from LF to CRLF (Windows-style) to ensure proper batch file execution on Windows systems. This can be done through your text editor's line ending conversion feature or by using Git configuration to normalize line endings for batch files.Source: Linters/SAST tools
docs/CASE_COCKPIT.md (1)
17-20:⚠️ Potential issue | 🟠 MajorDocumented command requires package installation step that is omitted from docs.
Line 19 documents
py -3 -m streamlit run src\aios_habit\case_cockpit.py, but this command will fail with relative import errors because the package is not installed. The.batand.ps1scripts correctly executepy -3 -m pip install -e .first to register the package; the docs omit this critical prerequisite. Either include the installation step in the documentation or reference the entry point defined inpyproject.toml:aios-case-cockpit(which requires package installation).🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/CASE_COCKPIT.md` around lines 17 - 20, The documentation for running the case cockpit via the direct streamlit command is incomplete because it omits the required package installation step that the .bat and .ps1 scripts correctly include. Update the documentation to either include the py -3 -m pip install -e . installation step before the streamlit run command, or replace it with the entry point command aios-case-cockpit (which requires the package to be installed first). Reference the pyproject.toml file where the aios-case-cockpit entry point is defined to ensure consistency with the project setup.
🧹 Nitpick comments (2)
tests/test_case_cockpit.py (1)
17-23: ⚡ Quick winAdd regression tests for privacy and upload-path hardening paths.
Please add tests that assert filename sanitization in upload persistence and
local_onlyevidence exclusion in prompt construction, so the hardening behavior is locked in.Also applies to: 24-32
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/test_case_cockpit.py` around lines 17 - 23, The current test_csv_ingest_reads_synthetic_file test only verifies basic CSV ingest functionality but lacks regression tests for the security hardening features. Add additional test functions that specifically assert filename sanitization behavior in upload persistence (verify that special characters or path traversal attempts in filenames are properly sanitized before being persisted) and test that evidence marked with local_only flag is properly excluded from prompt construction (verify that such evidence does not appear in the extracted text or other prompt-related outputs). These regression tests will lock in the hardening behavior and prevent future regressions of these security features.src/aios_habit/case_models.py (1)
17-17: ⚡ Quick winUse the shared UTC timestamp helper for persisted model defaults.
These defaults currently use local naive timestamps, while the project already exposes a canonical UTC formatter (
src/aios_habit/core.py:26-27). Keeping one format avoids mixed timestamp semantics in JSONL records.Suggested patch
from dataclasses import dataclass, field -from datetime import datetime from typing import List, Optional, Dict, Any +from .core import now @@ - created_at: str = field(default_factory=lambda: datetime.now().isoformat()) + created_at: str = field(default_factory=now) @@ - created_at: str = field(default_factory=lambda: datetime.now().isoformat()) - updated_at: str = field(default_factory=lambda: datetime.now().isoformat()) + created_at: str = field(default_factory=now) + updated_at: str = field(default_factory=now)Also applies to: 25-26
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/aios_habit/case_models.py` at line 17, The `created_at` field (and other timestamp fields on lines 25-26) are using local naive timestamps via `datetime.now().isoformat()`, which creates inconsistent timestamp formats in JSONL records. Replace the lambda default factories for all timestamp fields to use the canonical UTC timestamp helper that is already exposed in `src/aios_habit/core.py:26-27`, ensuring all persisted model defaults use consistent UTC timestamp semantics across the codebase.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@RUN_AIOS_CASE_COCKPIT.bat`:
- Around line 5-12: The dependency check on line 5 only verifies third-party
packages (streamlit, pandas, openpyxl) but does not check if the local
aios_habit package is installed. If the third-party packages are already present
but the local package is missing, the errorlevel check passes and line 8 (pip
install -e .) is skipped, causing the streamlit run command on line 12 to fail
at import time. Modify the pip show command to also verify the aios_habit
package is installed, or add a separate verification step that attempts to
import the aios_habit module to ensure both the third-party dependencies and the
local package are available before running the streamlit application.
In `@RUN_AIOS_HABIT_STUDIO.bat`:
- Around line 1-13: The RUN_AIOS_HABIT_STUDIO.bat file currently uses Unix line
endings (LF) which can cause batch parser failures on Windows systems. Convert
the entire file to use Windows line endings (CRLF) by saving the file with CRLF
line ending format in your editor, or using a command-line tool like dos2unix
with the DOS option to perform the conversion. Ensure that all line breaks in
the batch file are properly formatted as CRLF before committing.
In `@scripts/run_case_cockpit.ps1`:
- Around line 5-11: The dependency check at lines 5-8 only verifies that the
streamlit command exists in PATH using Get-Command, but does not confirm that
the aios_habit package is properly installed and importable in the Python
interpreter context that will execute case_cockpit.py. Replace or supplement the
Get-Command check for streamlit with a Python-based verification that attempts
to import both the streamlit module and the aios_habit package using py -3 to
ensure both dependencies are available in the same interpreter instance before
proceeding to launch the application.
In `@scripts/run_studio.ps1`:
- Around line 6-12: The current check only verifies that the streamlit command
exists but does not verify that the aios_habit package is installed, which will
cause the script to fail at runtime when studio.py tries to import from the
aios_habit module. Modify the dependency verification logic to check if both
streamlit and aios_habit are importable using Python import checks, and only
skip the pip install step if both are successfully importable. This ensures that
running py -3 -m streamlit run src\aios_habit\studio.py will have all required
dependencies available.
In `@src/aios_habit/case_cockpit.py`:
- Around line 205-207: The privacy filter for local_only evidence in the loop
currently only excludes content when the target contains "NotebookLM", but
local_only content should be excluded for all external targets. Modify the
condition in the if statement that checks both "NotebookLM" in target and
e.privacy_level == "local_only" to only check e.privacy_level == "local_only" so
that local_only evidence is consistently filtered out regardless of whether the
target is NotebookLM, GPT, Gemini, or Copilot.
- Around line 181-183: The code in the "Add Action" button handler does not
validate that new_action is not empty before appending it to
active_case.next_actions and persisting it with save_case(). Add a validation
check to ensure new_action contains non-empty, non-whitespace content before
proceeding with the append and save operations. Only execute the append and
save_case call if the new_action passes this validation check.
In `@src/aios_habit/case_ingest.py`:
- Around line 64-66: The uploaded filename from uploaded_file.name is used
directly without sanitization, creating a security risk for path traversal
attacks and file overwrite collisions. Extract only the base filename from
uploaded_file.name using os.path.basename() or pathlib.Path.name to remove any
directory path components before constructing the dest_path variable and writing
the file to disk. This ensures that only safe filenames without path separators
are written to the assets_dir directory.
In `@src/aios_habit/case_store.py`:
- Around line 31-32: The bare except Exception: pass blocks at lines 31-32 and
44-45 are silently suppressing parse and schema validation errors when loading
JSONL records, causing data loss without any notification or logging. Replace
these silent exception blocks with proper error handling that logs the specific
error details, includes information about which record failed, and either
re-raises the exception or handles it in a way that makes the failure visible to
the user instead of silently excluding persisted records from the application.
- Around line 60-63: The current implementation of save_cases and the related
code at lines 76-78 write directly to CASES_FILE, which is not atomic and can
leave partial files on interruption or cause data loss when multiple sessions
write concurrently. Refactor the file write operation to first write all JSON
lines to a temporary file, then atomically rename that temporary file to
CASES_FILE to ensure the entire write either succeeds completely or fails
without corrupting the target file.
In `@src/aios_habit/studio.py`:
- Around line 218-226: The approval logic in the button handler checks only that
evidence_ids is non-empty but does not validate that the referenced evidence
actually exists, and it unconditionally appends to MEMORY_PATH without checking
for duplicates. To fix this, before setting status to verified and calling
append_jsonl, add validation to confirm that each evidence_id referenced in data
actually exists in the evidence store, and implement a check to prevent
duplicate entries in MEMORY_PATH by either verifying the record does not already
exist before appending or by replacing an existing record if it does. Ensure the
success message is only shown after successful validation and append.
---
Outside diff comments:
In `@docs/CASE_COCKPIT.md`:
- Around line 17-20: The documentation for running the case cockpit via the
direct streamlit command is incomplete because it omits the required package
installation step that the .bat and .ps1 scripts correctly include. Update the
documentation to either include the py -3 -m pip install -e . installation step
before the streamlit run command, or replace it with the entry point command
aios-case-cockpit (which requires the package to be installed first). Reference
the pyproject.toml file where the aios-case-cockpit entry point is defined to
ensure consistency with the project setup.
In `@RUN_AIOS_CASE_COCKPIT.bat`:
- Around line 1-14: The RUN_AIOS_CASE_COCKPIT.bat file currently uses LF
(Unix-style) line endings, which can cause parsing issues in Windows batch
scripts. Convert all line endings in this file from LF to CRLF (Windows-style)
to ensure proper batch file execution on Windows systems. This can be done
through your text editor's line ending conversion feature or by using Git
configuration to normalize line endings for batch files.
---
Nitpick comments:
In `@src/aios_habit/case_models.py`:
- Line 17: The `created_at` field (and other timestamp fields on lines 25-26)
are using local naive timestamps via `datetime.now().isoformat()`, which creates
inconsistent timestamp formats in JSONL records. Replace the lambda default
factories for all timestamp fields to use the canonical UTC timestamp helper
that is already exposed in `src/aios_habit/core.py:26-27`, ensuring all
persisted model defaults use consistent UTC timestamp semantics across the
codebase.
In `@tests/test_case_cockpit.py`:
- Around line 17-23: The current test_csv_ingest_reads_synthetic_file test only
verifies basic CSV ingest functionality but lacks regression tests for the
security hardening features. Add additional test functions that specifically
assert filename sanitization behavior in upload persistence (verify that special
characters or path traversal attempts in filenames are properly sanitized before
being persisted) and test that evidence marked with local_only flag is properly
excluded from prompt construction (verify that such evidence does not appear in
the extracted text or other prompt-related outputs). These regression tests will
lock in the hardening behavior and prevent future regressions of these security
features.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 5a5fe4f4-9a2f-4109-878c-380725637b82
📒 Files selected for processing (22)
.gitignoreCHANGELOG.mdPROJECT_HANDOVER.mdREADME.mdRUN_AIOS_CASE_COCKPIT.batRUN_AIOS_HABIT_STUDIO.batdocs/CASE_COCKPIT.mddocs/INSTALL.mddocs/OPERATOR_RUNBOOK.mddocs/STUDIO_UI.mdpyproject.tomlscripts/run_case_cockpit.ps1scripts/run_studio.ps1src/aios_habit/case_actions.pysrc/aios_habit/case_cockpit.pysrc/aios_habit/case_graph.pysrc/aios_habit/case_ingest.pysrc/aios_habit/case_models.pysrc/aios_habit/case_store.pysrc/aios_habit/studio.pytests/test_aios_habit.pytests/test_case_cockpit.py
✅ Files skipped from review due to trivial changes (5)
- docs/STUDIO_UI.md
- PROJECT_HANDOVER.md
- CHANGELOG.md
- .gitignore
- README.md
🚧 Files skipped from review as they are similar to previous changes (1)
- tests/test_aios_habit.py
| py -3 -m pip show streamlit pandas openpyxl >nul 2>&1 | ||
| if %errorlevel% neq 0 ( | ||
| echo Dependencies missing. Installing... | ||
| py -3 -m pip install -e . | ||
| ) | ||
|
|
||
| echo Starting AIOS Case Cockpit... | ||
| py -3 -m streamlit run src\aios_habit\case_cockpit.py |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
cd /tmp && find . -maxdepth 3 -type f \( -name "setup.py" -o -name "pyproject.toml" -o -name "setup.cfg" \) 2>/dev/null | head -20Repository: Nakazasen/AIOS_habbit
Length of output: 47
🏁 Script executed:
ls -laRepository: Nakazasen/AIOS_habbit
Length of output: 3822
🏁 Script executed:
git ls-files | head -30Repository: Nakazasen/AIOS_habbit
Length of output: 1065
🏁 Script executed:
cat pyproject.tomlRepository: Nakazasen/AIOS_habbit
Length of output: 650
🏁 Script executed:
ls -la src/Repository: Nakazasen/AIOS_habbit
Length of output: 239
🏁 Script executed:
cat RUN_AIOS_CASE_COCKPIT.batRepository: Nakazasen/AIOS_habbit
Length of output: 388
🏁 Script executed:
find src -type f -name "*.py" | head -20Repository: Nakazasen/AIOS_habbit
Length of output: 612
Dependency check misses local package installation.
Line 5 only verifies third-party packages (streamlit, pandas, openpyxl). If those are already installed but the local aios_habit package is not, the check passes and Line 8 is skipped, causing Line 12 to fail at import time.
Suggested fix
- py -3 -m pip show streamlit pandas openpyxl >nul 2>&1
+ py -3 -c "import aios_habit, streamlit, pandas, openpyxl" >nul 2>&1
if %errorlevel% neq 0 (
echo Dependencies missing. Installing...
py -3 -m pip install -e .
)🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@RUN_AIOS_CASE_COCKPIT.bat` around lines 5 - 12, The dependency check on line
5 only verifies third-party packages (streamlit, pandas, openpyxl) but does not
check if the local aios_habit package is installed. If the third-party packages
are already present but the local package is missing, the errorlevel check
passes and line 8 (pip install -e .) is skipped, causing the streamlit run
command on line 12 to fail at import time. Modify the pip show command to also
verify the aios_habit package is installed, or add a separate verification step
that attempts to import the aios_habit module to ensure both the third-party
dependencies and the local package are available before running the streamlit
application.
| @echo off | ||
| cd /d "%~dp0" | ||
|
|
||
| echo Checking dependencies... | ||
| py -3 -m pip show streamlit >nul 2>&1 | ||
| if %errorlevel% neq 0 ( | ||
| echo Streamlit not found. Installing dependencies... | ||
| py -3 -m pip install -e . | ||
| ) | ||
|
|
||
| echo Starting AIOS Habit Studio... | ||
| py -3 -m streamlit run src\aios_habit\studio.py | ||
| pause |
There was a problem hiding this comment.
Convert file to Windows line endings (CRLF).
The batch file uses Unix line endings (LF only) which can cause GOTO/CALL label parsing failures on Windows due to documented batch parser boundary bugs. Windows batch files must use CRLF line endings for reliable parsing.
This is not a cosmetic issue—it can cause script malfunction on some Windows systems.
Fix: Convert the file to CRLF using one of these methods:
- In your editor: Save with "Windows (CRLF)" line ending
- Via Git config:
git config core.autocrlf trueto auto-convert on checkout - Via command-line tools:
dos2unix -c DOS RUN_AIOS_HABIT_STUDIO.bat
🧰 Tools
🪛 Blinter (1.0.112)
[error] 1-1: Unix line endings detected. Explanation: Batch file uses Unix line endings (LF-only) which can cause GOTO/CALL label parsing failures and script malfunction due to Windows batch parser 512-byte boundary bugs. Recommendation: Convert file to Windows line endings (CRLF). Use tools like dos2unix, notepad++, or configure git with 'git config core.autocrlf true'. Context: File uses Unix line endings (LF-only) - 13 LF sequences found
(E018)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@RUN_AIOS_HABIT_STUDIO.bat` around lines 1 - 13, The RUN_AIOS_HABIT_STUDIO.bat
file currently uses Unix line endings (LF) which can cause batch parser failures
on Windows systems. Convert the entire file to use Windows line endings (CRLF)
by saving the file with CRLF line ending format in your editor, or using a
command-line tool like dos2unix with the DOS option to perform the conversion.
Ensure that all line breaks in the batch file are properly formatted as CRLF
before committing.
| if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) { | ||
| Write-Host "Dependencies missing. Installing..." | ||
| py -3 -m pip install -e . | ||
| } | ||
|
|
||
| Write-Host "Starting AIOS Case Cockpit..." | ||
| py -3 -m streamlit run src\aios_habit\case_cockpit.py |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
# First, find and examine the case_cockpit.py file
find . -name "case_cockpit.py" -type fRepository: Nakazasen/AIOS_habbit
Length of output: 98
🏁 Script executed:
# Look for dependency files
find . -type f \( -name "setup.py" -o -name "setup.cfg" -o -name "pyproject.toml" -o -name "requirements.txt" \) | head -20Repository: Nakazasen/AIOS_habbit
Length of output: 82
🏁 Script executed:
# Examine the PowerShell script structure
cat -n scripts/run_case_cockpit.ps1Repository: Nakazasen/AIOS_habbit
Length of output: 530
🏁 Script executed:
# Examine case_cockpit.py imports
cat -n src/aios_habit/case_cockpit.py | head -50Repository: Nakazasen/AIOS_habbit
Length of output: 2409
🏁 Script executed:
# Examine pyproject.toml for dependencies
cat -n pyproject.tomlRepository: Nakazasen/AIOS_habbit
Length of output: 832
🏁 Script executed:
# Check if aios_habit package structure exists
find src/aios_habit -type f -name "*.py" | head -20Repository: Nakazasen/AIOS_habbit
Length of output: 612
Launcher should verify all dependencies are importable, not just streamlit command presence.
At lines 5–8, checking Get-Command "streamlit" is insufficient. The case_cockpit.py script imports from the local aios_habit package (lines 8–12), and the launcher can still fail at startup if the package itself isn't installed in that interpreter context, even if the streamlit command exists in PATH.
Suggested fix
-if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
+py -3 -c "import aios_habit, streamlit, pandas, openpyxl" 2>$null
+if ($LASTEXITCODE -ne 0) {
Write-Host "Dependencies missing. Installing..."
py -3 -m pip install -e .
}🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/run_case_cockpit.ps1` around lines 5 - 11, The dependency check at
lines 5-8 only verifies that the streamlit command exists in PATH using
Get-Command, but does not confirm that the aios_habit package is properly
installed and importable in the Python interpreter context that will execute
case_cockpit.py. Replace or supplement the Get-Command check for streamlit with
a Python-based verification that attempts to import both the streamlit module
and the aios_habit package using py -3 to ensure both dependencies are available
in the same interpreter instance before proceeding to launch the application.
| if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) { | ||
| Write-Host "Streamlit not found. Attempting to install dependencies..." | ||
| py -3 -m pip install -e . | ||
| } | ||
|
|
||
| Write-Host "Starting AIOS Habit Studio..." | ||
| py -3 -m streamlit run src\aios_habit\studio.py |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
# First, let's examine the run_studio.ps1 file
cat -n scripts/run_studio.ps1Repository: Nakazasen/AIOS_habbit
Length of output: 588
🏁 Script executed:
# Find where aios_habit is defined to understand if it's a real package
find . -type f -name "*.py" | head -20 | xargs grep -l "aios_habit" | head -5Repository: Nakazasen/AIOS_habbit
Length of output: 96
🏁 Script executed:
# Check for setup.py or pyproject.toml to understand package structure
ls -la | grep -E "setup\.(py|cfg)|pyproject\.toml|requirements"Repository: Nakazasen/AIOS_habbit
Length of output: 130
🏁 Script executed:
# Search for similar dependency checks in other scripts
fd -e "ps1" -o -e "sh" -o -e "bash" | xargs grep -l "pip install" 2>/dev/null | head -5Repository: Nakazasen/AIOS_habbit
Length of output: 182
🏁 Script executed:
# Check the pyproject.toml to understand the package structure
cat pyproject.tomlRepository: Nakazasen/AIOS_habbit
Length of output: 650
🏁 Script executed:
# Look at the studio.py file to understand what it imports
head -20 src/aios_habit/studio.pyRepository: Nakazasen/AIOS_habbit
Length of output: 755
🏁 Script executed:
# Check if there are other scripts with similar patterns
find . -name "run_*.ps1" -o -name "run_*.sh" | xargs cat 2>/dev/nullRepository: Nakazasen/AIOS_habbit
Length of output: 892
🏁 Script executed:
# Verify the structure of aios_habit package
ls -la src/aios_habit/ | head -15Repository: Nakazasen/AIOS_habbit
Length of output: 931
Studio launcher needs to verify both aios_habit and streamlit are importable before proceeding.
The current check at line 6 only verifies the streamlit command exists, which is insufficient. If streamlit is installed globally but aios-habit is not installed via pip install -e ., line 12 will fail when trying to import from aios_habit (see imports in studio.py: audit, discovery, export_pack, handover, models, phase_gate, profiles, storage).
Suggested fix
-# Check if streamlit is available
-if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) {
+py -3 -c "import aios_habit, streamlit" 2>$null
+if ($LASTEXITCODE -ne 0) {
Write-Host "Streamlit not found. Attempting to install dependencies..."
py -3 -m pip install -e .
}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if (-not (Get-Command "streamlit" -ErrorAction SilentlyContinue)) { | |
| Write-Host "Streamlit not found. Attempting to install dependencies..." | |
| py -3 -m pip install -e . | |
| } | |
| Write-Host "Starting AIOS Habit Studio..." | |
| py -3 -m streamlit run src\aios_habit\studio.py | |
| py -3 -c "import aios_habit, streamlit" 2>$null | |
| if ($LASTEXITCODE -ne 0) { | |
| Write-Host "Streamlit not found. Attempting to install dependencies..." | |
| py -3 -m pip install -e . | |
| } | |
| Write-Host "Starting AIOS Habit Studio..." | |
| py -3 -m streamlit run src\aios_habit\studio.py |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/run_studio.ps1` around lines 6 - 12, The current check only verifies
that the streamlit command exists but does not verify that the aios_habit
package is installed, which will cause the script to fail at runtime when
studio.py tries to import from the aios_habit module. Modify the dependency
verification logic to check if both streamlit and aios_habit are importable
using Python import checks, and only skip the pip install step if both are
successfully importable. This ensures that running py -3 -m streamlit run
src\aios_habit\studio.py will have all required dependencies available.
| except Exception: | ||
| pass |
There was a problem hiding this comment.
Stop silently dropping invalid JSONL records.
except Exception: pass hides parse/schema failures and silently excludes persisted records from the app, which is a data-loss failure mode.
Suggested patch
+import logging
@@
- except Exception:
- pass
+ except (json.JSONDecodeError, TypeError) as exc:
+ logging.warning("Skipping invalid case record in %s: %s", CASES_FILE, exc)
@@
- except Exception:
- pass
+ except (json.JSONDecodeError, TypeError) as exc:
+ logging.warning("Skipping invalid evidence record in %s: %s", EVIDENCE_FILE, exc)Also applies to: 44-45
🧰 Tools
🪛 Ruff (0.15.17)
[error] 31-32: try-except-pass detected, consider logging the exception
(S110)
[warning] 31-31: Do not catch blind exception: Exception
(BLE001)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_store.py` around lines 31 - 32, The bare except
Exception: pass blocks at lines 31-32 and 44-45 are silently suppressing parse
and schema validation errors when loading JSONL records, causing data loss
without any notification or logging. Replace these silent exception blocks with
proper error handling that logs the specific error details, includes information
about which record failed, and either re-raises the exception or handles it in a
way that makes the failure visible to the user instead of silently excluding
persisted records from the application.
Source: Linters/SAST tools
| with open(CASES_FILE, 'w', encoding='utf-8') as f: | ||
| for c in cases: | ||
| f.write(json.dumps(asdict(c), ensure_ascii=False) + '\n') | ||
|
|
There was a problem hiding this comment.
Avoid non-atomic whole-file rewrites for JSONL persistence.
Writing directly to the target file can leave partial files on interruption and can lose updates when multiple sessions write close together.
Suggested direction
+import tempfile
@@
+def _write_jsonl_atomic(path: Path, rows: list[dict]) -> None:
+ fd, tmp = tempfile.mkstemp(dir=path.parent, suffix=".tmp")
+ try:
+ with os.fdopen(fd, "w", encoding="utf-8") as f:
+ for row in rows:
+ f.write(json.dumps(row, ensure_ascii=False) + "\n")
+ os.replace(tmp, path)
+ finally:
+ if os.path.exists(tmp):
+ os.remove(tmp)
@@
- with open(CASES_FILE, 'w', encoding='utf-8') as f:
- for c in cases:
- f.write(json.dumps(asdict(c), ensure_ascii=False) + '\n')
+ _write_jsonl_atomic(CASES_FILE, [asdict(c) for c in cases])
@@
- with open(EVIDENCE_FILE, 'w', encoding='utf-8') as f:
- for e in items:
- f.write(json.dumps(asdict(e), ensure_ascii=False) + '\n')
+ _write_jsonl_atomic(EVIDENCE_FILE, [asdict(e) for e in items])Also applies to: 76-78
🧰 Tools
🪛 ast-grep (0.43.0)
[info] 61-61: use jsonify instead of json.dumps for JSON output
Context: json.dumps(asdict(c), ensure_ascii=False)
Note: Security best practice.
(use-jsonify)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_store.py` around lines 60 - 63, The current
implementation of save_cases and the related code at lines 76-78 write directly
to CASES_FILE, which is not atomic and can leave partial files on interruption
or cause data loss when multiple sessions write concurrently. Refactor the file
write operation to first write all JSON lines to a temporary file, then
atomically rename that temporary file to CASES_FILE to ensure the entire write
either succeeds completely or fails without corrupting the target file.
| if col1.button("Approve as Verified", key=f"app_{c_file.name}"): | ||
| if not data.get("evidence_ids"): | ||
| st.error("Cannot approve: Missing evidence_ids") | ||
| else: | ||
| data["status"] = "verified" | ||
| data["export_allowed"] = True | ||
| append_jsonl(MEMORY_PATH, data) | ||
| c_file.write_text(json.dumps(data, indent=2, ensure_ascii=False), encoding="utf-8") | ||
| st.success("Approved and added to Memory Vault.") |
There was a problem hiding this comment.
Review Queue approval can persist invalid and duplicate verified memories.
At Line 219-Line 225, approval only checks that evidence_ids is non-empty, then appends directly. This can (1) mark a memory as verified even when referenced evidence does not exist, and (2) append the same candidate repeatedly on re-approval.
Suggested fix
if col1.button("Approve as Verified", key=f"app_{c_file.name}"):
- if not data.get("evidence_ids"):
+ if not data.get("evidence_ids"):
st.error("Cannot approve: Missing evidence_ids")
else:
- data["status"] = "verified"
- data["export_allowed"] = True
- append_jsonl(MEMORY_PATH, data)
- c_file.write_text(json.dumps(data, indent=2, ensure_ascii=False), encoding="utf-8")
- st.success("Approved and added to Memory Vault.")
+ existing_evs = {e.get("evidence_id") for e in load_evidence()}
+ missing = [e for e in data.get("evidence_ids", []) if e not in existing_evs]
+ if missing:
+ st.error(f"Cannot approve: Missing evidence records: {missing}")
+ else:
+ data["status"] = "verified"
+ data["export_allowed"] = True
+ existing_mem_ids = {m.get("memory_id") for m in load_memory()}
+ if data.get("memory_id") in existing_mem_ids:
+ st.warning("Memory already exists in vault; skipping duplicate append.")
+ else:
+ append_jsonl(MEMORY_PATH, data)
+ c_file.write_text(json.dumps(data, indent=2, ensure_ascii=False), encoding="utf-8")
+ st.success("Approved and added to Memory Vault.")🧰 Tools
🪛 ast-grep (0.43.0)
[info] 224-224: use jsonify instead of json.dumps for JSON output
Context: json.dumps(data, indent=2, ensure_ascii=False)
Note: Security best practice.
(use-jsonify)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/studio.py` around lines 218 - 226, The approval logic in the
button handler checks only that evidence_ids is non-empty but does not validate
that the referenced evidence actually exists, and it unconditionally appends to
MEMORY_PATH without checking for duplicates. To fix this, before setting status
to verified and calling append_jsonl, add validation to confirm that each
evidence_id referenced in data actually exists in the evidence store, and
implement a check to prevent duplicate entries in MEMORY_PATH by either
verifying the record does not already exist before appending or by replacing an
existing record if it does. Ensure the success message is only shown after
successful validation and append.
There was a problem hiding this comment.
Actionable comments posted: 7
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/aios_habit/case_audit.py`:
- Around line 31-33: The audit function hardcodes the assets directory path
using Path.cwd() instead of using the store's configured assets root, which
causes inconsistency when running from different working directories. Replace
the hardcoded path construction for the assets_dir variable with a call to
retrieve the assets root from the case store configuration (similar to how
ingest/save operations obtain it), ensuring the audit uses the same asset
location as other store operations regardless of the current working directory.
- Around line 64-68: The leak detection logic in the conditional block checking
privacy_level and prompt_outputs is too simplistic - it only verifies if the
complete ev.extracted_text string appears within prompt_text as a full
substring, but this misses cases where the extracted text is truncated or
fragmented across the prompt. Instead of just checking `ev.extracted_text in
prompt_text`, implement a more robust detection mechanism that can identify
partial or fragmented versions of the sensitive text (such as checking if
substantial portions or key segments of ev.extracted_text appear within
prompt_text, or using fuzzy matching to catch truncated leaks), ensuring that
both complete and partial exposures of local_only evidence are properly caught
and flagged in the errors list.
- Around line 55-61: The path containment validation in the try block uses
string prefix matching with startswith() which is unreliable and can be bypassed
by similarly-prefixed sibling directories. Replace the string-based containment
check with proper path operations using Path methods like is_relative_to() or by
comparing the resolved paths using Path.relative_to() within a try-except block.
Additionally, replace the overly broad Exception catch with specific exception
types that Path operations can actually raise, such as ValueError or OSError, to
properly handle only the expected error conditions when resolving or validating
paths.
In `@src/aios_habit/case_graph.py`:
- Around line 31-35: The evidence node ID generation on line 34 uses raw
ev.evidence_id directly, which is inconsistent with the synthetic ID pattern
used for hypotheses, actions, and decisions elsewhere in the function. Replace
the raw ev.evidence_id in the ev_node assignment with a synthetic ID generated
via the nid function (similar to how hypotheses, actions, and decisions use
indexed synthetic IDs like nid("HYP"), nid("ACT"), nid("DEC")). This ensures all
node identifiers follow a consistent, safe alphanumeric format that aligns with
Mermaid best practices and prevents fragility if ID formats change in the
future.
In `@src/aios_habit/case_prompt.py`:
- Around line 19-27: The evidence items being processed in the loop are not
filtered by the active case's case_id, which can cause evidence from other cases
to leak into the prompt. Before iterating through evidence_items in the loop
that starts with "for e in evidence_items:", add a filter to only include items
where e.case_id matches the current case's case_id. This filtering should happen
before the loop begins so that only evidence belonging to the active case is
included in the prompt assembly.
- Around line 8-13: The cloud targets (gemini, gpt, copilot) in the conditional
statement currently allow bypassing local-only protections by setting
should_include_local_only equal to the include_local_only parameter. To fix this
security issue, change the assignment for cloud targets (gemini, gpt, copilot)
from should_include_local_only = include_local_only to should_include_local_only
= False, ensuring local-only protected content cannot be sent to cloud services
regardless of the include_local_only parameter value, consistent with the
privacy hardening already applied to notebooklm_safe.
In `@tests/test_case_cockpit_hardening.py`:
- Around line 154-155: The bat_content and ps1_content assignments on lines
154-155 use relative paths that depend on the current working directory, causing
failures when tests run from different locations. Make these paths absolute by
anchoring them to the test file location using __file__ and constructing paths
relative to the test file's directory, so the Path().read_text() calls work
regardless of the invocation directory.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: a80618ce-e3fc-4646-90ed-59ccfebea0ed
📒 Files selected for processing (23)
CASE_COCKPIT_ACCEPTANCE_CRITERIA.mdCHANGELOG.mdDEPRECATION_PAUSE_LIST.mdHARVEST_BACKLOG.mdINTEGRATION_DECISION_LOG.mdMIGRATION_POLICY.mdMONDAY_PILOT_CHECKLIST.mdPRODUCT_NORTH_STAR.mdPROJECT_HANDOVER.mdREPO_INHERITANCE_MAP.mdWORKLENS_ARCHITECTURE.mdWORKLENS_MASTER_ROADMAP.mddocs/CASE_COCKPIT.mddocs/inheritance_audit/ABW_NVIDIA_FUSION_CONTROL_AUDIT.mddocs/inheritance_audit/INHERITANCE_SUMMARY.mddocs/inheritance_audit/Nvidia_AUDIT.mddocs/inheritance_audit/skill_Anti_brain_wiki_note_AUDIT.mdsrc/aios_habit/case_audit.pysrc/aios_habit/case_cockpit.pysrc/aios_habit/case_graph.pysrc/aios_habit/case_ingest.pysrc/aios_habit/case_prompt.pytests/test_case_cockpit_hardening.py
✅ Files skipped from review due to trivial changes (11)
- DEPRECATION_PAUSE_LIST.md
- MIGRATION_POLICY.md
- INTEGRATION_DECISION_LOG.md
- HARVEST_BACKLOG.md
- docs/inheritance_audit/INHERITANCE_SUMMARY.md
- WORKLENS_ARCHITECTURE.md
- PRODUCT_NORTH_STAR.md
- MONDAY_PILOT_CHECKLIST.md
- REPO_INHERITANCE_MAP.md
- PROJECT_HANDOVER.md
- CHANGELOG.md
🚧 Files skipped from review as they are similar to previous changes (1)
- src/aios_habit/case_cockpit.py
| cwd = Path.cwd().resolve() | ||
| assets_dir = (cwd / "local_cases" / "assets").resolve() | ||
|
|
There was a problem hiding this comment.
Audit asset root should use store configuration, not current working directory.
Lines 31-33 hardcode Path.cwd()/local_cases/assets, but ingest/save uses the case store assets root. Running from a different CWD can produce false FAIL audits.
Suggested fix
from pathlib import Path
from typing import List, Dict, Optional, Any
from .case_models import Case, EvidenceItem
+from .case_store import ASSETS_DIR
@@
- cwd = Path.cwd().resolve()
- assets_dir = (cwd / "local_cases" / "assets").resolve()
+ assets_dir = ASSETS_DIR.resolve()🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_audit.py` around lines 31 - 33, The audit function
hardcodes the assets directory path using Path.cwd() instead of using the
store's configured assets root, which causes inconsistency when running from
different working directories. Replace the hardcoded path construction for the
assets_dir variable with a call to retrieve the assets root from the case store
configuration (similar to how ingest/save operations obtain it), ensuring the
audit uses the same asset location as other store operations regardless of the
current working directory.
| if ev.source_path and ev.source_path not in ("clipboard", "manual"): | ||
| try: | ||
| p = Path(ev.source_path).resolve() | ||
| if not str(p).startswith(str(assets_dir)): | ||
| errors.append(f"Evidence item {ev.evidence_id} asset path '{ev.source_path}' is outside assets directory '{assets_dir}'.") | ||
| except Exception as e: | ||
| errors.append(f"Evidence item {ev.evidence_id} has invalid asset path: {e}") |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
cd /tmp && find . -name "case_audit.py" -type f 2>/dev/null | head -5Repository: Nakazasen/AIOS_habbit
Length of output: 47
🏁 Script executed:
git ls-files | grep -i case_auditRepository: Nakazasen/AIOS_habbit
Length of output: 94
🏁 Script executed:
cat -n src/aios_habit/case_audit.py | sed -n '50,70p'Repository: Nakazasen/AIOS_habbit
Length of output: 1588
🏁 Script executed:
cat -n src/aios_habit/case_audit.py | head -40Repository: Nakazasen/AIOS_habbit
Length of output: 1737
🏁 Script executed:
cat -n src/aios_habit/case_audit.py | sed -n '25,65p'Repository: Nakazasen/AIOS_habbit
Length of output: 2574
Use path-aware containment checks instead of string prefix matching.
Line 58 uses startswith() on path strings, which can be bypassed by prefix-like siblings (e.g., /local_cases/assets_backup/ would match if checking only string prefix). This is not robust path validation. Additionally, line 60 catches a blind Exception instead of the specific exceptions that should be handled.
Suggested fix
if ev.source_path and ev.source_path not in ("clipboard", "manual"):
try:
p = Path(ev.source_path).resolve()
- if not str(p).startswith(str(assets_dir)):
- errors.append(f"Evidence item {ev.evidence_id} asset path '{ev.source_path}' is outside assets directory '{assets_dir}'.")
- except Exception as e:
+ p.relative_to(assets_dir)
+ except ValueError:
+ errors.append(
+ f"Evidence item {ev.evidence_id} asset path '{ev.source_path}' is outside assets directory '{assets_dir}'."
+ )
+ except OSError as e:
errors.append(f"Evidence item {ev.evidence_id} has invalid asset path: {e}")🧰 Tools
🪛 Ruff (0.15.17)
[warning] 60-60: Do not catch blind exception: Exception
(BLE001)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_audit.py` around lines 55 - 61, The path containment
validation in the try block uses string prefix matching with startswith() which
is unreliable and can be bypassed by similarly-prefixed sibling directories.
Replace the string-based containment check with proper path operations using
Path methods like is_relative_to() or by comparing the resolved paths using
Path.relative_to() within a try-except block. Additionally, replace the overly
broad Exception catch with specific exception types that Path operations can
actually raise, such as ValueError or OSError, to properly handle only the
expected error conditions when resolving or validating paths.
Source: Linters/SAST tools
| if ev.privacy_level == "local_only" and prompt_outputs: | ||
| for target, prompt_text in prompt_outputs.items(): | ||
| if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"): | ||
| if has_text and ev.extracted_text in prompt_text: | ||
| errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.") |
There was a problem hiding this comment.
Leak detection misses partial local-only text exposure.
Line 67 checks only full-string inclusion (ev.extracted_text in prompt_text). If prompt generation includes a truncated snippet, this audit can pass despite leakage.
Suggested fix
if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"):
- if has_text and ev.extracted_text in prompt_text:
+ leak_probe = (ev.extracted_text or "")[:200]
+ if has_text and leak_probe and leak_probe in prompt_text:
errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.")📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if ev.privacy_level == "local_only" and prompt_outputs: | |
| for target, prompt_text in prompt_outputs.items(): | |
| if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"): | |
| if has_text and ev.extracted_text in prompt_text: | |
| errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.") | |
| if ev.privacy_level == "local_only" and prompt_outputs: | |
| for target, prompt_text in prompt_outputs.items(): | |
| if target.lower() in ("notebooklm_safe", "gemini", "gpt", "copilot"): | |
| leak_probe = (ev.extracted_text or "")[:200] | |
| if has_text and leak_probe and leak_probe in prompt_text: | |
| errors.append(f"local_only evidence '{ev.evidence_id}' raw extracted_text was leaked in target '{target}'.") |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_audit.py` around lines 64 - 68, The leak detection logic
in the conditional block checking privacy_level and prompt_outputs is too
simplistic - it only verifies if the complete ev.extracted_text string appears
within prompt_text as a full substring, but this misses cases where the
extracted text is truncated or fragmented across the prompt. Instead of just
checking `ev.extracted_text in prompt_text`, implement a more robust detection
mechanism that can identify partial or fragmented versions of the sensitive text
(such as checking if substantial portions or key segments of ev.extracted_text
appear within prompt_text, or using fuzzy matching to catch truncated leaks),
ensuring that both complete and partial exposures of local_only evidence are
properly caught and flagged in the errors list.
| for ev in evidence_items: | ||
| if ev.case_id == case.case_id: | ||
| label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}" | ||
| ev_node = f'{ev.evidence_id}["{label}"]' | ||
| lines.append(f'C -->|contains| {ev_node}') |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
# First, let's examine the actual file and context around the mentioned lines
cat -n src/aios_habit/case_graph.py | head -60Repository: Nakazasen/AIOS_habbit
Length of output: 2300
🏁 Script executed:
# Find the safe_mermaid_label function definition
rg "def safe_mermaid_label" -A 10 src/Repository: Nakazasen/AIOS_habbit
Length of output: 915
🏁 Script executed:
# Check how evidence_id is defined/used
rg "evidence_id" src/aios_habit/ -B 2 -A 2Repository: Nakazasen/AIOS_habbit
Length of output: 21358
🏁 Script executed:
# Look at the full generate_case_mermaid function to understand the complete context
rg "def generate_case_mermaid" -A 50 src/aios_habit/case_graph.pyRepository: Nakazasen/AIOS_habbit
Length of output: 1098
🏁 Script executed:
# Find the nid() function definition
rg "def nid" -A 5 src/Repository: Nakazasen/AIOS_habbit
Length of output: 381
🌐 Web query:
Mermaid node ID syntax requirements restrictions characters allowed
💡 Result:
In Mermaid flowchart syntax, node IDs serve as the internal identifiers for nodes and are distinct from their displayed labels [1]. While Mermaid does not strictly define a single character set for IDs, the parser is sensitive to certain keywords and patterns that can cause rendering errors or conflicts [2][3][1][4]. Core Requirements and Restrictions: Reserved Words: You should avoid using Mermaid reserved words as bare node IDs, as these can conflict with the parser and cause the diagram to break [4][5]. Reserved words include: end, graph, subgraph, click, style, linkStyle, class, classDef, default, call, href, and interpolate [4][5]. If you must use these as labels, define the node with a safe ID and provide the reserved word as a quoted label (e.g., safeId["end"]) [4][5]. Starting Characters: Node IDs that begin with specific characters like 'o' or 'x' can sometimes be misinterpreted by the parser as edge types (e.g., leading to unintended edge styling or shapes) [5]. It is best practice to use descriptive alphanumeric IDs or camelCase [5][6]. Special Characters: While some special characters (like '.' or '/') may be permitted in certain contexts, they can lead to ambiguous or broken syntax [3][7]. To avoid issues with spaces or special characters in labels, always use double quotes [8][1][4]. Best Practices for Node IDs: Use Descriptive, Alphanumeric IDs: Adopt naming conventions like camelCase (e.g., orderService) rather than generic labels like s1 or single letters [5][6]. This improves readability and reduces the risk of collisions with parser keywords [7][6]. Separate IDs from Labels: Always explicitly define nodes if they require complex text or reserved words, using the format ID["Display Text"] [1][5]. This ensures the ID remains safe for internal referencing while the label displays correctly [1]. If a diagram fails to render, check for node IDs that match reserved keywords or contain unexpected special characters, and verify that all labels with spaces or symbols are enclosed in double quotes [1][4].
Citations:
- 1: https://mermaid.ai/docs/build-and-edit/write-diagram-syntax
- 2: https://github.com/mermaid-js/mermaid/blob/develop/docs/intro/syntax-reference.md
- 3: Bug/4645 graph node containing keyword mermaid-js/mermaid#4657
- 4: https://github.com/product-on-purpose/pm-skills/blob/main/skills/utility-mermaid-diagrams/references/syntax-guide.md
- 5: https://github.com/awesome-skills/mermaid-syntax-skill/blob/main/SKILL.md
- 6: https://github.com/tech-leads-club/agent-skills/blob/main/packages/skills-catalog/skills/(tooling)/mermaid-studio/SKILL.md
- 7: How to manage node id names in flowchart? mermaid-js/mermaid#2650
- 8: https://github.com/mermaid-js/mermaid/blob/9745f325/docs/syntax/flowchart.md
Use synthetic node IDs for evidence items to maintain consistency and safety.
Line 34 uses raw ev.evidence_id directly as the Mermaid node identifier, while hypotheses (line 38), actions (line 42), and decisions (line 46) use indexed synthetic IDs. Evidence node IDs generated via nid("EVD") produce format "EVD-{hex}" with hyphens. While this may render, Mermaid best practices favor pure alphanumeric IDs, and the current approach creates an inconsistent pattern and fragility if ID format changes.
Suggested fix
- for ev in evidence_items:
+ ev_index = 0
+ for ev in evidence_items:
if ev.case_id == case.case_id:
label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}"
- ev_node = f'{ev.evidence_id}["{label}"]'
- lines.append(f'C -->|contains| {ev_node}')
+ ev_node_id = f"EV{ev_index}"
+ ev_index += 1
+ lines.append(f'{ev_node_id}["{label}"]')
+ lines.append(f'C -->|contains| {ev_node_id}')📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| for ev in evidence_items: | |
| if ev.case_id == case.case_id: | |
| label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}" | |
| ev_node = f'{ev.evidence_id}["{label}"]' | |
| lines.append(f'C -->|contains| {ev_node}') | |
| ev_index = 0 | |
| for ev in evidence_items: | |
| if ev.case_id == case.case_id: | |
| label = f"{safe_mermaid_label(ev.source_type)}: {safe_mermaid_label(ev.title)}" | |
| ev_node_id = f"EV{ev_index}" | |
| ev_index += 1 | |
| lines.append(f'{ev_node_id}["{label}"]') | |
| lines.append(f'C -->|contains| {ev_node_id}') |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_graph.py` around lines 31 - 35, The evidence node ID
generation on line 34 uses raw ev.evidence_id directly, which is inconsistent
with the synthetic ID pattern used for hypotheses, actions, and decisions
elsewhere in the function. Replace the raw ev.evidence_id in the ev_node
assignment with a synthetic ID generated via the nid function (similar to how
hypotheses, actions, and decisions use indexed synthetic IDs like nid("HYP"),
nid("ACT"), nid("DEC")). This ensures all node identifiers follow a consistent,
safe alphanumeric format that aligns with Mermaid best practices and prevents
fragility if ID formats change in the future.
| if target_lower == "notebooklm_safe": | ||
| should_include_local_only = False | ||
| elif target_lower in ("gemini", "gpt", "copilot"): | ||
| should_include_local_only = include_local_only | ||
| elif target_lower == "local_ai": | ||
| should_include_local_only = include_local_only |
There was a problem hiding this comment.
Cloud targets can bypass local_only protections via include_local_only=True.
For gemini/gpt/copilot, should_include_local_only = include_local_only permits explicit leakage to cloud prompts. That conflicts with public-safe privacy hardening.
Suggested fix
- elif target_lower in ("gemini", "gpt", "copilot"):
- should_include_local_only = include_local_only
+ elif target_lower in ("gemini", "gpt", "copilot"):
+ should_include_local_only = False🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_prompt.py` around lines 8 - 13, The cloud targets
(gemini, gpt, copilot) in the conditional statement currently allow bypassing
local-only protections by setting should_include_local_only equal to the
include_local_only parameter. To fix this security issue, change the assignment
for cloud targets (gemini, gpt, copilot) from should_include_local_only =
include_local_only to should_include_local_only = False, ensuring local-only
protected content cannot be sent to cloud services regardless of the
include_local_only parameter value, consistent with the privacy hardening
already applied to notebooklm_safe.
| for e in evidence_items: | ||
| if e.privacy_level == "local_only": | ||
| if not should_include_local_only: | ||
| excluded_any = True | ||
| lines.append(f"- [{e.source_type}] {e.title}: [EXCLUDED FOR PRIVACY - local_only]") | ||
| continue | ||
|
|
||
| snippet = e.extracted_text[:200] if e.extracted_text else "" | ||
| lines.append(f"- [{e.source_type}] {e.title}: {snippet}...") |
There was a problem hiding this comment.
Scope evidence to the active case before prompt assembly.
The current flow summarizes every item passed in, without filtering by case.case_id. That can leak another case’s evidence into a prompt pack.
Suggested fix
-def summarize_evidence_for_prompt(evidence_items: List[EvidenceItem], target: str, include_local_only: bool = False) -> str:
+def summarize_evidence_for_prompt(
+ evidence_items: List[EvidenceItem],
+ target: str,
+ include_local_only: bool = False,
+ case_id: str | None = None,
+) -> str:
@@
- for e in evidence_items:
+ for e in evidence_items:
+ if case_id is not None and e.case_id != case_id:
+ continue
if e.privacy_level == "local_only":
@@
- prompt += summarize_evidence_for_prompt(evidence_items, target, include_local_only)
+ prompt += summarize_evidence_for_prompt(
+ evidence_items, target, include_local_only, case_id=case.case_id
+ )Also applies to: 36-40
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/aios_habit/case_prompt.py` around lines 19 - 27, The evidence items being
processed in the loop are not filtered by the active case's case_id, which can
cause evidence from other cases to leak into the prompt. Before iterating
through evidence_items in the loop that starts with "for e in evidence_items:",
add a filter to only include items where e.case_id matches the current case's
case_id. This filtering should happen before the loop begins so that only
evidence belonging to the active case is included in the prompt assembly.
| bat_content = Path("RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8") | ||
| ps1_content = Path("scripts/run_case_cockpit.ps1").read_text(encoding="utf-8") |
There was a problem hiding this comment.
Make launcher path assertions independent of working directory.
Lines 154-155 depend on current working directory. This can fail in IDE/CI invocations that run tests from another path.
Suggested fix
def test_launcher_contents_check():
- bat_content = Path("RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
- ps1_content = Path("scripts/run_case_cockpit.ps1").read_text(encoding="utf-8")
+ repo_root = Path(__file__).resolve().parents[1]
+ bat_content = (repo_root / "RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8")
+ ps1_content = (repo_root / "scripts" / "run_case_cockpit.ps1").read_text(encoding="utf-8")📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| bat_content = Path("RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8") | |
| ps1_content = Path("scripts/run_case_cockpit.ps1").read_text(encoding="utf-8") | |
| def test_launcher_contents_check(): | |
| repo_root = Path(__file__).resolve().parents[1] | |
| bat_content = (repo_root / "RUN_AIOS_CASE_COCKPIT.bat").read_text(encoding="utf-8") | |
| ps1_content = (repo_root / "scripts" / "run_case_cockpit.ps1").read_text(encoding="utf-8") |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/test_case_cockpit_hardening.py` around lines 154 - 155, The bat_content
and ps1_content assignments on lines 154-155 use relative paths that depend on
the current working directory, causing failures when tests run from different
locations. Make these paths absolute by anchoring them to the test file location
using __file__ and constructing paths relative to the test file's directory, so
the Path().read_text() calls work regardless of the invocation directory.
Hardens AIOS Habit public MVP with privacy guardrails, real phase gates, tests, docs, CI, and Python source formatting integrity.
Summary by CodeRabbit