fix(TER-9): vibe dispatch + provider stores (NO-GO wholesale — extract only) - #6
Conversation
- Built comprehensive CLI based on deepcli architecture with deepterm integration - Created core module with Mistralai API wrapper, session management, and chat dispatching - Implemented harvesters module for code extraction, search, and analysis - Added tools module with file, git, network, and Termux utilities - Created integrations module for DeepSeek, ArchW1z, and Synthegration compatibility - Set up proper dev/staging/prod directory structure in sandbox - Added comprehensive test suite with 30+ passing tests - Created setup.py and requirements.txt for package installation - Added detailed README.md documentation - Maintained backward compatibility with existing cli.py interface All for One; and, .One for All! Closes #8f2bf066-28ab-421e-b613-ea634776e672 Co-authored-by: timerloggedout-spec <timerloggedout-spec@users.noreply.github.com>
…ing CLIs - Removed deepseek_integration.py, archwiz_integration.py, synthegration_integration.py - These were unnecessary because we should call existing CLIs as subprocesses - Updated main.py and mistralai_cli.py to properly select and run providers - multi-ai-cli now serves as the main entry point that can: - Use native Mistral implementation - Call existing deepcli as subprocess - Call existing deepcli-tui as subprocess - Updated README.md to reflect correct architecture and usage - All 30 tests still passing The key insight: existing DeepSeek CLI and TUI already work perfectly. We don't need to integrate their internals - just call them as subprocesses. All for One; and, .One for All! Co-authored-by: timerloggedout-spec <timerloggedout-spec@users.noreply.github.com>
…ponses
- Updated CodeHarvester to follow cli-synthegration Codex pattern
- Uses lightweight regex (not BeautifulSoup) for code block extraction
- Implements Pointer class for content-addressable references
- Implements CodexIndex with hierarchical taxonomy (language -> project -> session)
- Stores code blobs in ~/.mistralai-cli/codex/blobs/{hash}.blob
- Enables autonomous agentic CLI interactions to directly reference code blocks
- All 31 tests passing
Key pattern from cli-synthegration:
- CODE_BLOCK_PATTERN = r''
- Pointer(session_id, message_index, block_index, content_hash)
- Blob storage with SHA256 hash (first 16 chars)
- Taxonomy: language -> project -> role
This allows MistralAI to have the same code block isolation and
reference capabilities as the existing DeepSeek/ArchW1z/Synthegration
architecture.
All for One; and, .One for All!
Co-authored-by: timerloggedout-spec <timerloggedout-spec@users.noreply.github.com>
…ders
Created comprehensive provider system replicating Codex pattern for:
**Existing Providers:**
- mistral (native implementation with full Codex)
- deepseek (calls existing deepcli as subprocess)
- claude (uses existing backend)
- gemini (uses existing backend)
- colab (uses existing backend)
**ChapitoAI Providers:**
- ai_studio
- anthropic
- duckduckgo
- grok
- kimi
- openai
- perplexity
- qwen
**Key Features:**
- BaseProvider class with Codex harvesting support
- ProviderConfig for unified configuration
- CodexIndex with hierarchical taxonomy (language -> project -> session)
- Pointer class for content-addressable code blocks
- CodeBlock class for extracted code with hashing
- All providers support: send_message, create_session, get_history
- All providers support: harvest_code, search_code, get_code_by_hash
- Automatic code harvesting from session responses
**Codex Pattern Replicated:**
- CODE_BLOCK_PATTERN = r''
- SHA256 hashing (first 16 chars)
- Blob storage in ~/.multi-ai-cli/codex/{provider}/blobs/
- Taxonomy: language -> project -> role
- Content-addressable via hash
All 31 tests passing.
All for One; and, .One for All!
Co-authored-by: timerloggedout-spec <timerloggedout-spec@users.noreply.github.com>
…affold providers non-live - multi-ai-cli/core/core.py: stderr log instead of silent pass; pass provider/account to update_all - archwiz/dispatch_pipeline.py: resolve mistral/deepcli/multi-ai-cache stores; log step failures - scaffold Chapito-style providers: is_available() -> False
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
📝 WalkthroughWalkthroughThe change adds a unified multi-provider CLI with Mistral and DeepSeek integrations, provider registration, session management, code harvesting and search, analysis tools, Termux utilities, packaging, documentation, and tests. ChangesUnified provider and session flow
Estimated code review effort: 5 (Critical) | ~120 minutes Sequence Diagram(s)sequenceDiagram
participant User
participant UnifiedCLI
participant ProviderRegistry
participant Provider
participant CodexIndex
User->>UnifiedCLI: select provider and send message
UnifiedCLI->>ProviderRegistry: get_provider(provider_name)
ProviderRegistry->>Provider: construct provider
UnifiedCLI->>Provider: send_message(message, session_id)
Provider-->>UnifiedCLI: response and session data
UnifiedCLI->>CodexIndex: index conversation code
CodexIndex-->>UnifiedCLI: search and retrieval metadata
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 8
Note
Due to the large number of review comments, Critical severity comments were prioritized as inline comments.
🟠 Major comments (19)
multi-ai-cli/harvesters/analyzer.py-246-256 (1)
246-256: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
re.DOTALLmakes single-line comment patterns consume the rest of the file.
comment_pattern = r'//.*|/\*.*?\*/|#.*|--.*'is matched withre.findall(comment_pattern, code, re.DOTALL).re.DOTALLmakes.match newlines too, so the greedy//.*and#.*alternatives each match from the first marker to the end of the string, not to the end of the line. For code with multiple line comments, this produces one giant match instead of several, socomment_countandcomplexity["comments"]are wrong for any language that falls through to_analyze_generic.Scope
.to exclude newlines for the line-comment alternatives, keepingre.DOTALLsemantics only for the block-comment alternative.🐛 Proposed fix
- comment_pattern = r'//.*|/\*.*?\*/|#.*|--.*' + comment_pattern = r'//[^\n]*|/\*.*?\*/|#[^\n]*|--[^\n]*' comments = re.findall(comment_pattern, code, re.DOTALL)🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/harvesters/analyzer.py` around lines 246 - 256, Update comment_pattern in the generic analysis flow to prevent the //, #, and -- alternatives from matching across newlines while preserving multiline matching for /* ... */ blocks. Adjust the pattern used by the comment-count calculation before re.findall so each single-line comment is counted separately and result.complexity["comments"] remains accurate.multi-ai-cli/providers/base.py-219-238 (1)
219-238: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
_savediscards per-pointer provider attribution.
_savewrites'provider': self.provider or 'unknown'for every pointer._from_flatreads a per-pointerproviderfield at line 203, andindex_conversationaccepts aproviderargument that can differ fromself.provider(line 275). WhenCodexIndexis created without a provider,base_dirresolves to the sharedcodex/globaldirectory, so one index holds blocks from several providers. Each save then rewrites all of them with the same value, andget_by_providerreturns wrong results after the next load.Store the provider on
Pointer(or in a hash-to-provider map) and write that value.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/base.py` around lines 219 - 238, Update _save and the pointer creation flow used by index_conversation to preserve each pointer’s originating provider, rather than using self.provider for every serialized entry. Store the provider on Pointer or in an equivalent per-content-hash mapping, then have _save write that per-pointer value so _from_flat and get_by_provider retain correct attribution across reloads.multi-ai-cli/providers/claude.py-36-49 (1)
36-49: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift
send_messageacceptssession_idand history, then discards both. Both providers declare asession_idparameter and both callbackend.send_message(message, []). Neither passes the session, and both pass an empty history list. Every message is therefore stateless, whileChatDispatcher.sendforwardssession_idand reports it in the result, so callers believe the session was applied. The shared root cause is that theClaudeWebBackendandColabBackendcalls drop the session and the conversation turns.Load the prior turns for the given
session_idthroughSessionManagerand pass them as the second argument. If the backends cannot accept a session yet, raise an error for a non-defaultsession_idinstead of accepting and discarding it. Also constructSessionManagerand the backend once in__init__rather than on every call.
multi-ai-cli/providers/claude.py#L36-L49: remove the unusedsession_id = "default"assignment at lines 45-47, and pass the resolved history toClaudeWebBackend.send_message.multi-ai-cli/providers/colab.py#L27-L35: read thesession_idparameter, and pass the resolved history toColabBackend.send_message.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/claude.py` around lines 36 - 49, Update multi-ai-cli/providers/claude.py#L36-L49 and multi-ai-cli/providers/colab.py#L27-L35 so each provider constructs SessionManager and its backend once in __init__, loads prior turns for the supplied session_id, and passes that history to the backend send_message call. In claude.py, remove the default session_id reassignment; in colab.py, use the existing session_id parameter. If non-default sessions cannot be supported by either backend, reject them instead of discarding the session.multi-ai-cli/providers/colab.py-17-25 (1)
17-25: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick winColab config points at the DeepSeek cookie file.
Line 23 sets
cookie_path="~/deepseek-cli/cookies_2.json".multi-ai-cli/providers/deepseek.pyline 35 sets the same path for DeepSeek. Colab useshttps://colab.research.google.com, so DeepSeek cookies cannot authenticate it.Any code that loads
cookie_pathfor Colab reads credentials that belong to a different origin and may send them to Google. Set the Colab cookie path to a dedicated file.Line 24 also holds a stray blank line inside the constructor call.
🐛 Proposed fix
return ProviderConfig( name="colab", api_url="https://colab.research.google.com", - cookie_path="~/deepseek-cli/cookies_2.json", - + cookie_path="~/.multi-ai-tokens/colab_cookies.json", )🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/colab.py` around lines 17 - 25, Update ColabProvider.get_default_config so cookie_path points to a Colab-specific cookie file rather than the shared DeepSeek path, and remove the stray blank line inside the ProviderConfig constructor.multi-ai-cli/providers/deepseek.py-56-74 (1)
56-74: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
send_messagecan send an empty session ID and return an empty response.
create_sessionreturns""when_run_deepclifails or when the output holds nosession_id(lines 71-74).send_messagedoes not check the result, so line 62 runsdeepcli send --session "" <message>.
_run_deepclialso returns{}after a non-zero exit code. Line 66 then returns"". The caller cannot tell that failure from an empty model answer.Raise an error when session creation fails and when the command fails.
🐛 Proposed fix
def send_message(self, message: str, session_id: str = None, **kwargs) -> str: """Send a message via deepcli.""" if not session_id: session_id = self.create_session() + if not session_id: + raise RuntimeError("deepcli did not return a session id") # Use deepcli send command result = self._run_deepcli(["send", "--session", session_id, message])🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/deepseek.py` around lines 56 - 74, Update send_message and create_session to detect failed DeepSeek operations instead of propagating empty identifiers or responses. In send_message, raise an error when create_session returns an empty session ID and when _run_deepcli indicates command failure, including the existing dict failure result; in create_session, raise an error when _run_deepcli fails or provides neither session_id nor id, while preserving successful response extraction.multi-ai-cli/providers/deepseek.py-39-54 (1)
39-54: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winAdd a timeout to the
deepclisubprocess call.Line 45 calls
subprocess.runwith notimeoutand with an inherited stdin. Ifdeepcli.pyhangs, waits for input, or blocks on a network call,_run_deepclinever returns. Every provider method routes through it, so the CLI hangs with no way to recover except a signal.
self.config.timeoutalready carries a value for this purpose (ProviderConfig.timeout, default 30).Line 53 also uses a bare
except, which catchesKeyboardInterrupt. Narrow it tojson.JSONDecodeError.🐛 Proposed fix
- cmd = ["python3", str(self.deepcli_path)] + args - result = subprocess.run(cmd, capture_output=True, text=True) + cmd = ["python3", str(self.deepcli_path), *args] + timeout = self.config.timeout if self.config else 30 + try: + result = subprocess.run( + cmd, + capture_output=True, + text=True, + stdin=subprocess.DEVNULL, + timeout=timeout, + check=False, + ) + except subprocess.TimeoutExpired: + console.print(f"[red]deepcli timed out after {timeout}s[/red]") + return {} if result.returncode != 0: console.print(f"[red]deepcli error: {result.stderr}[/red]") return {} try: return json.loads(result.stdout) - except: + except json.JSONDecodeError: return {"output": result.stdout}🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/deepseek.py` around lines 39 - 54, Update _run_deepcli to pass self.config.timeout to subprocess.run so deepcli cannot block indefinitely, and ensure stdin does not remain inherited if the command may wait for input. Replace the bare exception around json.loads with json.JSONDecodeError only, preserving the existing fallback output behavior for invalid JSON.multi-ai-cli/providers/base.py-249-295 (1)
249-295: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winHash computation differs between
index_conversationandextract_from_messages.Line 252 hashes
match.group(2). Line 294 hashesmatch.group(2).strip(). The same code block therefore produces two differentcontent_hashvalues.
BaseProvider.harvest_code(lines 432-435) calls both methods for the same messages. It stores the blob under the unstripped hash and returnsCodeBlockobjects that carry the stripped hash.get_code_by_hashthen returnsNonefor every returned block, and thePointerin the taxonomy cannot be resolved from the returned hash.Normalize the code text once and reuse it in both methods.
🐛 Proposed fix for
index_conversationfor blk_idx, match in enumerate(self.CODE_BLOCK_PATTERN.finditer(content)): lang = (match.group(1) or 'text').lower() - code = match.group(2) + code = match.group(2).strip() ch = hashlib.sha256(code.encode()).hexdigest()[:16]🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/base.py` around lines 249 - 295, Normalize each extracted code block identically in index_conversation and extract_from_messages by applying the same trimming before hashing, storing blobs, and constructing CodeBlock results. Update the visible hashing flow around CODE_BLOCK_PATTERN and ensure BaseProvider.harvest_code returns hashes that resolve through get_code_by_hash and the taxonomy pointers.multi-ai-cli/providers/base.py-184-217 (1)
184-217: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winRestore pointer identity when the index is loaded.
_from_flatadds eachPointertoself.taxonomybut never writes it toself.hash_to_pointer._rebuild_hash_indexthen runs and inserts a placeholderPointer("imported", 0, 0, ch)for every hash found inblobs/.
search(line 327) andget_by_provider(line 370) readhash_to_pointerand reportsession_id,message_index, andblock_index. After any process restart, every persisted block is reported withsession_id="imported"and indices0, so citations and pointer keys are wrong.Populate
hash_to_pointerin_from_flat.🐛 Proposed fix
path = ptr_data.get('path', ['uncategorized']) self.taxonomy.add_pointer(p, path) + if p.content_hash: + self.hash_to_pointer[p.content_hash] = p ts = ptr_data.get('ts')🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/base.py` around lines 184 - 217, Update _from_flat to store each reconstructed Pointer in self.hash_to_pointer keyed by p.content_hash before or alongside adding it to self.taxonomy, so _rebuild_hash_index preserves the persisted pointer instead of creating an imported placeholder. Keep the existing pointer reconstruction and provider indexing behavior unchanged.multi-ai-cli/main.py-14-20 (1)
14-20: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winThe provider shortcut never runs from the installed console script.
multi-ai-cli/setup.pyline 25 points themulti-ai-cliconsole script atmulti_ai_cli.main:cli. That target is the imported Click group, so the--providerhandling underif __name__ == "__main__"is skipped for installed users. It runs only forpython main.py. Move the dispatch logic into amain()function and point the entry point at it.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/main.py` around lines 14 - 20, Move the provider dispatch currently guarded by __name__ == "__main__" in main.py into a callable main() function, and update setup.py’s multi-ai-cli console entry point to target main:main rather than the imported cli group. Preserve the existing direct-script behavior and ensure installed console-script invocations execute the --provider handling.multi-ai-cli/mistralai_cli.py-142-164 (1)
142-164: 📐 Maintainability & Code Quality | 🟠 Major | 🏗️ Heavy liftThe CLI provider registry duplicates the
providerspackage registry.
get_available_providershardcodes three entries here, whilemulti-ai-cli/providers/__init__.pymaintains the real registry withget_provider_typesand per-provideris_available. The two lists already disagree: this function omits Gemini, Claude, and Colab, and it addsdeepseek-tui, which is not a registered provider.tools infoandprovider selecttherefore report a provider set that does not match dispatch. Read from theproviderspackage and keep only the subprocess launch details here.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/mistralai_cli.py` around lines 142 - 164, Update get_available_providers to derive the provider set and availability from the providers package registry, including all registered providers and excluding unregistered entries such as deepseek-tui. Preserve only the CLI-specific subprocess launch metadata (such as paths) locally, while reusing get_provider_types and each provider’s is_available behavior so tools info and provider select match dispatch.multi-ai-cli/mistralai_cli.py-662-669 (1)
662-669: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winThe interactive shell re-invokes a module path that does not exist.
Lines 666 and 683 run
python3 -m multi_ai_cli.mistralai_cli. The package directory ismulti-ai-cli, which is not an importable module name, andmulti-ai-cli/setup.pyusesfind_packages(), which cannot pick up a hyphenated directory. Every interactive command therefore fails with "No module named multi_ai_cli". Invoke the Click group in-process instead of spawning a subprocess.🐛 Proposed alternative
- # Parse and execute command - import subprocess - result = subprocess.run( - ["python3", "-m", "multi_ai_cli.mistralai_cli", "--"] + cmd.split(), - capture_output=True, - text=True - ) - if result.stdout: - console.print(result.stdout) - if result.stderr: - console.print(f"[red]{result.stderr}[/red]", file=sys.stderr) + # Execute the command in-process + try: + cli.main(args=cmd.split(), standalone_mode=False) + except click.ClickException as exc: + exc.show()Also applies to: 679-686
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/mistralai_cli.py` around lines 662 - 669, Replace the subprocess-based command execution in the interactive shell branches around the command handling and exit paths with an in-process invocation of the existing Click command group, avoiding the invalid `multi_ai_cli.mistralai_cli` module path. Preserve the current command argument parsing and output/error handling while ensuring both affected execution paths invoke the Click group directly.multi-ai-cli/core/core.py-173-174 (1)
173-174: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick winDo not print session cookie material.
Line 174 writes the first 30 characters of the session cookie to stdout. Terminal scrollback, CI logs, and shell captures then contain credential material. Remove the debug print, or log only a non-reversible indicator such as the cookie length.
🔒️ Proposed fix
- if cookie: - print(f"[DEBUG] create_session using cookie: {cookie[:30]}...") + if cookie: + console.print("[dim]create_session using a provided session cookie[/]")🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/core/core.py` around lines 173 - 174, Remove the cookie value output from the create_session flow around the if cookie block; do not print any session cookie material, and if diagnostic logging is required, report only a non-reversible attribute such as its length.multi-ai-cli/core/chat_dispatcher.py-32-33 (1)
32-33: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winPrint only the response from the legacy CLI.
multi-ai-cli/cli.pycallsChatDispatcher.send(...)and passes the result directly toclick.echoat lines 21 and 33, sochatandexecuteprint a Python dict instead of text. UseLegacyChatDispatcher.send(...)or readresult.get('response', '')before echoing, including the path covered bychat --session ....🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/core/chat_dispatcher.py` around lines 32 - 33, The ChatDispatcher.send result is printed as a Python dict by the CLI instead of only its response text. Update the chat and execute output paths in cli.py, including the session-specific chat path, to use LegacyChatDispatcher.send(...) or extract result.get('response', '') before passing the value to click.echo.multi-ai-cli/main.py-24-56 (1)
24-56: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winThe
--providerfallback path passes unknown options to Click.
sys.argvis never rewritten. If the provider is unknown, unavailable, or its script file is missing, control reaches line 54 and callscli(). Click then parses the originalsys.argv, sees--provider, and exits with "no such option". The user gets a confusing error instead of a clear message. Strip the consumed arguments before the fallback, and report the failure reason.🐛 Proposed fix
providers = get_available_providers() if provider_name in providers: prov_info = providers[provider_name] if prov_info.get("available", True): import subprocess if provider_name == "deepseek": deepcli_path = os.path.expanduser("~/deepcli/deepcli.py") if os.path.exists(deepcli_path): cmd = ["python3", deepcli_path] + args result = subprocess.run(cmd) sys.exit(result.returncode) elif provider_name == "deepseek-tui": tui_path = os.path.expanduser("~/deepcli-tui/tui.py") if os.path.exists(tui_path): cmd = ["python3", tui_path] + args result = subprocess.run(cmd) sys.exit(result.returncode) - # If we get here, run the normal CLI - cli() + # If we get here, drop the consumed arguments and run the normal CLI + console.print(f"[yellow]Provider '{provider_name}' is not runnable; falling back to the CLI[/yellow]") + sys.argv = [sys.argv[0]] + args + cli()🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/main.py` around lines 24 - 56, Update the --provider handling in the main argument-dispatch flow to remove the consumed --provider and provider name arguments before calling cli(), so Click only receives the provider-specific arguments. When the provider is unknown, unavailable, or its script path is missing, report a clear failure reason before using the fallback CLI.multi-ai-cli/mistralai_cli.py-377-383 (1)
377-383: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winAdd the missing storage and index-loading methods.
CodeHarvester.save_to_storage()andSearchEngine.load_index()are undefined, soharvest code,harvest text, andsearch code --indexwill call undefined methods. Add these methods or replace the unsupported commands with an existing harvester/search API path.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/mistralai_cli.py` around lines 377 - 383, Add the missing CodeHarvester.save_to_storage() and SearchEngine.load_index() implementations used by the harvest and search command flows, or redirect those commands to existing supported harvester/search APIs. Ensure harvest code/text and search code --index no longer invoke undefined methods, while preserving their current output and index-loading behavior.Source: Linters/SAST tools
multi-ai-cli/requirements.txt-4-4 (1)
4-4: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick winRaise the
curl_cffifloor above the SSRF advisory.
curl_cffi>=0.5.0permits versions affected by GHSA-qw2m-4pqf-rmpp / PYSEC-2026-2431 (redirect-based SSRF with TLS impersonation bypass); update the floor to at least0.15.0. Since>=0.5.0also includes versions before the bundled libcurl patch, the requirement should be tightened to satisfy both advisories.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/requirements.txt` at line 4, Update the curl_cffi requirement in requirements.txt from the vulnerable 0.5.0 floor to a minimum of 0.15.0, ensuring versions affected by both cited advisories are excluded.Source: Linters/SAST tools
multi-ai-cli/tools/git_utils.py-116-148 (1)
116-148: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winFix the no-op
git addwhenall_files=False.
commit()defaults toall_files=False. In that branch,add_cmdis["git", "add"]with no pathspec. Runninggit addwithout arguments stages nothing; it does not add tracked-file modifications. Consequently, callingGitUtils.commit(message)(the default call) never stages the caller's changes and typically fails at thegit commitstep with "nothing to commit," unless something was staged elsewhere beforehand.Decide the intended default behavior and pass an explicit pathspec.
🐛 Proposed fix
- add_cmd = ["git", "add", "."] if all_files else ["git", "add"] + add_cmd = ["git", "add", "-A"] if all_files else ["git", "add", "-u"]🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tools/git_utils.py` around lines 116 - 148, Update the add command in commit() so the all_files=False path includes an explicit pathspec that stages the intended changes, while preserving the existing all_files=True behavior. Ensure the default GitUtils.commit(message) call stages caller changes before running git commit.multi-ai-cli/tools/network_utils.py-29-104 (1)
29-104: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick winStop mutating session headers per request; pass
headersper call instead.
get,post,put, anddeleteall callself.session.headers.update(headers)before issuing the request. This permanently merges one-off headers into the session's persistent header set. curl_cffi's per-requestheaders=argument merges headers for that single call only and leaves session headers unchanged; passing headers this way instead avoids the leak. As written, any credential or one-off header passed to one call (e.g., anAuthorizationheader for API A) persists and is sent on later calls to unrelated URLs that don't specify headers, which can leak secrets across domains.
clear_headers()(Line 163-165) doesn't fix this either: it only re-applies the default header dict via.update(), so custom headers set outside the defaults are never actually removed.🔒 Proposed fix
def get(self, url: str, headers: Dict = None, params: Dict = None, timeout: int = 30) -> Optional[Dict]: """Send GET request.""" try: - if headers: - self.session.headers.update(headers) - - response = self.session.get(url, params=params, timeout=timeout) + response = self.session.get(url, params=params, headers=headers, timeout=timeout)def post(self, url: str, data: Dict = None, json_data: Dict = None, headers: Dict = None, timeout: int = 30) -> Optional[Dict]: """Send POST request.""" try: - if headers: - self.session.headers.update(headers) - - response = self.session.post(url, data=data, json=json_data, timeout=timeout) + response = self.session.post(url, data=data, json=json_data, headers=headers, timeout=timeout)Apply the same change to
putanddelete, and clear the underlying multi-dict fully inclear_headers():def clear_headers(self): """Clear all custom headers.""" - self._configure_session() + self.session.headers.clear() + self._configure_session()🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tools/network_utils.py` around lines 29 - 104, Update get, post, put, and delete to pass headers directly to their respective session request calls instead of mutating self.session.headers with update. Ensure headers remain scoped to the individual request and do not persist across later calls. Also update clear_headers to fully clear the underlying header multidict before reapplying the default headers.multi-ai-cli/tools/git_utils.py-238-261 (1)
238-261: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick winGuard against git option injection in
clone().
clone(repository, directory, branch)passes these values intogit cloneas bare positional arguments, which git can parse as options when-prefixes them. Use--before each untrusted positional argument, such asgit clone --branch ... -- <repository> -- <directory>. Apply the same treatment tocheckout(),push(), andpull()wherebranch/remoteare bare positional arguments.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tools/git_utils.py` around lines 238 - 261, Update clone() to prevent option injection by separating git options from untrusted repository and directory arguments with the appropriate -- delimiter, while preserving branch handling. Apply the same argument-boundary protection to checkout(), push(), and pull() for their untrusted branch and remote positional arguments.
🟡 Minor comments (14)
multi-ai-cli/README.md-246-269 (1)
246-269: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick winDocument the required Termux smoke tests.
The PR objective requires Termux smoke tests before merge. The development section only documents
pytest tests/. Add the required Termux commands and expected results for Vibe dispatch, explicit store paths, missing-store handling, and unavailable scaffold providers.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/README.md` around lines 246 - 269, Update the Development testing documentation near the “Running tests” section to include the required Termux smoke-test commands and expected outcomes for Vibe dispatch, explicit store paths, missing-store handling, and unavailable scaffold providers, while retaining the existing pytest command.multi-ai-cli/tests/test_core.py-86-116 (1)
86-116: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick winRestore process environment state after each test.
Line 86 replaces
HOMEwithout restoring it. Lines 110-116 deleteMISTRALAI_TOKEN, including a token that existed before the test. Line 104 also assumes that no token is set. Isolate these environment variables so test order and developer configuration cannot change test results.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tests/test_core.py` around lines 86 - 116, Update the affected tests in the cache test and TestMistralCore methods to isolate HOME and MISTRALAI_TOKEN using pytest’s environment-management fixture or equivalent save-and-restore handling. Preserve each variable’s pre-test value, restore it afterward, and explicitly remove MISTRALAI_TOKEN within the initialization test so its SystemExit assertion is independent of the developer’s environment.multi-ai-cli/tests/test_tools.py-176-182 (1)
176-182: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick winAssert that
clear_headers()removes custom headers.The test only checks that
User-Agentexists after the call. Add an assertion thatX-Test-Headeris absent. Otherwise a regression that retains custom headers passes this test.Proposed fix
net.clear_headers() # After clear, default headers should be restored assert "User-Agent" in net.session.headers + assert "X-Test-Header" not in net.session.headers🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tests/test_tools.py` around lines 176 - 182, Update test_clear_headers to assert that the custom X-Test-Header is absent from net.session.headers after clear_headers(), while preserving the existing assertion that default User-Agent headers are restored.multi-ai-cli/README.md-35-41 (1)
35-41: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick winRemove the feature-branch checkout from source installation.
Line 39 forces users onto
vibe/mistralai-vibe-code-wrapper-6055d2. This can move an installed checkout away from the intended release or current PR revision. Keep source installation branch-agnostic.Proposed fix
cd multi-ai-cli -git checkout vibe/mistralai-vibe-code-wrapper-6055d2 pip install -e .🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/README.md` around lines 35 - 41, Remove the git checkout command from the “Install from source” instructions in the README, leaving the install flow branch-agnostic with only the directory change and editable pip install steps.multi-ai-cli/tests/test_tools.py-147-158 (1)
147-158: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winMake the initial branch deterministic.
A bare
git initcan create a repository withtrunkif Git’sinit.defaultBranchis configured. Initialize the test repository with an explicit branch name, then assert that exact name.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tests/test_tools.py` around lines 147 - 158, Update test_get_current_branch to initialize the temporary repository with an explicit branch name, such as main, using the git init invocation, and replace the flexible branch assertion with an assertion for that exact name.multi-ai-cli/tests/test_core.py-24-35 (1)
24-35: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winMake
test_load_configload the temporary config.Line 24 creates
config_path, but Line 34 callsload_config()without assigning that path tocore.core.CONFIG_FILE. The test reads the process configuration instead of the fixture. It can pass without validating configuration loading.Proposed fix
+ import core.core as core_module + original_config_file = core_module.CONFIG_FILE + core_module.CONFIG_FILE = Path(config_path) result = load_config() - assert isinstance(result, dict) + assert result == config + core_module.CONFIG_FILE = original_config_file🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tests/test_core.py` around lines 24 - 35, Update test_load_config to assign the temporary config_path to core.core.CONFIG_FILE before calling load_config(), ensuring the test loads the fixture rather than the process configuration; preserve the existing cleanup and assertions.multi-ai-cli/harvesters/analyzer.py-78-78 (1)
78-78: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
_analyze_pythonand_calculate_python_complexitymissasync deffunctions and methods.Line 78 checks only
isinstance(node, ast.FunctionDef), soasync deffunctions are excluded fromresult.functions. Line 97 has the same gap for class methods, so async methods are missing fromclass_info["methods"]. Line 271 in_calculate_python_complexityuses the same check, so async functions are also excluded from thefunctionsandcyclomaticmetrics.Extend each check to
isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef))(andisinstance(item, (ast.FunctionDef, ast.AsyncFunctionDef))at line 97) to cover modern async Python code.🐛 Proposed fix
for node in ast.walk(tree): - if isinstance(node, ast.FunctionDef): + if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)): func_info = { "name": node.name, "line": node.lineno, "args": [arg.arg for arg in node.args.args], "docstring": ast.get_docstring(node) or "", } result.functions.append(func_info)# Find methods in class for item in node.body: - if isinstance(item, ast.FunctionDef): + if isinstance(item, (ast.FunctionDef, ast.AsyncFunctionDef)): class_info["methods"].append({ "name": item.name, "line": item.lineno, })for node in ast.walk(tree): - if isinstance(node, ast.FunctionDef): + if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)): metrics["functions"] += 1Also applies to: 97-101, 271-271
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/harvesters/analyzer.py` at line 78, Update the function-type checks in _analyze_python, its class-method collection, and _calculate_python_complexity to recognize both ast.FunctionDef and ast.AsyncFunctionDef. Preserve the existing handling and metrics while ensuring async functions populate result.functions, async methods populate class_info["methods"], and async functions contribute to functions and cyclomatic.multi-ai-cli/providers/claude.py-56-59 (1)
56-59: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
get_historyreturns an empty list, which disables Claude code harvesting.
BaseProvider.harvest_code(multi-ai-cli/providers/base.pylines 428-429) callsget_historywhen the caller passes no messages. For Claude, that returns[], soindex_conversationindexes nothing and reports success. The PR lists Claude as a live candidate, so this gap is silent.Raise
NotImplementedErroruntil the backend supports history retrieval. That makes the gap visible instead of returning a valid-looking empty result.Do you want me to open an issue to track the Claude history implementation?
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/claude.py` around lines 56 - 59, Update ClaudeProvider.get_history to raise NotImplementedError instead of returning an empty list, making unsupported history retrieval explicit while preserving its method signature and documenting the unavailable backend behavior through the exception.multi-ai-cli/providers/ai_studio.py-25-35 (1)
25-35: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
get_historyreturns[]while the other methods raise.
send_messageandcreate_sessionraiseRuntimeErrorand state that the provider is a scaffold.get_historyreturns an empty list instead, soBaseProvider.harvest_coderuns to completion and indexes nothing. The scaffold state stays hidden on that path.The same inconsistency exists in the other scaffold providers.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/ai_studio.py` around lines 25 - 35, Update get_history in the ai_studio scaffold provider to raise the same scaffold-provider RuntimeError pattern as send_message and create_session instead of returning an empty list. Apply the same consistency fix to get_history in the other scaffold providers, preserving is_available behavior and using each provider’s existing scaffold messaging conventions.multi-ai-cli/providers/__init__.py-139-146 (1)
139-146: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick winBare
exceptin the availability paths catchesKeyboardInterruptandSystemExit. Each availability check usesexcept:with no exception class.BaseExceptionsubclasses such asKeyboardInterrupt,SystemExit, andMemoryErrorare caught and converted into aFalseresult. A user who presses Ctrl+C during a provider scan sees the scan continue instead of stopping. Ruff reports E722 at every site. Replaceexcept:withexcept Exception:at each location.
multi-ai-cli/providers/__init__.py#L139-L146: narrow theexcept:at line 144 that wrapsprovider_class()andprovider.is_available().multi-ai-cli/providers/mistral.py#L60-L65: narrow theexcept:at line 64; the guarded expression is a plain attribute read, soexcept AttributeErroris sufficient here.multi-ai-cli/providers/claude.py#L61-L70: narrow theexcept:at line 69 that wraps the backend imports andbackend.is_available().multi-ai-cli/providers/colab.py#L45-L54: narrow theexcept:at line 53 that wraps the backend imports andbackend.is_available().🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/__init__.py` around lines 139 - 146, Replace the bare availability-check handlers with appropriately narrow exceptions: in multi-ai-cli/providers/__init__.py lines 139-146, use Exception around provider construction and is_available(); in multi-ai-cli/providers/mistral.py lines 60-65, use AttributeError for the guarded attribute read; and in multi-ai-cli/providers/claude.py lines 61-70 and multi-ai-cli/providers/colab.py lines 45-54, use Exception around backend imports and backend.is_available().Source: Linters/SAST tools
multi-ai-cli/providers/__init__.py-1-5 (1)
1-5: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick winMove the module docstring above the import.
Line 2 contains an import. The string at line 3 therefore is a plain expression statement, not the module docstring.
providers.__doc__isNone, and documentation tools show no description for this package.🐛 Proposed fix
#!/usr/bin/env python3 -from typing import Dict, List, Any """Providers module - Unified interface for all AI providers.""" + +from typing import Dict, List🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/__init__.py` around lines 1 - 5, Move the module docstring in the providers module above the typing and local imports, keeping the shebang first so Python recognizes the string as providers.__doc__.multi-ai-cli/providers/ai_studio.py-19-23 (1)
19-23: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winReplace the invalid AI Studio base URL.
https://ai.studiois not Google AI Studio’s website (https://aistudio.google.com) and is not the documented programmatic API base endpoint for Gemini AI Studio (https://generativelanguage.googleapis.com). Update this stub before wiring the provider.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/providers/ai_studio.py` around lines 19 - 23, Update the ProviderConfig returned by the AI Studio provider configuration to use the documented Gemini AI Studio API base endpoint, https://generativelanguage.googleapis.com, instead of the invalid https://ai.studio URL; leave the provider name and model unchanged.archwiz/dispatch_pipeline.py-77-79 (1)
77-79: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick winSet an explicit encoding when reading the store.
open(store)uses the platform default encoding. Session JSON is written as UTF-8 by the producers. On a non-UTF-8 locale, a session with non-ASCII content fails to load and the dispatch silently stops at line 86.🐛 Proposed fix
- with open(store) as f: + with open(store, encoding="utf-8") as f: msgs = json.load(f)🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@archwiz/dispatch_pipeline.py` around lines 77 - 79, Update the store-reading open call in the dispatch pipeline’s JSON loading block to explicitly use UTF-8, matching the encoding used by session producers while preserving the existing json.load flow and error handling.multi-ai-cli/tools/termux_utils.py-230-247 (1)
230-247: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winParse
pkg list-installedoutput asname/repo, not as plain package names.The installed-package lines use
package-name/repository-info version architectureformat, soline.split()[0]returns tokens likebash/stable,nowinstead of parsed package names if callers need exact package names.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@multi-ai-cli/tools/termux_utils.py` around lines 230 - 247, Update get_installed_packages to parse each installed-package line’s first token as name/repository metadata and extract only the package name before the slash. Preserve the existing header skipping, blank-line filtering, return behavior, and exception handling.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 5d55d1bc-8324-464f-ab48-0f51ced66589
📒 Files selected for processing (39)
archwiz/dispatch_pipeline.pymulti-ai-cli/README.mdmulti-ai-cli/cli.pymulti-ai-cli/core/__init__.pymulti-ai-cli/core/chat_dispatcher.pymulti-ai-cli/core/core.pymulti-ai-cli/harvesters/__init__.pymulti-ai-cli/harvesters/analyzer.pymulti-ai-cli/harvesters/code_harvester.pymulti-ai-cli/harvesters/extractor.pymulti-ai-cli/harvesters/search_engine.pymulti-ai-cli/main.pymulti-ai-cli/mistralai_cli.pymulti-ai-cli/providers/__init__.pymulti-ai-cli/providers/ai_studio.pymulti-ai-cli/providers/anthropic.pymulti-ai-cli/providers/base.pymulti-ai-cli/providers/claude.pymulti-ai-cli/providers/colab.pymulti-ai-cli/providers/deepseek.pymulti-ai-cli/providers/duckduckgo.pymulti-ai-cli/providers/gemini.pymulti-ai-cli/providers/grok.pymulti-ai-cli/providers/kimi.pymulti-ai-cli/providers/mistral.pymulti-ai-cli/providers/openai.pymulti-ai-cli/providers/perplexity.pymulti-ai-cli/providers/qwen.pymulti-ai-cli/requirements.txtmulti-ai-cli/setup.pymulti-ai-cli/tests/__init__.pymulti-ai-cli/tests/test_core.pymulti-ai-cli/tests/test_harvesters.pymulti-ai-cli/tests/test_tools.pymulti-ai-cli/tools/__init__.pymulti-ai-cli/tools/file_utils.pymulti-ai-cli/tools/git_utils.pymulti-ai-cli/tools/network_utils.pymulti-ai-cli/tools/termux_utils.py
| from pathlib import Path | ||
| from typing import Optional, List, Dict, Any | ||
| from curl_cffi import requests as curl_requests | ||
| import requests as http_requests |
There was a problem hiding this comment.
🩺 Stability & Availability | 🔴 Critical | ⚡ Quick win
requests is imported but is not a declared dependency.
Line 13 imports requests unconditionally at module import time. In multi-ai-cli/requirements.txt line 11, requests is commented out. On a clean install, importing core.core raises ImportError and the whole CLI fails. The alias http_requests is not used anywhere in this file. Remove the import, or declare requests in requirements.txt.
🐛 Proposed fix
from curl_cffi import requests as curl_requests
-import requests as http_requests
from rich.console import Console📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| import requests as http_requests | |
| import requests as curl_requests | |
| from rich.console import Console |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@multi-ai-cli/core/core.py` at line 13, Remove the unused http_requests import
from core.py so importing core.core does not require the undeclared requests
dependency; do not modify dependency declarations since this symbol is not used.
| def search(self, term: str, language: str = None) -> List[Dict]: | ||
| """Search code blocks by term and optional language filter. | ||
|
|
||
| Args: | ||
| term: Search term | ||
| language: Optional language filter | ||
|
|
||
| Returns: | ||
| List of search results with code content | ||
| """ | ||
| results = [] | ||
|
|
||
| # Search all blobs for the term | ||
| for ch, blob_path in self.blobs.items(): | ||
| if Path(blob_path).exists(): | ||
| code = Path(blob_path).read_text() | ||
| if term.lower() in code.lower(): | ||
| # Get pointer | ||
| p = self.hash_to_pointer.get(ch) | ||
| if p: | ||
| # Filter by language if specified | ||
| if language: | ||
| # Get the language from taxonomy | ||
| # For now, check if the code block is in the specified language | ||
| # by checking the first part of the taxonomy path | ||
| lang_node = self.taxonomy.children.get(language.lower()) | ||
| if lang_node: | ||
| # Check if this hash is in the language node | ||
| if not self._is_hash_in_node(lang_node, ch): | ||
| continue | ||
|
|
||
| results.append({ | ||
| 'pointer': p.to_key(), | ||
| 'hash': ch, | ||
| 'code': code[:200] + '...' if len(code) > 200 else code, | ||
| 'timestamp': self.time_index.get(ch, '').isoformat(), | ||
| 'session_id': p.session_id, | ||
| 'message_index': p.message_index, | ||
| 'block_index': p.block_index, | ||
| }) | ||
|
|
||
| return results |
There was a problem hiding this comment.
🩺 Stability & Availability | 🔴 Critical | ⚡ Quick win
search() and search_by_language() crash when a blob has no time_index entry.
At line 429, self.time_index.get(ch, '').isoformat() falls back to the empty string '' when ch is absent from time_index, then immediately calls .isoformat() on it. The same pattern repeats at line 481 in search_by_language(). _rebuild_hash_index() (lines 221-231) adds entries to self.blobs and self.hash_to_pointer for every .blob file found on disk, but never adds a corresponding time_index entry. Any blob recovered this way (for example after codex_index.json is missing or stale) triggers AttributeError: 'str' object has no attribute 'isoformat' the first time it matches a search term.
_save() (line 244) already uses the safe pattern — self.time_index.get(p.content_hash, datetime.now(timezone.utc)).isoformat() — that same fallback approach should be applied here. This bug also propagates to SearchEngine.search() and SearchEngine.search_by_language() in multi-ai-cli/harvesters/search_engine.py (lines 63-95, 97-128), since they call self.codex.search(...) / self.codex.search_by_language(...) directly.
🐛 Proposed fix
- results.append({
- 'pointer': p.to_key(),
- 'hash': ch,
- 'code': code[:200] + '...' if len(code) > 200 else code,
- 'timestamp': self.time_index.get(ch, '').isoformat(),
+ ts = self.time_index.get(ch)
+ results.append({
+ 'pointer': p.to_key(),
+ 'hash': ch,
+ 'code': code[:200] + '...' if len(code) > 200 else code,
+ 'timestamp': ts.isoformat() if ts else '',Apply the equivalent change at line 481 in search_by_language().
Also applies to: 447-487
🧰 Tools
🪛 Pylint (4.0.6)
[refactor] 407-433: Too many nested blocks (7/5)
(R1702)
🪛 Ruff (0.16.0)
[warning] 394-394: PEP 484 prohibits implicit Optional
Convert to T | None
(RUF013)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@multi-ai-cli/harvesters/code_harvester.py` around lines 394 - 435, Update the
timestamp construction in CodeHarvester.search() and search_by_language() to use
a datetime fallback, such as the existing timezone-aware current-time pattern
used by _save(), before calling isoformat(). Apply the same safe handling to the
corresponding SearchEngine.search() and search_by_language() result paths if
they construct timestamps, while preserving existing results for blobs with
time_index entries.
There was a problem hiding this comment.
@coderabbitai please apply all your suggested and proposed fixes.
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
ast-grep outline multi-ai-cli/harvesters/code_harvester.py --items all --type function --match 'search|search_by_language|_rebuild_hash_index'
ast-grep outline multi-ai-cli/harvesters/search_engine.py --items all --type function --match 'search|search_by_language'
ast-grep outline multi-ai-cli/tests/test_harvesters.py --items all --type function --match 'search|language'
printf '\n--- Codex search implementation ---\n'
sed -n '200,510p' multi-ai-cli/harvesters/code_harvester.py
printf '\n--- SearchEngine implementation ---\n'
sed -n '1,155p' multi-ai-cli/harvesters/search_engine.py
printf '\n--- Relevant tests ---\n'
sed -n '1,220p' multi-ai-cli/tests/test_harvesters.pyLength of output: 26373
🏁 Script executed:
#!/bin/bash
set -euo pipefail
python -m py_compile \
multi-ai-cli/harvesters/code_harvester.py \
multi-ai-cli/harvesters/search_engine.py \
multi-ai-cli/tests/test_harvesters.py
rg -n -C2 "time_index\\.get\\([^)]*,\\s*['\\\"]['\\\"]\\)\\.isoformat\\(\\)" \
multi-ai-cli/harvesters/code_harvester.py \
multi-ai-cli/harvesters/search_engine.py \
|| trueLength of output: 1601
@timerloggedout-spec, validation confirms that the issue is still present at Line 429 and Line 481 in multi-ai-cli/harvesters/code_harvester.py.
The source files compile successfully. The planned change replaces both unsafe fallbacks and adds regression coverage for recovered blobs with no timestamp.
You are interacting with an AI system.
| class CodeHarvester: | ||
| """Harvests code from files, directories, session responses, and other sources. | ||
|
|
||
| This is the main interface for code harvesting, which uses CodexIndex | ||
| internally for session-based code extraction. | ||
| """ | ||
|
|
||
| # Language extensions mapping | ||
| LANGUAGE_EXTENSIONS = { | ||
| '.py': 'python', | ||
| '.js': 'javascript', | ||
| '.ts': 'typescript', | ||
| '.java': 'java', | ||
| '.c': 'c', | ||
| '.cpp': 'cpp', | ||
| '.h': 'c', | ||
| '.hpp': 'cpp', | ||
| '.go': 'go', | ||
| '.rs': 'rust', | ||
| '.rb': 'ruby', | ||
| '.php': 'php', | ||
| '.swift': 'swift', | ||
| '.kt': 'kotlin', | ||
| '.scala': 'scala', | ||
| '.sh': 'bash', | ||
| '.bash': 'bash', | ||
| '.zsh': 'bash', | ||
| '.sql': 'sql', | ||
| '.html': 'html', | ||
| '.css': 'css', | ||
| '.json': 'json', | ||
| '.yaml': 'yaml', | ||
| '.yml': 'yaml', | ||
| '.xml': 'xml', | ||
| '.md': 'markdown', | ||
| } | ||
|
|
||
| def __init__(self, codex_dir: Path = None): | ||
| """Initialize the harvester.""" | ||
| self.codex = CodexIndex(codex_dir) | ||
| self.snippets: List[CodeBlock] = [] | ||
|
|
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🔴 Critical | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Description: Confirm save_to_storage is not defined elsewhere in the repo.
rg -n 'def save_to_storage' --type=pyRepository: timerloggedout-spec/termux-monorepo
Length of output: 173
🏁 Script executed:
#!/bin/bash
set -euo pipefail
echo "== Locate relevant files =="
git ls-files | rg '(^|/)multi-ai-cli/(mistralai_cli\.py|harvesters/code_harvester\.py)$' || true
echo
echo "== save_to_storage references =="
rg -n 'save_to_storage|CodeHarvester\(' --type=py . || true
echo
echo "== CodeHarvester class outline =="
ast-grep outline multi-ai-cli/harvesters/code_harvester.py --view expanded | sed -n '1,220p' || true
echo
echo "== Relevant CodeHarvester section =="
sed -n '542,833p' multi-ai-cli/harvesters/code_harvester.py
echo
echo "== Harness/CLI relevant section =="
sed -n '320,400p' multi-ai-cli/mistralai_cli.py
echo
echo "== python runtime static/import probe (no repo execution) =="
python3 - <<'PY'
import ast
from pathlib import Path
paths = [
Path('multi-ai-cli/harvesters/code_harvester.py'),
Path('multi-ai-cli/mistralai_cli.py'),
]
for path in paths:
tree = ast.parse(path.read_text(), filename=str(path))
print(f"\n--- {path} ---")
for node in ast.walk(tree):
if isinstance(node, ast.ClassDef):
methods = [n.name for n in node.body if isinstance(n, ast.FunctionDef)]
print(f"class {node.name}: methods={methods}")
elif isinstance(node, ast.AnnAssign):
target = ast.unparse(node.target)
value = ast.unparse(node.value) if node.value else ''
if 'CODE_HARVESTER' in target or 'CodeHarvester' in value:
print(f"ANNASSIGN {target}={value}")
elif isinstance(node, ast.Attribute) and node.attr == 'save_to_storage':
print(f"ATTR save_to_storage at line {node.lineno}, full={ast.unparse(node)}")
PYRepository: timerloggedout-spec/termux-monorepo
Length of output: 18483
Add save_to_storage to CodeHarvester before using the CLI harvest commands.
mistralai_cli.py calls harvester.save_to_storage(...) in both code and text, but CodeHarvester only stores extracted snippets in self.snippets and does not define this method. Without it, the commands raise AttributeError after harvesting completes. Add a persistence method that saves the snippets, for example through self.codex._save() or a harvester-specific storage writer.
🧰 Tools
🪛 Ruff (0.16.0)
[warning] 550-577: Mutable default value for class attribute
(RUF012)
[warning] 579-579: PEP 484 prohibits implicit Optional
Convert to T | None
(RUF013)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@multi-ai-cli/harvesters/code_harvester.py` around lines 542 - 583, The
CodeHarvester class lacks the save_to_storage method invoked by the CLI
commands, causing an AttributeError after harvesting. Add
CodeHarvester.save_to_storage to persist the collected self.snippets through the
existing CodexIndex storage mechanism, such as self.codex._save(), while
preserving the current harvesting behavior.
| if result.stderr: | ||
| console.print(f"[red]{result.stderr}[/red]", file=sys.stderr) |
There was a problem hiding this comment.
🩺 Stability & Availability | 🔴 Critical | ⚡ Quick win
Standard print keyword arguments are passed to rich.console.Console.print. Console.print accepts sep, end, style, and rendering options only. It does not accept file or flush. Each of these calls raises TypeError on the affected path.
multi-ai-cli/mistralai_cli.py#L115-L116: removefile=sys.stderrand print through aConsole(stderr=True)instance; apply the same change at lines 127-128, 672-673, and 689-690.multi-ai-cli/core/core.py#L264-L274: removeflush=Truefrom the streamingconsole.printcall at line 271 and callconsole.file.flush()if an explicit flush is required.
🧰 Tools
🪛 Pylint (4.0.6)
[error] 116-116: Unexpected keyword argument 'file' in method call
(E1123)
📍 Affects 2 files
multi-ai-cli/mistralai_cli.py#L115-L116(this comment)multi-ai-cli/core/core.py#L264-L274
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@multi-ai-cli/mistralai_cli.py` around lines 115 - 116, Update the
Console.print calls in multi-ai-cli/mistralai_cli.py at lines 115-116, 127-128,
672-673, and 689-690 to remove file=sys.stderr and route stderr output through a
Console(stderr=True) instance; update the streaming console.print call in
multi-ai-cli/core/core.py at lines 264-274 to remove flush=True and use
console.file.flush() when explicit flushing is required.
Source: Linters/SAST tools
| packages=find_packages(), | ||
| python_requires=">=3.8", | ||
| install_requires=requirements, | ||
| entry_points={ | ||
| "console_scripts": [ | ||
| "multi-ai-cli = multi_ai_cli.main:cli", | ||
| "multi-ai = multi_ai_cli.main:cli", | ||
| ], |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🔴 Critical | 🏗️ Heavy lift
No importable multi_ai_cli package exists. The source directory is named multi-ai-cli, which is not a valid module name, and find_packages() produces only the generic top-level packages core, providers, harvesters, tools, and tests. Every reference to multi_ai_cli fails at runtime.
multi-ai-cli/setup.py#L20-L27: move the sources into amulti_ai_cli/package directory, then point both console scripts at a real module path inside it.multi-ai-cli/main.py#L14-L20: move the--providerdispatch logic out of the__main__block into amain()function and make that function the console-script target.multi-ai-cli/mistralai_cli.py#L662-L669: stop spawningpython3 -m multi_ai_cli.mistralai_cli; invoke the Click group in-process, and apply the same change at lines 679-686.
📍 Affects 3 files
multi-ai-cli/setup.py#L20-L27(this comment)multi-ai-cli/main.py#L14-L20multi-ai-cli/mistralai_cli.py#L662-L669
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@multi-ai-cli/setup.py` around lines 20 - 27, Reorganize the sources under an
importable multi_ai_cli/ package, update setup.py’s package discovery and both
console_scripts targets to real modules, and move the __main__ provider dispatch
in main.py into a callable main() entry point. In mistralai_cli.py, replace both
subprocess calls that launch python3 -m multi_ai_cli.mistralai_cli with direct
in-process invocation of the Click group. Apply these changes at
multi-ai-cli/setup.py lines 20-27, multi-ai-cli/main.py lines 14-20, and
multi-ai-cli/mistralai_cli.py lines 662-669; also update the second
corresponding call at multi-ai-cli/mistralai_cli.py lines 679-686.
Feedback for CoPilot / master-staging gateObserved: Considerations / recommendations
— Grok (continuing project; ready for Termux smoke confirmation or next TER-10 branch when you give the green light) |
|
ArchW1z disposition: 🔴 NO-GO — bug farm, not merge candidate Critical unresolved defects (Critical-Eval §6):
Action: Do not merge wholesale. Extract individual fixes onto Specs: |
#6 - Extract CodexIndex and Pointer logic into archwiz/codex.py - Define BaseProvider interface in archwiz/providers/base.py - Implement DeepSeekProvider wrapper in archwiz/providers/deepseek.py - Use archwiz.config for all paths
Automated disposition (production prioritization)Status remains: 🔴 NO-GO wholesale. Valuable pieces exist (silent dispatch logging already partially superseded by merged #5; provider-aware store resolution; scaffold Do not merge this branch as a unit. Extract focused commits onto Recommended extract targets:
Once extracts land, close this PR. — Grok production automation |
|
Note Docstrings generation - SUCCESS |
…ovider-aware-stores-mark-scaffold` Docstrings generation was requested by @timerloggedout-spec. * #6 (comment) The following files were modified: * `archwiz/dispatch_pipeline.py` * `multi-ai-cli/cli.py` * `multi-ai-cli/core/chat_dispatcher.py` * `multi-ai-cli/core/core.py` * `multi-ai-cli/harvesters/analyzer.py` * `multi-ai-cli/harvesters/code_harvester.py` * `multi-ai-cli/harvesters/extractor.py` * `multi-ai-cli/harvesters/search_engine.py` * `multi-ai-cli/mistralai_cli.py` * `multi-ai-cli/providers/__init__.py` * `multi-ai-cli/providers/ai_studio.py` * `multi-ai-cli/providers/anthropic.py` * `multi-ai-cli/providers/base.py` * `multi-ai-cli/providers/claude.py` * `multi-ai-cli/providers/colab.py` * `multi-ai-cli/providers/deepseek.py` * `multi-ai-cli/providers/duckduckgo.py` * `multi-ai-cli/providers/gemini.py` * `multi-ai-cli/providers/grok.py` * `multi-ai-cli/providers/kimi.py` * `multi-ai-cli/providers/mistral.py` * `multi-ai-cli/providers/openai.py` * `multi-ai-cli/providers/perplexity.py` * `multi-ai-cli/providers/qwen.py` * `multi-ai-cli/tools/file_utils.py` * `multi-ai-cli/tools/git_utils.py` * `multi-ai-cli/tools/network_utils.py` * `multi-ai-cli/tools/termux_utils.py`
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
sha: 8483f33 @jules opsSweep (heyVern lane) — high-perf unattended advance. PR #6 · Instructions
Monikers: docs/ops/AGENT-MONIKERS.md · Read AGENTS.md. |
|
Closing as superseded rather than abandoned: current |
Summary
TER-9 work on the vibe/mistral wrapper branch: surface silent dispatch failures, provider-aware session resolution, and mark scaffold providers non-live.
Status: 🔴 NO-GO wholesale
Disposition: Valuable patches exist; do not merge the branch as a unit. Extract focused commits onto
master-stagingafter review. CodeRabbit “feature dump” summary overstates product scope.Base:
master-stagingImplements: CE-15 (extract only) / TER-9
Intended changes (keep when extracting)
[archwiz dispatch]from vibecore.py_cache_saveupdate_all(..., provider=, account=, store_path=)with ordered store searchis_available() -> Falsefor non-live backends; live candidates called out explicitlyNon-goals
~/.archwiz/sessions/…migrationValidation (Termux — required before any extract merge)
update_allmust find file[archwiz dispatch]get_available_providers()reports scaffolds as unavailableFollow-ups
master-stagingAgent notes
grok-archw1zdevinorchatgptif available (anti-monopoly roster)