Repository navigation
feat(almas): implement model-based refinement, memory digest, and expand unit tests - #12
Conversation
This lands the UIX feature set from the maki fork into noon main: - AgentLoop caches tool definitions and hashes with incremental updates - task plugin gains auto_tier and structured_output validation - new ALMAS plugin with role/model routing - shared noon.route_tier library for prompt-based tier selection - noon.json.to_ton / from_ton for token-efficient tool output - fix noon.agent.usage_cost for current providers model layout - restore maki alias for backward-compatible Lua plugins Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
…docs - Rename noon.json.to_ton/from_ton to to_toon/from_toon for consistent spelling. - Remove unused ctx parameter from noon.agent.usage_cost. - Pass resolved model spec to noon.agent.tools in ALMAS roles. - Regenerate site docs to reflect usage_cost, to_toon/from_toon, almas, and task auto_tier. Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
…and unit tests Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
|
Warning Review limit reached
Next review available in: 47 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (5)
📝 WalkthroughWalkthroughThe changes add agent tool-cache invalidation, Lua TOON and usage-cost APIs, automatic model-tier routing, and a bundled ChangesAgent tool cache invalidation
Lua APIs and model-tier routing
ALMAS bundled workflow
Estimated code review effort: 5 (Critical) | ~120 minutes Sequence Diagram(s)sequenceDiagram
participant Caller
participant AlmasTool
participant Supervisor
participant RoleAgent
participant Memory
Caller->>AlmasTool: submit goal and options
AlmasTool->>Memory: load prior goal digest
AlmasTool->>Supervisor: request validated role plan
Supervisor-->>AlmasTool: ordered role steps
AlmasTool->>RoleAgent: execute each step with optional retrieved context
RoleAgent-->>AlmasTool: result and estimated cost
AlmasTool->>Memory: save learnings digest
AlmasTool-->>Caller: markdown report and completion summary
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…tests - Add missing plugins/almas/mem.lua for persistent learnings. - Fix retrieve.lua hash_token to use bit32.bxor/band for Luau. - Remove unused max_concurrent option in almas/init.lua. - Enable almas in DEFAULT_BUILTINS and noon-docgen SECTIONS. - Regenerate site docs to include almas. - Add noon.agent.usage_cost to task plugin spec expectations. - Re-apply messages text-extraction test fixes for new role prefixes. - Re-apply task structured-output nudge retry fixes. - Add usage_cost acceptance test without context. Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
There was a problem hiding this comment.
Code Review
This pull request introduces the ALMAS (Autonomous LLM-based Multi-Agent Software Engineering) plugin, which decomposes software development goals into specialized role agents running on cost-aware model tiers. To support this, several core enhancements were added, including a cost-aware model-tier router (route_tier), TOON (Token-Oriented Object Notation) encoding/decoding for token-efficient context blocks, and a usage_cost API for estimating completion costs. Additionally, tool definition caching in the UI was optimized using a cheap variable content hash and registry snapshot pinning. Reviewers identified two key areas for improvement: refactoring the retrieval logic in plugins/almas/retrieve.lua to prevent duplicate grep calls when falling back from lexical to vector search, and addressing a potential regression in the task tool caused by the removal of the retry nudge loop for structured outputs.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
I am having trouble creating individual review comments. Click here to see my feedback.
plugins/almas/retrieve.lua (122-128)
The current implementation of M.retrieve is inefficient. It calls retrieve_lexical, and if that fails, it calls retrieve_vector. Both functions perform grep calls for keywords from the goal. If retrieve_lexical finds no results, retrieve_vector will be called and will repeat the same failing grep calls.
I suggest refactoring to avoid this duplication by combining the logic into a single function. This new implementation of M.retrieve retrieves all snippets and ranks them, making retrieve_lexical and retrieve_vector obsolete. I've also added a check to avoid processing duplicate grep results.
function M.retrieve(ctx, goal, role, k)
k = k or MAX_SNIPPETS
local kws = keywords(goal)
if #kws == 0 then
return nil
end
local goal_vec = embed(goal)
local scored = {}
local seen_outputs = {}
for _, kw in ipairs(kws) do
local ok, out = pcall(function()
return noon.agent.call_tool(ctx, "grep", { pattern = kw, head_limit = 3 })
end)
if ok and out and #out > 0 and not seen_outputs[out] then
scored[#scored + 1] = {
sim = cosine(goal_vec, embed(out)),
text = "## " .. kw .. "\n" .. out:sub(1, 400),
}
seen_outputs[out] = true
end
end
if #scored == 0 then
return nil
end
table.sort(scored, function(a, b)
return a.sim > b.sim
end)
local picked = {}
for i = 1, math.min(k, #scored) do
picked[#picked + 1] = scored[i].text
end
return table.concat(picked, "\n\n"):sub(1, MAX_BLOCK_BYTES)
end
plugins/task/init.lua (237)
The removal of the retry loop that nudged the model to call structured_output is a significant behavioral change for the task tool. Previously, it would retry up to MAX_STRUCTURED_RETRIES times. Now, it will fail immediately if the subagent doesn't produce the structured output on the first attempt.
While the new almas plugin implements its own single-nudge logic, this change makes the general-purpose task tool less robust and could be a regression for other use cases that relied on this behavior. Was this simplification intentional? If so, this change in behavior should probably be documented for users of the task tool.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e9f13c8d4b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| local supervisor_tier = input.model_tier or "strong" | ||
| local goal = refine.refine_goal(ctx, input.goal, supervisor_tier) | ||
|
|
||
| local slug = memory.slug(input.goal) |
There was a problem hiding this comment.
Define ALMAS memory access before indexing it
Every invocation of the new almas tool reaches this before planning, but no bundled module defines a global memory table; the existing memory plugin only registers the memory tool and helpers. In a normal plugin environment this raises an attempt to index a nil value, so ALMAS never runs. Use the memory tool or explicitly require/export helpers before reading prior learnings.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in c6b4a89 by explicitly requiring the bundled mem module before use.
|
|
||
| local results = {} | ||
| local total_cost = 0.0 | ||
| for i, step in ipairs(steps) do |
There was a problem hiding this comment.
Stop executing default supervised plans autonomously
When input.mode is omitted or set to "supervised" (the schema advertises the default as reviewing the plan), this loop still launches every role agent immediately. That spends model/tool calls and can let developer steps use general_sub write-capable tools before any plan review; branch on mode before entering this execution loop.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in c6b4a89: supervised mode now returns the plan for review and only autonomous mode executes role agents.
| if not model_tier and input.auto_tier then | ||
| model_tier = route_tier(input.prompt) |
There was a problem hiding this comment.
Apply the configured auto_tier option
The plugin registers opts.auto_tier, but the handler only checks the per-call input.auto_tier. When a user enables the plugin option in config and invokes task without an explicit auto_tier field, model_tier remains nil and the subagent inherits the parent model instead of using route_tier; include opts.auto_tier in this condition.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in c6b4a89: configured auto_tier is applied when the call does not explicitly override it.
Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
There was a problem hiding this comment.
Actionable comments posted: 4
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@plugins/almas/init.lua`:
- Around line 6-10: Add the missing local require for the plugin’s mem.lua
module alongside the existing imports in plugins/almas/init.lua, binding it to
the memory symbol used by handler. Ensure handler can call memory.slug,
memory.load, and memory.save without relying on a global.
- Around line 6-10: Remove the unused ToolView and route_tier imports from the
module initialization imports; retain refine, retrieve, and roles because they
are still used.
In `@plugins/almas/roles.lua`:
- Line 41: Update the tier selection expression in the roles configuration to
prioritize opts.auto_tier and route_tier(prompt) before opts.model_tier, while
retaining r.tier as the final fallback. This ensures automatic routing overrides
the planner-provided tier and matches the task behavior.
In `@plugins/almas/tests/spec.lua`:
- Around line 71-97: The retrieve_vector_fallback test currently allows the
second lexical keyword search to return a match, so retrieve.retrieve can
succeed without invoking retrieve_vector. Update the noon.agent.call_tool mock
in retrieve_vector_fallback to return an empty result for every lexical grep
call, while returning the vector fallback match for the appropriate vector
query, ensuring the assertions exercise the actual fallback path.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 6560dfa5-d4eb-4005-81fc-ac2ae65c14e3
⛔ Files ignored due to path filters (1)
Cargo.lockis excluded by!**/*.lock
📒 Files selected for processing (26)
noon-agent/src/template.rsnoon-agent/src/tools/registry.rsnoon-config/src/lib.rsnoon-docgen/src/gen_tools.rsnoon-lua/Cargo.tomlnoon-lua/benches/perf.rsnoon-lua/src/api/agent.rsnoon-lua/src/api/json.rsnoon-lua/src/loader.rsnoon-lua/src/runtime.rsnoon-lua/tests/spec.rsnoon-lua/tests/task_policy.rsnoon-ui/src/agent/agent_loop.rsplugins/almas/init.luaplugins/almas/mem.luaplugins/almas/refine.luaplugins/almas/retrieve.luaplugins/almas/roles.luaplugins/almas/tests/spec.luaplugins/lib/noon/route_tier.luaplugins/lib/tests/spec.luaplugins/task/init.luaplugins/task/tests/spec.luasite/docs/content/configuration/_index.mdsite/docs/content/lua-api/_index.mdsite/docs/content/tools/_index.md
📜 Review details
⏰ Context from checks skipped due to timeout. (3)
- GitHub Check: Test
- GitHub Check: Build (Windows)
- GitHub Check: Lint (Windows)
🧰 Additional context used
📓 Path-based instructions (6)
site/**/*.md
📄 CodeRabbit inference engine (AGENTS.md)
Keep user documentation warm, simple, concise, easy for non-native English speakers, with no em dashes, emojis, or AI-like tone.
Files:
site/docs/content/configuration/_index.mdsite/docs/content/tools/_index.mdsite/docs/content/lua-api/_index.md
plugins/**/*.lua
📄 CodeRabbit inference engine (AGENTS.md)
Keep built-in Lua plugins in
./pluginscompatible with their documented plugin responsibilities and APIs.
Files:
plugins/almas/refine.luaplugins/task/tests/spec.luaplugins/lib/noon/route_tier.luaplugins/almas/roles.luaplugins/lib/tests/spec.luaplugins/almas/mem.luaplugins/almas/tests/spec.luaplugins/almas/retrieve.luaplugins/almas/init.luaplugins/task/init.lua
**/*.rs
📄 CodeRabbit inference engine (AGENTS.md)
**/*.rs: Avoid trivial comments in Rust source files.
Keep Rust code minimal and avoid unnecessary bloat; follow KISS, DRY, and SRP.
Do not introduce unnecessary state, variables, fields, or arguments; each line should justify its existence.
Follow Rust idioms and best practices, using current stable Rust features where appropriate.
Use descriptive variable and function names.
Do not use wildcard imports.
Import types at the top of the file and use short names instead of fully qualified inline paths.
Keep constants at the top of the file, immediately after imports.
Use explicit error handling withResult<T, E>instead of panics.
Usecolor_eyrewhen the specific error type is not important.
Use customthiserrorerror types for domain-specific errors.
Place unit tests in the same Rust file inside#[cfg(test)]modules.
Be mindful of allocations in hot paths.
Prefer structured logging with useful fields.
Provide helpful error messages.
Do not add tautological or otherwise pointless tests.
Ensure tests are deterministic and do not rely on arbitrary sleeps or other flaky timing behavior.
Do not use inline magic numbers or strings; in tests, define shared constant error/status messages and assert against those constants.
Add#[derive(Copy)]only to structs containing one primitive field.
Runcargo clippy --all --tests -- -D warningsandcargo nextest run --workspaceto validate changes.
Files:
noon-lua/tests/spec.rsnoon-lua/src/loader.rsnoon-agent/src/template.rsnoon-docgen/src/gen_tools.rsnoon-lua/benches/perf.rsnoon-config/src/lib.rsnoon-lua/tests/task_policy.rsnoon-lua/src/runtime.rsnoon-lua/src/api/agent.rsnoon-lua/src/api/json.rsnoon-ui/src/agent/agent_loop.rsnoon-agent/src/tools/registry.rs
**/*.{rs,toml}
📄 CodeRabbit inference engine (AGENTS.md)
Use
#[test_case]for parameterized tests and snake_case for test names.
Files:
noon-lua/tests/spec.rsnoon-lua/src/loader.rsnoon-lua/Cargo.tomlnoon-agent/src/template.rsnoon-docgen/src/gen_tools.rsnoon-lua/benches/perf.rsnoon-config/src/lib.rsnoon-lua/tests/task_policy.rsnoon-lua/src/runtime.rsnoon-lua/src/api/agent.rsnoon-lua/src/api/json.rsnoon-ui/src/agent/agent_loop.rsnoon-agent/src/tools/registry.rs
**/Cargo.toml
📄 CodeRabbit inference engine (AGENTS.md)
**/Cargo.toml: Declare dependencies in the workspace-rootCargo.toml, then useworkspace = truein individual packages.
Prefer existing dependencies before adding new ones, and prefer well-maintained crates from crates.io.
Files:
noon-lua/Cargo.toml
noon-docgen/src/**/*.rs
📄 CodeRabbit inference engine (AGENTS.md)
Regenerate generated documentation with
just gen-docsafter changing documentation generator code.
Files:
noon-docgen/src/gen_tools.rs
🪛 Luacheck (1.2.0)
plugins/almas/tests/spec.lua
[warning] 52-52: unused argument 'ctx'
(W212)
[warning] 77-77: unused argument 'ctx'
(W212)
[warning] 77-77: unused argument 'args'
(W212)
[warning] 107-107: unused argument 'ctx'
(W212)
[warning] 116-116: unused argument 'ctx'
(W212)
plugins/almas/retrieve.lua
[warning] 122-122: unused argument 'role'
(W212)
plugins/almas/init.lua
[warning] 6-6: unused variable 'ToolView'
(W211)
[warning] 10-10: unused variable 'route_tier'
(W211)
[warning] 139-139: variable 'res' is never accessed
(W231)
🔇 Additional comments (24)
noon-lua/Cargo.toml (1)
26-26: LGTM!Also applies to: 85-88
noon-lua/src/api/json.rs (1)
130-162: LGTM!Also applies to: 164-175
noon-lua/src/api/agent.rs (1)
30-30: LGTM!Also applies to: 146-176, 629-629, 1004-1014
noon-lua/benches/perf.rs (1)
1-64: LGTM!noon-lua/src/runtime.rs (1)
1204-1205: LGTM!site/docs/content/lua-api/_index.md (1)
71-71: LGTM!Also applies to: 814-842, 2078-2081, 2167-2212, 4801-4813
plugins/lib/noon/route_tier.lua (1)
1-90: LGTM!plugins/lib/tests/spec.lua (1)
1300-1321: LGTM!plugins/task/tests/spec.lua (1)
18-18: LGTM!noon-lua/tests/task_policy.rs (1)
119-119: LGTM!Also applies to: 148-148, 352-355
site/docs/content/configuration/_index.md (1)
198-198: LGTM!plugins/task/init.lua (1)
154-160: 🎯 Functional Correctness
plugins.task.auto_tieris already used here. The handler falls back toopts.auto_tierwheninput.auto_tieris unset.> Likely an incorrect or invalid review comment.noon-agent/src/template.rs (1)
3-3: LGTM!Also applies to: 41-48, 86-102
noon-agent/src/tools/registry.rs (1)
479-487: LGTM!Also applies to: 587-602
noon-ui/src/agent/agent_loop.rs (2)
11-11: LGTM!Also applies to: 50-61, 106-106, 159-160, 186-186, 313-322
277-322: 🗄️ Data Integrity & IntegrationInclude config-derived inputs in
ToolsCache
build_toolsdepends onToolFilter::from_config(&self.config, model, &[]), sodisabled_toolscan change the tool set.ToolsCachedoesn’t include anyself.config-derived hash, so a live config update could reuse a stale tool list. IfAgentConfigis immutable for the lifetime ofAgentLoop, this is fine; otherwise, add config state to the cache key.noon-lua/src/loader.rs (1)
94-97: LGTM!noon-config/src/lib.rs (1)
56-56: LGTM!noon-docgen/src/gen_tools.rs (1)
37-37: LGTM!noon-lua/tests/spec.rs (1)
7-7: LGTM!site/docs/content/tools/_index.md (1)
10-10: LGTM!Also applies to: 158-181
plugins/almas/mem.lua (1)
1-47: LGTM!plugins/almas/refine.lua (1)
6-68: LGTM!plugins/almas/retrieve.lua (1)
12-128: LGTM!
| case("retrieve_vector_fallback", function() | ||
| local retrieve = require("retrieve") | ||
|
|
||
| -- Mock noon.agent.call_tool to return empty for first call, then a match | ||
| local old_call = noon.agent.call_tool | ||
| local calls = 0 | ||
| noon.agent.call_tool = function(ctx, name, args) | ||
| if name == "grep" then | ||
| calls = calls + 1 | ||
| if calls == 1 then | ||
| return "" | ||
| else | ||
| return "some match with vector similarity" | ||
| end | ||
| end | ||
| return nil | ||
| end | ||
|
|
||
| local dummy_ctx = {} | ||
| local block = retrieve.retrieve(dummy_ctx, "add retry helper", "developer", 2) | ||
|
|
||
| -- Restore original | ||
| noon.agent.call_tool = old_call | ||
|
|
||
| assert(block ~= nil, "retrieved block should fallback to vector and not be nil") | ||
| assert(block:find("vector similarity") ~= nil, "retrieved block should contain mock results") | ||
| end) |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win
retrieve_vector_fallback does not actually exercise the vector fallback.
keywords("add retry helper") yields {"retry","helper"}. The mock returns "" on the first grep (call 1) and a match on the second (call 2). Inside retrieve_lexical, the second keyword's non-empty result is picked, so retrieve returns via the lexical path and retrieve_vector is never called — yet the assertions still pass. To genuinely test the fallback, force lexical to return nothing (e.g. always return ""/empty for the lexical pass) so retrieve must fall through to retrieve_vector.
🧰 Tools
🪛 Luacheck (1.2.0)
[warning] 77-77: unused argument 'ctx'
(W212)
[warning] 77-77: unused argument 'args'
(W212)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@plugins/almas/tests/spec.lua` around lines 71 - 97, The
retrieve_vector_fallback test currently allows the second lexical keyword search
to return a match, so retrieve.retrieve can succeed without invoking
retrieve_vector. Update the noon.agent.call_tool mock in
retrieve_vector_fallback to return an empty result for every lexical grep call,
while returning the vector fallback match for the appropriate vector query,
ensuring the assertions exercise the actual fallback path.
noon.agent.session expects `tools` to be a JSON array. An empty Lua table
`{}` serializes to an object, so leave the key out to let the Rust side
use the default empty array for model refinement and digest sessions.
Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
Signed-off-by: w0wl0lxd <w0wl0lxd@tuta.com>
This PR adds model-based prompt refinement (HALO), digested learning summaries for the project memory loop, and expands test coverage for retriever and roles in the ALMAS plugin.