Skip to content

test: comprehensive QA for genie v2 — 76 new tests exploring limits - #556

Closed
namastex888 wants to merge 1 commit into
devfrom
test/v2-qa-limits
Closed

namastex888 wants to merge 1 commit into
devfrom
test/v2-qa-limits

Conversation

@namastex888

Copy link
Copy Markdown
Contributor

Summary

  • 76 new tests across 9 files — unit boundaries, concurrency, stress, edge cases
  • Covers parseWishGroups() (had zero coverage), wish state machine edge cases (cycles, diamond deps, corruption), file lock concurrency, team-chat/mailbox boundaries, provider adapter shell injection safety, and stress tests (500 agents, 100 concurrent messages, 50-group wish)
  • 768 total tests pass, bun run check exits 0

Bugs Found

Issue Severity Description
#546 High Built-in agent spawn ignores team worktree CWD
#547 Critical mailbox.send() loses messages under concurrent writes (no file lock)
#548 Medium No cycle detection in wish state dependency graph
#549 Design ready->done transition skips in_progress state
#550 Refactor File lock code duplicated 3x — extract shared utility
#551 Medium Team name not validated against git branch naming rules
#552 Enhancement No failed state in wish state machine
#553 Low getState() reads without lock (TOCTOU)
#554 Medium parseWishGroups() case-sensitive — silent failure on format variations
#555 Medium team-chat JSONL postMessage() has no file lock

Test plan

  • bun test — 768 pass, 0 fail
  • bun run check — exits 0 (typecheck + lint + dead-code + test)
  • All new tests run in CI alongside existing suite

Covers parseWishGroups (zero coverage), wish state edge cases (cycles,
diamonds, corruption), file lock concurrency, provider adapter shell
injection and flag selection, team manager idempotency, mailbox race
condition (confirms G-10 bug), and stress tests (500 agents, 100
concurrent chat, 50-group wish lifecycle).

768 tests total, 0 failures, quality gate green.
@coderabbitai

coderabbitai Bot commented Mar 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 29fe4e95-6a0a-484c-9d8a-82e86173fb0e

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch test/v2-qa-limits
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request dramatically enhances the quality assurance for Genie v2 by introducing a large suite of new tests. The primary goal was to uncover hidden issues and validate the robustness of core functionalities under various conditions, including high load, concurrency, and malformed inputs. This effort successfully identified several critical and high-severity bugs, improving the overall stability and reliability of the system.

Highlights

  • Comprehensive Test Suite Expansion: Added 76 new tests across 9 files, significantly increasing coverage for critical components like parseWishGroups() (which previously had zero coverage), wish state machine edge cases, file lock concurrency, team-chat/mailbox boundaries, and provider adapter shell injection safety.
  • Critical Bug Identification: The new tests identified and documented several high-severity bugs, including a critical issue where mailbox.send() loses messages under concurrent writes due to a missing file lock, and a high-severity bug where built-in agent spawn ignores team worktree CWD.
  • Robustness and Edge Case Handling: New tests specifically target unit boundaries, concurrency, stress scenarios (e.g., 500 agents, 100 concurrent messages, 50-group wishes), and various edge cases such as circular dependencies, corrupted data files, and shell injection vulnerabilities.
  • Improved Wish State Management Validation: Extensive testing of the wish state machine now covers complex dependency graphs (cycles, diamond dependencies), state transitions, and error handling for corrupted state files.
Changelog
Activity
  • 768 total tests pass after the changes.
  • bun run check exits with 0, indicating no type errors, linting issues, or dead code.
  • All new tests are configured to run in CI alongside the existing test suite.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a comprehensive suite of 76 new tests, significantly improving quality assurance for genie v2. The tests cover stress scenarios, concurrency issues, and numerous edge cases across various modules, including parseWishGroups() which previously had no coverage. The new tests are well-designed and have successfully uncovered several bugs as detailed in the pull request description. My review identifies one opportunity for improvement in the new tests for parseWishGroups() to encourage a more robust, case-insensitive implementation. Overall, this is an excellent contribution that greatly enhances the project's stability and correctness.

Comment on lines +348 to +352
it('should return empty for lowercase "### group 1:" (regex is case-sensitive)', () => {
const content = '### group 1: lowercase\n**depends-on:** none\n';
const groups = parseWishGroups(content);
expect(groups).toEqual([]);
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This test correctly documents the current case-sensitive parsing of group headings, which is noted as a bug (#554). For improved robustness and user experience, the parsing should be case-insensitive, similar to how depends-on is handled.

I suggest updating this test to assert the desired case-insensitive behavior. This would cause the test to fail initially, guiding the implementation fix in parseWishGroups (by adding the i flag to the groupPattern regex).

Suggested change
it('should return empty for lowercase "### group 1:" (regex is case-sensitive)', () => {
const content = '### group 1: lowercase\n**depends-on:** none\n';
const groups = parseWishGroups(content);
expect(groups).toEqual([]);
});
it('should parse group headings case-insensitively', () => {
const content = '### group 1: lowercase\n**depends-on:** none\n';
const groups = parseWishGroups(content);
expect(groups).toEqual([{ name: '1', dependsOn: [] }]);
});

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 125ba5ace1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

const parseTime = Date.now() - parseStart;

expect(entries.length).toBe(500);
expect(parseTime).toBeLessThan(1000); // Parse time < 1s

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Remove machine-specific latency assertions from stress tests

These expectations make pass/fail depend on fixed wall-clock budgets (<1s, <500ms, <10s) rather than correctness, so the suite can fail on slower CI runners or loaded hosts even when behavior is functionally correct. Since each test already has explicit timeout limits, these extra micro-performance thresholds introduce avoidable flakiness across environments.

Useful? React with 👍 / 👎.

Comment on lines +359 to +362
if (!existsSync(wishPath)) {
// Skip if running from a different CWD
console.log('Skipping U-DC-09: WISH.md not found at', wishPath);
return;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid silently skipping parser assertions on missing local file

This early return means the core assertions for U-DC-09 are skipped whenever that repo-local WISH path is absent (for example, different working directories or minimal checkouts), so the test can pass without exercising the intended behavior. Using a checked-in fixture would keep coverage deterministic instead of silently turning this into a no-op.

Useful? React with 👍 / 👎.

}

// At minimum, one message should survive
expect(messages.length).toBeGreaterThanOrEqual(1);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Make mailbox concurrency data-loss check fail explicitly

This assertion allows the test to pass even when most concurrent messages are lost (e.g., 1 out of 10 survives), which defeats the stated goal of catching mailbox write races and gives false confidence in a corruption-prone path. If this behavior is known-broken, it should be marked expected-failure/todo; otherwise the test should require all writes to persist.

Useful? React with 👍 / 👎.

@namastex888

Copy link
Copy Markdown
Contributor Author

Superseded by #560 which included updated versions of these tests adapted for the v2 fixes. Dev now has 734 tests covering the same areas.

@namastex888
namastex888 deleted the test/v2-qa-limits branch March 29, 2026 22:10
namastex888 added a commit that referenced this pull request Apr 29, 2026
…#1537)

Group 2 of the omni-host-fingerprint-trust wish (D5 follow-up). First
genie-side piece of per-host fingerprint trust: generate a local
ed25519 keypair, register the public key with the local omni server
via POST /api/v2/trust/handshake, and persist the returned host_id
locally so subsequent groups (request signing, verification) can
attach `X-Genie-Host-Id` to outgoing requests.

Builds on omni #555/#556/#558 (the schema + handshake endpoint + trust
CRUD endpoints).

CLI surface
===========
  genie omni handshake               One-time registration (idempotent on pubkey)
  genie omni handshake --rotate      New keypair + revoke old in a single round-trip
  genie omni handshake --hostname X  Override os.hostname() for the omni record

Files written
=============
  ~/.genie/keys/genie-host.ed25519       PKCS#8 PEM, 0600 perms (private)
  ~/.genie/keys/genie-host.ed25519.pub   base64url of raw 32-byte pubkey
  ~/.genie/keys/host.json                { hostId, pubkey, hostname, registeredAt, rotatedFrom? }

Sanity checks
=============
  - Refuses to write keys inside a git working tree (`assertNotInsideGitRepo`)
    so an accidental `genie omni handshake` from a project root doesn't
    stage the secret key for the next commit. Walk up to fs root or 16
    levels, whichever comes first.
  - `--rotate` requires an existing host record. Generates the new keypair,
    registers it, then revokes the OLD record. Order matters: revoke fails
    after register, so we never lose access. If revoke fails post-register,
    the new key is live and we surface the manual recovery command.

Auth: bearer token from genie config or $OMNI_API_KEY. The first handshake
always uses bearer because that's the only way to bootstrap trust for a
brand-new host. Subsequent signed requests (Group 3) can authenticate
themselves.

What's NOT in
=============
  - Signing outgoing requests (Group 3): the keypair lives here, but
    `omni-registration.ts` doesn't read it yet.
  - Verification middleware on omni (Group 4, security review gate):
    the host record is stored, but no incoming request is verified yet.

Tests
=====
  9 tests pinning:
    - keyPaths respects $GENIE_HOME (test isolation)
    - assertNotInsideGitRepo throws on git tree, passes on plain dir
    - generateAndPersistKeypair → 0600 perms + 43-char base64url pubkey
    - host.json round-trip (load null, write/load, malformed → null)
    - regenerating overwrites the keypair

The HTTP path is exercised by the omni-side tests in #556/#558 — we
don't re-test the omni contract here, just the local filesystem
invariants.

Tracked under omni-host-fingerprint-trust wish, Group 2.

Co-authored-by: Genie <genie@namastex.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant