Skip to content

test: add fuzz coverage for orchestrator inputs - #53

Merged
seonghobae merged 1 commit into
mainfrom
codex/fix-contextual-orchestrator-security-gates
Jul 10, 2026
Merged

test: add fuzz coverage for orchestrator inputs#53
seonghobae merged 1 commit into
mainfrom
codex/fix-contextual-orchestrator-security-gates

Conversation

@seonghobae

Copy link
Copy Markdown
Contributor

Summary

  • add Hypothesis property tests and Atheris harnesses for request parsing, agent config parsing, redaction, and offline orchestration
  • pin fuzz CI dependencies with hashes and fix the security workflow audit-tool install pins
  • add fuzz optional dependency metadata and keep normal pytest collection away from native Atheris harness modules

Evidence

  • CodeGraph: explored current-head untrusted-input and workflow surfaces
  • git diff --check
  • Ruby YAML parse for .github/workflows/security.yml and .github/workflows/fuzz.yml
  • TOML parse for pyproject.toml
  • uv run --with pytest --with hypothesis python -m pytest tests/fuzz -q => 7 passed
  • uv run --with pytest --with hypothesis python -m pytest -q => 197 passed

Failure mapping

  • fixes FuzzingID by adding coverage-guided fuzzing plus deterministic property tests
  • fixes PinnedDependenciesID in .github/workflows/security.yml by pinning pip-audit==2.10.1 and cyclonedx-bom==7.3.0
  • fixes the old PR test(fuzz): add coverage-guided + property-based fuzzing #47 runtime failures by using available atheris==3.0.0 and ensuring Hypothesis is installed in the review/coverage environment

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head 73de9e6263e16828d77ff1948be4cc1a72a32177.

  • Head SHA: 73de9e6263e16828d77ff1948be4cc1a72a32177

  • Workflow run: 29084702066

  • Workflow attempt: 1

Coverage evidence

Coverage Evidence

  • Head SHA: 73de9e6263e16828d77ff1948be4cc1a72a32177
  • Required test evidence: supported repository test suites must pass.
  • Required docstring evidence: repository-owned docstring gates must pass when configured; otherwise docstring coverage is advisory.

Python project dependencies (.)

Using CPython 3.12.3 interpreter at: /usr/bin/python3
Creating virtual environment at: .venv
Resolved 28 packages in 266ms
Checked in 0.00ms
  • Result: PASS

Python coverage with missing-line report (.)

Downloading pygments (1.2MiB)
 Downloaded pygments
Installed 6 packages in 10ms
============================= test session starts ==============================
platform linux -- Python 3.12.3, pytest-9.1.1, pluggy-1.6.0
rootdir: /home/runner/work/contextual-orchestrator/contextual-orchestrator
configfile: pyproject.toml
collected 190 items / 1 error

==================================== ERRORS ====================================
_________ ERROR collecting pr-head/tests/fuzz/test_fuzz_properties.py __________
ImportError while importing test module '/home/runner/work/contextual-orchestrator/contextual-orchestrator/pr-head/tests/fuzz/test_fuzz_properties.py'.
Hint: make sure your test modules/packages have valid Python names.
Traceback:
/usr/lib/python3.12/importlib/__init__.py:90: in import_module
    return _bootstrap._gcd_import(name[level:], package, level)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
tests/fuzz/test_fuzz_properties.py:17: in <module>
    from hypothesis import given, settings, strategies as st
E   ModuleNotFoundError: No module named 'hypothesis'
=========================== short test summary info ============================
ERROR tests/fuzz/test_fuzz_properties.py
!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!
=============================== 1 error in 2.08s ===============================
  • Result: FAIL (exit 2)

Python docstring coverage advisory

RESULT: PASSED (minimum: 80.0%, actual: 92.5%)
  • Result: PASS

Coverage Decision

  • Result: FAIL
  • Test evidence: not proven passing
  • Docstring evidence: not proven passing when configured
  • Failure count: 1

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Workflow (2 files)"]
  S1 --> I1["GitHub Actions review job"]
  I1 --> R1["Review risk: Workflow (2 files)"]
  R1 --> V1["actionlint plus required checks"]
  Evidence --> S2["Changed file (34 files)"]
  S2 --> I2["repository behavior"]
  I2 --> R2["Review risk: Changed file (34 files)"]
  R2 --> V2["required checks"]
  Evidence --> S3["Docs (2 files)"]
  S3 --> I3["operator or user guidance"]
  I3 --> R3["Review risk: Docs (2 files)"]
  R3 --> V3["docs review"]
  Evidence --> S4["Test: test_fuzz_properties.py"]
  S4 --> I4["regression suite"]
  I4 --> R4["Review risk: Test: test_fuzz_properties.py"]
  R4 --> V4["targeted test run"]
Loading

@github-actions

github-actions Bot commented Jul 10, 2026

Copy link
Copy Markdown

OpenCode Review Overview

  • Head SHA: 734a9c5b339bddf3421b90a8250410c68047fd6c
  • Workflow run: 29085669950
  • Workflow attempt: 1
  • Gate result: APPROVE (approval step)

Pull request overview

OpenCode reviewed the current-head bounded evidence and found no blocking issues.

Findings

No blocking findings.

Summary

Approval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Verification posture: CodeGraph evidence was initialized and bounded current-head evidence reviewed for changed-file evidence including .github/workflows/fuzz.yml, .github/workflows/security.yml, conftest.py, contextual_orchestrator/server.py, docs/fuzzing.md, and 32 more.
Linter/static: workflow/static review evidence is bounded by the current-head GitHub Checks gate and changed-file evidence.
TDD/regression: coverage execution evidence and focused changed hunks were reviewed from bounded-review-evidence.md.
Coverage: coverage execution evidence reports supported repository test suites passed.
Docstring coverage: coverage execution evidence reports configured repository docstring gates passed or docstring coverage was advisory.
DAG: CodeGraph/source-backed behavior map connects .github/workflows/fuzz.yml to the affected review, runtime, or workflow path and required checks.
PoC/execution: coverage-evidence job executed on the current head and reported PASS.
DDD/domain: workflow and repository-governance invariants were reviewed against changed files in bounded evidence.
CDD/context: CodeGraph evidence, changed-file history, and focused hunks were reviewed from bounded-review-evidence.md.
Similar issues: changed-file history evidence was reviewed for comparable local precedents.
Claim/concept check: bounded evidence, repository source, current-head workflow evidence, and, where numeric, scientific, statistical, or literature-backed claims are affected, original-paper/formula evidence and parameter-recovery expectations were used for claims.
Standards search: standards and external-source checks are delegated to configured OpenCode web_search/Context7/DeepWiki sources when applicable; no evidence-backed standards blocker is present in bounded evidence.
Compatibility/convention: changed workflow/script conventions, object naming, and reserved-word safety for schema/API/config/code surfaces were checked in bounded evidence.
Breaking-change/backcompat: deployment evidence and changed-file history were checked for backward-compatibility risk.
Performance: changed surfaces were checked for performance risk in bounded evidence.
Developer experience: changed automation, review, test, setup, and maintenance surfaces were checked for helpful or obstructive DX impact in bounded evidence.
User experience: connected user, operator, API, CLI, documentation, review-comment, status-check, rendering, and workflow-reader behavior was checked for contradictions against code, docs, and tests in bounded evidence.
Visual/DOM: Playwright visual, DOM locator, ARIA snapshot, console, and responsive evidence were checked when a web UI surface was present; for non-web surfaces, API/CLI/log/docs/workflow interaction evidence was reviewed instead.
Accessibility/i18n: accessibility, localization, and human-readable text surfaces were checked where UI, CLI, API message, docs, logs, or review text changed.
Supply-chain/license: dependency, package, model, container, and external-tool changes were checked in bounded evidence.
Packaging: package, build, test, lint, and security contracts were checked in bounded evidence.
Security/privacy: workflow-token, review-gate, and repository-automation security/privacy boundaries were checked in bounded evidence.

  • Result: APPROVE
  • Reason: Tests pass, coverage is acceptable, and no unresolved issues remain.
  • Head SHA: 734a9c5b339bddf3421b90a8250410c68047fd6c
  • Workflow run: 29085669950
  • Workflow attempt: 1

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Workflow (2 files)"]
  S1 --> I1["GitHub Actions review job"]
  I1 --> R1["Review risk: Workflow (2 files)"]
  R1 --> V1["actionlint plus required checks"]
  Evidence --> S2["Changed file (36 files)"]
  S2 --> I2["repository behavior"]
  I2 --> R2["Review risk: Changed file (36 files)"]
  R2 --> V2["required checks"]
  Evidence --> S3["Docs (2 files)"]
  S3 --> I3["operator or user guidance"]
  I3 --> R3["Review risk: Docs (2 files)"]
  R3 --> V3["docs review"]
  Evidence --> S4["Test: test_fuzz_properties.py"]
  S4 --> I4["regression suite"]
  I4 --> R4["Review risk: Test: test_fuzz_properties.py"]
  R4 --> V4["targeted test run"]
Loading

@seonghobae
seonghobae force-pushed the codex/fix-contextual-orchestrator-security-gates branch from 73de9e6 to 05a57ea Compare July 10, 2026 10:05
@seonghobae
seonghobae force-pushed the codex/fix-contextual-orchestrator-security-gates branch from 05a57ea to 734a9c5 Compare July 10, 2026 10:13

@opencode-agent opencode-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode reviewed the current-head bounded evidence and found no blocking issues.

Findings

No blocking findings.

Summary

Approval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Verification posture: CodeGraph evidence was initialized and bounded current-head evidence reviewed for changed-file evidence including .github/workflows/fuzz.yml, .github/workflows/security.yml, conftest.py, contextual_orchestrator/server.py, docs/fuzzing.md, and 32 more.
Linter/static: workflow/static review evidence is bounded by the current-head GitHub Checks gate and changed-file evidence.
TDD/regression: coverage execution evidence and focused changed hunks were reviewed from bounded-review-evidence.md.
Coverage: coverage execution evidence reports supported repository test suites passed.
Docstring coverage: coverage execution evidence reports configured repository docstring gates passed or docstring coverage was advisory.
DAG: CodeGraph/source-backed behavior map connects .github/workflows/fuzz.yml to the affected review, runtime, or workflow path and required checks.
PoC/execution: coverage-evidence job executed on the current head and reported PASS.
DDD/domain: workflow and repository-governance invariants were reviewed against changed files in bounded evidence.
CDD/context: CodeGraph evidence, changed-file history, and focused hunks were reviewed from bounded-review-evidence.md.
Similar issues: changed-file history evidence was reviewed for comparable local precedents.
Claim/concept check: bounded evidence, repository source, current-head workflow evidence, and, where numeric, scientific, statistical, or literature-backed claims are affected, original-paper/formula evidence and parameter-recovery expectations were used for claims.
Standards search: standards and external-source checks are delegated to configured OpenCode web_search/Context7/DeepWiki sources when applicable; no evidence-backed standards blocker is present in bounded evidence.
Compatibility/convention: changed workflow/script conventions, object naming, and reserved-word safety for schema/API/config/code surfaces were checked in bounded evidence.
Breaking-change/backcompat: deployment evidence and changed-file history were checked for backward-compatibility risk.
Performance: changed surfaces were checked for performance risk in bounded evidence.
Developer experience: changed automation, review, test, setup, and maintenance surfaces were checked for helpful or obstructive DX impact in bounded evidence.
User experience: connected user, operator, API, CLI, documentation, review-comment, status-check, rendering, and workflow-reader behavior was checked for contradictions against code, docs, and tests in bounded evidence.
Visual/DOM: Playwright visual, DOM locator, ARIA snapshot, console, and responsive evidence were checked when a web UI surface was present; for non-web surfaces, API/CLI/log/docs/workflow interaction evidence was reviewed instead.
Accessibility/i18n: accessibility, localization, and human-readable text surfaces were checked where UI, CLI, API message, docs, logs, or review text changed.
Supply-chain/license: dependency, package, model, container, and external-tool changes were checked in bounded evidence.
Packaging: package, build, test, lint, and security contracts were checked in bounded evidence.
Security/privacy: workflow-token, review-gate, and repository-automation security/privacy boundaries were checked in bounded evidence.

  • Result: APPROVE
  • Reason: Tests pass, coverage is acceptable, and no unresolved issues remain.
  • Head SHA: 734a9c5b339bddf3421b90a8250410c68047fd6c
  • Workflow run: 29085669950
  • Workflow attempt: 1

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Workflow (2 files)"]
  S1 --> I1["GitHub Actions review job"]
  I1 --> R1["Review risk: Workflow (2 files)"]
  R1 --> V1["actionlint plus required checks"]
  Evidence --> S2["Changed file (36 files)"]
  S2 --> I2["repository behavior"]
  I2 --> R2["Review risk: Changed file (36 files)"]
  R2 --> V2["required checks"]
  Evidence --> S3["Docs (2 files)"]
  S3 --> I3["operator or user guidance"]
  I3 --> R3["Review risk: Docs (2 files)"]
  R3 --> V3["docs review"]
  Evidence --> S4["Test: test_fuzz_properties.py"]
  S4 --> I4["regression suite"]
  I4 --> R4["Review risk: Test: test_fuzz_properties.py"]
  R4 --> V4["targeted test run"]
Loading

@seonghobae
seonghobae dismissed github-actions[bot]’s stale review July 10, 2026 10:28

Dismissed stale OpenCode coverage-evidence failure review from prior head 73de9e6 after current head 734a9c5 passed coverage-evidence, Fuzz, Security, Strix, code scanning, and received current-head OpenCode approval.

@seonghobae
seonghobae merged commit 2ca8312 into main Jul 10, 2026
28 checks passed
@seonghobae
seonghobae deleted the codex/fix-contextual-orchestrator-security-gates branch July 10, 2026 10:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant