feat(content-graph): parse structured non-PDF attachments - #1038
Conversation
|
PR governance metadata gate is not ready for
|
OpenCode Review Overview
Pull request overviewOpenCode reviewed the current-head bounded evidence and found no blocking issues. FindingsNo blocking findings. SummaryApproval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Changed-File Evidence Mapflowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file: .gitattributes"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file: .gitattributes"]
R1 --> V1["required checks"]
Evidence --> S2["Backend (7 files)"]
S2 --> I2["API and service runtime"]
I2 --> R2["Review risk: Backend (7 files)"]
R2 --> V2["backend tests"]
Evidence --> S3["Docs (12 files)"]
S3 --> I3["operator or user guidance"]
I3 --> R3["Review risk: Docs (12 files)"]
R3 --> V3["docs review"]
|
There was a problem hiding this comment.
Pull request overview
OpenCode reviewed the current-head bounded evidence and found no blocking issues.
Findings
No blocking findings.
Summary
Approval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Verification posture: CodeGraph evidence was initialized and bounded current-head evidence reviewed for changed-file evidence including backend/services/attachment_parser.py, backend/services/content_graph/parser.py, backend/services/email_import_service.py, backend/tests/test_attachment_parser.py, backend/tests/test_content_graph_parser.py, and 11 more.
Linter/static: workflow/static review evidence is bounded by the current-head GitHub Checks gate and changed-file evidence.
TDD/regression: coverage execution evidence and focused changed hunks were reviewed from bounded-review-evidence.md.
Coverage: coverage execution evidence reports supported repository test suites passed.
Docstring coverage: coverage execution evidence reports configured repository docstring gates passed or docstring coverage was advisory.
DAG: CodeGraph/source-backed behavior map connects backend/services/attachment_parser.py to the affected review, runtime, or workflow path and required checks.
PoC/execution: coverage-evidence job executed on the current head and reported PASS.
DDD/domain: workflow and repository-governance invariants were reviewed against changed files in bounded evidence.
CDD/context: CodeGraph evidence, changed-file history, and focused hunks were reviewed from bounded-review-evidence.md.
Similar issues: changed-file history evidence was reviewed for comparable local precedents.
Claim/concept check: bounded evidence, repository source, current-head workflow evidence, and, where numeric, scientific, statistical, or literature-backed claims are affected, original-paper/formula evidence and parameter-recovery expectations were used for claims.
Standards search: standards and external-source checks are delegated to configured OpenCode web_search/Context7/DeepWiki sources when applicable; no evidence-backed standards blocker is present in bounded evidence.
Compatibility/convention: changed workflow/script conventions, object naming, and reserved-word safety for schema/API/config/code surfaces were checked in bounded evidence.
Breaking-change/backcompat: deployment evidence and changed-file history were checked for backward-compatibility risk.
Performance: changed surfaces were checked for performance risk in bounded evidence.
Developer experience: changed automation, review, test, setup, and maintenance surfaces were checked for helpful or obstructive DX impact in bounded evidence.
User experience: connected user, operator, API, CLI, documentation, review-comment, status-check, rendering, and workflow-reader behavior was checked for contradictions against code, docs, and tests in bounded evidence.
Visual/DOM: Playwright visual, DOM locator, ARIA snapshot, console, and responsive evidence were checked when a web UI surface was present; for non-web surfaces, API/CLI/log/docs/workflow interaction evidence was reviewed instead.
Accessibility/i18n: accessibility, localization, and human-readable text surfaces were checked where UI, CLI, API message, docs, logs, or review text changed.
Supply-chain/license: dependency, package, model, container, and external-tool changes were checked in bounded evidence.
Packaging: package, build, test, lint, and security contracts were checked in bounded evidence.
Security/privacy: workflow-token, review-gate, and repository-automation security/privacy boundaries were checked in bounded evidence.
- Result: APPROVE
- Reason: Implements structured attachment parsing for JSON, CSV, XML, and iCalendar with tests passing and security considered.
- Head SHA:
961df95836aead1e3d3f02ee48d574cd68e6a66f - Workflow run: 29144537655
- Workflow attempt: 1
Changed-File Evidence Map
flowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file: .gitattributes"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file: .gitattributes"]
R1 --> V1["required checks"]
Evidence --> S2["Backend (7 files)"]
S2 --> I2["API and service runtime"]
I2 --> R2["Review risk: Backend (7 files)"]
R2 --> V2["backend tests"]
Evidence --> S3["Docs (12 files)"]
S3 --> I3["operator or user guidance"]
I3 --> R3["Review risk: Docs (12 files)"]
R3 --> V3["docs review"]
Summary
Implements the Project #1 P0 slice for #1021 under the ContextualWisdomLab/.github#363 project protocol.
backend/services/content_graph/parser beyond text/html, Markdown, and plain text to structured JSON, CSV, XML, and iCalendar attachmentsapplication/json,text/csv,application/xml,text/calendar, and extension fallback paths reach email import/content graph persistencedocs/research/non-pdf-dom-evidence/, including original PDF files for WHATWG HTML and two content-extraction papers plus original RFC/CommonMark standard text filesdocs/planning/naruon-platform-plan.md, which Project feat: setup backend foundation and archive extractor #1 and docs: cross-agent protocol for operating the GitHub Project (roadmap source of truth) .github#363 reference as the roadmap detail sourceGovernance and split decision
No new repo, submodule, package, domain, or product name was created in this slice. The repo-local architecture note says
content_graphshould remain internal until the parser API, segment schema, failure taxonomy, and golden corpus tests stabilize and a second consumer exists. This PR keeps that boundary while making extraction evidence-oriented enough to evaluate later as a standalone product.Git LFS is not used. The largest preserved PDF is about 16 MB.
Validation
python3 -m pytest -q tests/test_content_graph_parser.py tests/test_attachment_parser.py tests/test_email_parser.py tests/test_email_import_service.pypython3 -m ruff check services/content_graph/parser.py services/attachment_parser.py services/email_import_service.py tests/test_content_graph_parser.py tests/test_attachment_parser.py tests/test_email_parser.py tests/test_email_import_service.pygit diff --cached --checkReferences