From 1d4f59a077784f9951f5b7f2eb559096ccf517bb Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 10:36:39 +0900 Subject: [PATCH 01/15] docs: design standards-aligned file ontology --- ...026-07-21-semantic-file-ontology-design.md | 156 ++++++++++++++++++ 1 file changed, 156 insertions(+) create mode 100644 docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md diff --git a/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md b/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md new file mode 100644 index 0000000..b5f7a44 --- /dev/null +++ b/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md @@ -0,0 +1,156 @@ +# 국제 표준 기반 하이브리드 파일 온톨로지 설계 + +## 목적 + +`semantic-data-portal`에 파일의 의미 정체성과 물리 저장 위치를 분리하는 상위 카탈로그를 추가한다. 사용자는 파일이 로컬 디스크, Synology 동기화 폴더, AWS S3, S3 호환 저장소, Azure Blob 중 어디에 있든 프로젝트·시스템·업무 단계·산출물 종류·주제로 검색하고 실제 위치를 찾을 수 있어야 한다. + +포털은 naruon 문서 지식 그래프 위의 카탈로그·거버넌스 계층이라는 기존 경계를 유지한다. 원문과 원문 조각은 파일을 읽고 LLM에 보내는 동안에만 사용하며 포털 그래프나 GitHub에는 저장하지 않는다. + +## 승인된 범위 + +- 기존 Apache AGE 그래프와 pgvector 의미 검색을 재사용한다. +- 파일 내용은 로컬에서 추출하며, 사용자가 승인한 최소 원문 조각만 OpenAI Responses API로 전송한다. +- 파일을 이동·삭제·복사하지 않는다. 수집은 읽기 전용이다. +- Synology는 필수 기반시설이 아니라 `filesystem` 저장소의 한 배치 형태다. +- 저장소 공급자는 `filesystem`, `s3`, `s3_compatible`, `azure_blob`을 지원한다. +- 공급자 자격증명과 OpenAI API key는 그래프·문서·로그·GitHub에 저장하지 않는다. +- LLM 결과는 근거 조각, 신뢰도, `proposed` 검토 상태를 가진 후보 주장으로 저장한다. LLM이 온톨로지의 정답을 직접 확정하지 않는다. +- OCR, HWP, legacy DOC/XLS 파서는 1차 파일럿에 포함하지 않는다. 실제 12개 파일은 PDF, DOCX, PPTX, XLSX로 구성되어 있다. + +## 국제 표준 프로파일 + +애플리케이션 프로파일 이름은 `CWL File Knowledge Profile 0.1`로 한다. + +| 목적 | 표준 | 적용 | +|---|---|---| +| 공통 그래프 모델 | RDF 1.1 | JSON-LD 및 Turtle 용어의 기준 모델 | +| 업무 개념과 제약 | OWL 2 | CWL 클래스와 객체 속성 정의 | +| 통제 어휘·동의어 | SKOS | 프로젝트, 시스템, 단계, 산출물, 주제 후보 | +| 구조 검증 | SHACL 1.0 | 파일 자산·배포 위치·근거 주장 shape | +| 카탈로그·배포·버전 | DCAT 3 | `dcat:Resource`, `dcat:Distribution`, 버전 관계 | +| 문서 메타데이터 | DCMI Terms | 제목, 형식, 생성·수정 시각, 식별자, 주제 | +| 출처·파생 관계 | PROV-O | `prov:Entity`, `prov:wasDerivedFrom`, 생성 근거 | +| 체크섬 | SPDX terms | SHA-256 checksum 표현 | +| 교환 | JSON-LD 1.1 | 포털 API의 표준 내보내기 | +| 외부 질의 호환 | SPARQL 1.1 | 내보낸 RDF를 표준 RDF store에서 질의 가능 | + +RDF 1.2와 SHACL 1.2는 2026-07-21 현재 안정된 Recommendation 기준선이 아니므로 채택하지 않고 향후 호환 대상으로만 추적한다. ISO/IEC 11179는 초기 파일 발견에 필요한 범위를 넘는 메타데이터 등록소 운영 부담이 있어 1차 범위에서 제외한다. + +## 핵심 모델 + +### FileAsset + +콘텐츠 SHA-256으로 식별되는 의미 자산이다. 하나의 파일이 이름이나 저장소만 달리해 복제되어도 같은 `FileAsset`이다. + +- RDF type: `cwl:FileAsset`, `dcat:Resource`, `prov:Entity` +- 필수 값: SHA-256, 제목, media type, byte size +- 선택 값: 생성·수정 시각, 이전 버전, 파생 원본 +- 그래프 node kind: `file_asset` +- node id: `urn:sha256:<64 lowercase hex>` + +### Distribution + +`FileAsset`의 물리적 접근 위치다. 한 자산은 0개 이상의 `dcat:Distribution`을 가진다. + +- 공급자: `filesystem`, `s3`, `s3_compatible`, `azure_blob` +- 공통 값: locator IRI, endpoint id, availability, ETag, version id, checksum +- 공급자 값: bucket 또는 container, object key/blob name +- Synology 경로는 별도 공급자 타입을 만들지 않고 `filesystem` distribution으로 기록한다. +- 기본 JSON-LD 응답은 locator를 숨긴다. 명시적으로 요청하고 정책 검사를 통과한 경우에만 실제 위치를 포함한다. + +### SemanticAssertion + +LLM 또는 규칙이 제안한 파일과 업무 개념 사이의 관계다. + +- 허용 관계: `belongsToProject`, `usesSystem`, `hasWorkPhase`, `hasArtifactType`, `hasTopic`, `wasDerivedFrom`, `previousVersion` +- 필수 값: 대상 개념, 근거 참조(chunk SHA-256과 문자 offset), 신뢰도(0..1), 추출 방법, 검토 상태 +- 모든 LLM 주장의 초기 상태는 `proposed`다. +- SHACL 호환 검증은 필수 필드·관계 allowlist·신뢰도 범위·근거 참조 존재를 검사한다. +- 주장은 그래프 탐색과 검색에 사용하되 응답에 검토 상태를 항상 노출한다. + +## 저장소 어댑터 + +핵심 수집기는 다음 세 연산만 요구한다. + +1. `list(prefix)` — 후보 object를 열거한다. +2. `stat(object)` — 크기, 수정 시각, ETag/version을 조회한다. +3. `read(object, max_bytes)` — 제한된 크기로 원문 bytes를 읽는다. + +구현은 다음과 같다. + +- `FilesystemReader`: `pathlib`로 로컬, UNC, Synology 동기화 폴더를 읽는다. +- `S3Reader`: 호출자가 주입한 boto3 호환 client를 사용한다. 사용자 지정 endpoint를 준 client면 S3 호환 저장소도 같은 구현을 쓴다. +- `AzureBlobReader`: 호출자가 주입한 Azure `ContainerClient` 호환 객체를 사용한다. + +SDK를 포털의 필수 의존성으로 추가하지 않는다. 실제 클라우드 배포 환경이 이미 사용하는 SDK client를 주입하며, CI는 같은 public method를 가진 fake client로 계약을 검증한다. + +## 문서 추출과 LLM 흐름 + +1. reader가 bytes와 저장소 메타데이터를 반환한다. +2. SHA-256을 계산해 `FileAsset`을 결정하고 중복 위치를 합친다. +3. DOCX/PPTX/XLSX는 Python 표준 `zipfile`·XML parser로 텍스트를 추출한다. +4. PDF는 `pypdf`로 텍스트를 추출한다. 이미지 전용 PDF는 `needs_ocr`로 보고하고 전송하지 않는다. +5. 텍스트를 고정 크기 조각으로 나누고 파일당 최대 전송 문자 수를 적용한다. +6. OpenAI Responses API의 strict JSON Schema 출력으로 의미 후보와 짧은 근거 인용을 받는다. +7. 근거 인용이 입력 조각에 실제로 존재하는지 확인한 뒤 인용문 대신 chunk SHA-256과 문자 offset만 남긴다. +8. 여러 조각의 후보를 `(관계, 정규화 label)` 기준으로 합치고 가장 높은 신뢰도와 그 근거 참조를 유지한다. +9. SHACL 호환 검증을 통과한 후보만 그래프에 `proposed` assertion으로 기록한다. +10. 파일 node의 embedding text는 제목과 제안된 개념 label만 사용한다. 원문과 근거 인용문은 저장하지 않는다. + +OpenAI API key는 extractor에 credential registry lookup으로 주입한다. 애플리케이션 코드가 `os.getenv()`로 key를 읽지 않는다. HTTP transport는 Python 표준 라이브러리를 사용하고 요청에 `store: false`와 pinned model id를 포함한다. CI에서는 HTTP 호출을 fake transport로 대체한다. + +## API와 정책 + +- `POST /file-assets`: 관리자만 검증된 파일 메타데이터와 후보 주장을 ingest한다. +- `GET /file-assets/{asset_id}`: 인증된 reader가 자산과 의미 관계를 조회한다. +- `GET /file-assets/{asset_id}/jsonld`: 기본적으로 저장소 locator를 제거한 JSON-LD를 반환한다. +- `GET /file-assets/{asset_id}/validate`: SHACL 호환 validation report를 반환한다. +- 기존 `POST /graph/query`와 `POST /search/semantic`을 그대로 사용해 관계 탐색과 의미 검색을 제공한다. + +모든 write/read route는 기존 `policy.evaluate()` 경로를 재사용한다. 원문 조각, API key, cloud credential은 request/response, graph property, audit detail에 포함하지 않는다. + +## 효성중공업 VOC 파일럿 + +초기 root는 다음 로컬 동기화 폴더다. + +`D:\SynologyDrive\업무자료\Download_정리_2026-07-15\01_문서` + +파일명에 `효성중공업` 또는 `중공업VOC`가 포함된 최상위 파일 12개를 읽기 전용으로 수집한다. + +- DOCX 4개 +- XLSX 4개 +- PPTX 2개 +- PDF 2개 +- SHA-256 기준 고유 `FileAsset` 10개 +- 같은 SHA-256을 가진 XLSX 3개는 자산 하나와 distribution 세 개로 모델링한다. + +실제 파일명, locator, 추출 원문, LLM 응답 전문은 GitHub에 commit하지 않는다. 파일럿 결과는 사용자 로컬의 Git 제외 경로에 JSON-LD/요약 manifest로 저장한다. + +## 실패와 안전 동작 + +- 허용 root 밖 path, 상위 경로 탈출, symlink/reparse point는 읽지 않는다. +- 파일 크기 제한을 넘으면 `too_large`로 기록하고 읽지 않는다. +- 지원하지 않는 형식은 `unsupported_format`, 추출 문자가 없으면 `needs_ocr`로 기록한다. +- API timeout, rate limit, 불완전 응답은 해당 파일을 `extraction_failed`로 남기며 기존 그래프를 덮어쓰지 않는다. +- storage credential과 OpenAI key가 없으면 fail closed 한다. +- 쓰기·삭제·이동·원격 object mutation 기능은 구현하지 않는다. + +## 검증 기준 + +- 같은 bytes의 서로 다른 locator가 한 `FileAsset`으로 합쳐진다. +- 각 공급자 reader가 list/stat/read 계약을 만족한다. +- 지원 문서에서 텍스트가 추출되고 원문은 graph property에 남지 않는다. +- OpenAI 요청은 strict schema, `store: false`, 최소 조각만 포함한다. +- 허용되지 않은 관계나 검증 가능한 근거 참조가 없는 후보는 validation에서 거부된다. +- JSON-LD가 DCAT/DCTERMS/PROV/SKOS/SPDX/CWL context와 locator redaction을 지킨다. +- 기존 전체 pytest suite와 새 단일 파일 지식 테스트가 통과한다. +- Codegraph를 동기화하고 변경 영향과 관련 테스트를 확인한다. + +## 제외 사항 + +- 파일 이동·삭제·중복 파일 제거 +- OCR 및 HWP/legacy Office parsing +- S3/Azure SDK를 필수 dependency로 추가 +- cloud object 쓰기 또는 lifecycle 관리 +- 자동 온톨로지 승인 +- 원문·embedding의 GitHub 업로드 From 0f051f77b54a6991d025669b7e1fa0fe68ba126a Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 10:42:06 +0900 Subject: [PATCH 02/15] docs: plan semantic file ontology implementation --- .../2026-07-21-semantic-file-ontology.md | 460 ++++++++++++++++++ 1 file changed, 460 insertions(+) create mode 100644 docs/superpowers/plans/2026-07-21-semantic-file-ontology.md diff --git a/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md b/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md new file mode 100644 index 0000000..f040d4f --- /dev/null +++ b/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md @@ -0,0 +1,460 @@ +# Semantic File Ontology Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add a standards-aligned, provider-neutral file knowledge catalog with read-only filesystem/S3/S3-compatible/Azure Blob ingestion, evidence-bound OpenAI semantic extraction, and a Hyosung Heavy Industries VOC pilot. + +**Architecture:** Keep Apache AGE and pgvector as the persistence/search implementation. Represent content-addressed `FileAsset` nodes separately from one-or-many DCAT `Distribution` nodes, project validated LLM assertions into the existing graph, and export JSON-LD. Read document bytes through small injected readers; raw text never enters graph properties or Git. + +**Tech Stack:** Python 3.10+, FastAPI, Pydantic 2, existing graph store, stdlib `pathlib`/`zipfile`/`urllib`, pypdf 6.14.2, W3C RDF/OWL/SKOS/SHACL/DCAT/PROV/JSON-LD vocabularies. + +## Global Constraints + +- File operations are read-only; no move, delete, copy, or remote mutation. +- Providers are exactly `filesystem`, `s3`, `s3_compatible`, and `azure_blob`; Synology is a filesystem deployment. +- Object bytes are capped at 20 MiB; LLM text is chunked at 6,000 characters with 300-character overlap and capped at 24,000 characters per file. +- The pinned default model is `gpt-5-mini-2025-08-07`; Requests API calls set `store: false` and a 60-second timeout. +- OpenAI API keys and cloud credentials come from an injected credential registry/client, never `os.getenv()`, graph properties, logs, CLI arguments, or GitHub fixtures. +- LLM assertions remain `proposed`; structural validation does not imply steward approval. +- Raw chunks and evidence quotations are not persisted. Persist only chunk SHA-256 and character offsets after verifying the quotation exists. +- Existing policy authorization, AGE/pgvector stores, semantic search, and graph traversal are reused. +- CI uses fake transports and fake cloud clients; it makes no OpenAI or cloud call. +- No OCR, HWP, DOC, XLS parser, mandatory cloud SDK, new database, or automatic ontology approval. + +--- + +### Task 1: Machine-readable CWL profile and file contracts + +**Files:** +- Create: `ontology/cwl-file-profile.ttl` +- Create: `ontology/cwl-file-shapes.ttl` +- Create: `src/sdp/file_ontology.py` +- Create: `tests/test_file_knowledge.py` + +**Interfaces:** +- Produces: `StorageDistribution`, `SemanticAssertion`, `FileAsset`, `validate_file_asset()`, `file_asset_jsonld()`. +- Consumes: Pydantic 2 only. + +- [ ] **Step 1: Write failing contract and JSON-LD tests** + +```python +def test_file_asset_jsonld_uses_standards_and_redacts_locator(): + asset = sample_asset() + payload = file_asset_jsonld(asset) + assert payload["@type"] == ["dcat:Resource", "prov:Entity", "cwl:FileAsset"] + assert "dcat:accessURL" not in payload["dcat:distribution"][0] + assert payload["spdx:checksum"]["spdx:checksumValue"] == asset.sha256 + +def test_file_asset_validation_rejects_unverifiable_evidence(): + asset = sample_asset(assertion_overrides={"evidence_chunk_sha256": "", "evidence_start": 4, "evidence_end": 2}) + report = validate_file_asset(asset) + assert report["conforms"] is False +``` + +- [ ] **Step 2: Run tests and confirm import failure** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -q` + +Expected: FAIL because `sdp.file_ontology` does not exist. + +- [ ] **Step 3: Add the minimal contracts and validation** + +```python +StorageProvider = Literal["filesystem", "s3", "s3_compatible", "azure_blob"] +AssertionRelation = Literal[ + "belongsToProject", "usesSystem", "hasWorkPhase", "hasArtifactType", + "hasTopic", "wasDerivedFrom", "previousVersion", +] + +class StorageDistribution(BaseModel): + id: str + provider: StorageProvider + locator: str + endpoint_id: str + available: bool = True + bucket: str | None = None + container: str | None = None + object_key: str | None = None + version_id: str | None = None + etag: str | None = None + +class SemanticAssertion(BaseModel): + relation: AssertionRelation + target_kind: Literal["business_project", "system", "work_phase", "artifact_type", "topic", "file_asset"] + target_label: str + confidence: float = Field(ge=0, le=1) + evidence_chunk_sha256: str + evidence_start: int = Field(ge=0) + evidence_end: int = Field(gt=0) + method: str = "openai" + review_status: Literal["proposed", "approved", "rejected"] = "proposed" + +class FileAsset(BaseModel): + sha256: str + title: str + media_type: str + byte_size: int = Field(ge=0) + modified_at: datetime | None = None + distributions: list[StorageDistribution] + assertions: list[SemanticAssertion] = Field(default_factory=list) + + @property + def asset_id(self) -> str: + return f"urn:sha256:{self.sha256}" +``` + +Validation must require lowercase 64-hex SHA-256, nonempty distributions, provider-specific bucket/container/object-key values, locator without query/fragment, relation/target-kind consistency, 64-hex chunk hash, and `evidence_end > evidence_start`. `file_asset_jsonld(asset, include_locations=False)` maps DCAT/DCTERMS/PROV/SKOS/SPDX/CWL terms and omits `dcat:accessURL` by default. + +- [ ] **Step 4: Add OWL/SKOS/DCAT/SHACL Turtle files and run tests** + +The profile defines `cwl:FileAsset` as a subclass of `dcat:Resource` and `prov:Entity`; `cwl:BusinessProject`, `cwl:System`, `cwl:WorkPhase`, and `cwl:ArtifactType` as SKOS concept subclasses; and the seven assertion properties as OWL object properties. The shapes require checksum, title, distribution, confidence, review status, and evidence reference fields. + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -q` + +Expected: PASS. + +- [ ] **Step 5: Commit** + +```bash +git add ontology src/sdp/file_ontology.py tests/test_file_knowledge.py +git commit -m "feat: define standards-aligned file ontology" +``` + +### Task 2: Read-only storage readers + +**Files:** +- Create: `src/sdp/storage_readers.py` +- Modify: `tests/test_file_knowledge.py` + +**Interfaces:** +- Produces: `ObjectRef`, `ObjectReader`, `FilesystemReader`, `S3Reader`, `AzureBlobReader`. +- Consumes: `StorageDistribution` from Task 1 and caller-injected cloud clients. + +- [ ] **Step 1: Write failing reader contract tests** + +```python +def test_filesystem_reader_stays_inside_root_and_reads_bytes(tmp_path): + path = tmp_path / "효성중공업 VOC.txt" + path.write_text("C-Cube PoC", encoding="utf-8") + reader = FilesystemReader(tmp_path) + ref = next(reader.list(name_pattern=r"효성중공업|중공업VOC")) + assert reader.read(ref, max_bytes=1024) == "C-Cube PoC".encode() + assert ref.distribution.provider == "filesystem" + +def test_s3_and_azure_readers_use_injected_clients(): + assert list(S3Reader(FakeS3(), "voc").list("reports/"))[0].object_key == "reports/a.docx" + assert list(AzureBlobReader(FakeContainer()).list("reports/"))[0].object_key == "reports/a.docx" +``` + +- [ ] **Step 2: Run focused tests and confirm failure** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k reader -q` + +Expected: FAIL because reader classes do not exist. + +- [ ] **Step 3: Implement the three readers** + +Define `ObjectReader.list(prefix: str = "", *, name_pattern: str | None = None) -> Iterable[ObjectRef]` and `ObjectReader.read(ref: ObjectRef, *, max_bytes: int) -> bytes`. `FilesystemReader` walks with `followlinks=False`, resolves each file, and checks `relative_to(root)` before opening. `S3Reader` uses `get_paginator("list_objects_v2")` and `get_object()["Body"].read(max_bytes + 1)`. `AzureBlobReader` uses `list_blobs(name_starts_with=prefix)` and `get_blob_client(name).download_blob().readall()`. + +Each reader rejects an object larger than `max_bytes`; filesystem traversal prunes symlink/junction directories and temporary Office files beginning `~$`. S3-compatible storage uses `S3Reader(client, bucket, provider="s3_compatible", endpoint_url="https://objects.example")`. + +- [ ] **Step 4: Run reader tests** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k reader -q` + +Expected: PASS. + +- [ ] **Step 5: Commit** + +```bash +git add src/sdp/storage_readers.py tests/test_file_knowledge.py +git commit -m "feat: add provider-neutral read-only storage readers" +``` + +### Task 3: Safe document text extraction + +**Files:** +- Create: `src/sdp/document_semantics.py` +- Modify: `tests/test_file_knowledge.py` +- Modify: `pyproject.toml` +- Modify: `requirements.txt` +- Modify: `requirements-dev.txt` +- Modify: `requirements-test.in` +- Modify: `requirements-test.txt` + +**Interfaces:** +- Produces: `DocumentText`, `TextChunk`, `extract_document_text()`, `chunk_text()`. +- Consumes: raw bytes and filename; does not persist text. + +- [ ] **Step 1: Write failing extraction/chunk tests** + +```python +def test_openxml_text_is_extracted_without_office_dependency(): + payload = make_openxml("word/document.xml", "효성중공업 VOC C-Cube") + assert "효성중공업" in extract_document_text("meeting.docx", payload).text + +def test_chunks_are_bounded_and_content_addressed(): + chunks = chunk_text("가" * 13000, max_chars=6000, overlap=300, max_total_chars=24000) + assert max(len(chunk.text) for chunk in chunks) <= 6000 + assert all(len(chunk.sha256) == 64 for chunk in chunks) +``` + +- [ ] **Step 2: Run tests and confirm failure** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k 'openxml or chunks or pdf' -q` + +Expected: FAIL because extraction functions do not exist. + +- [ ] **Step 3: Add `pypdf==6.14.2` and regenerate hashes** + +Add `pypdf==6.14.2` to `[project.dependencies]` and `requirements-test.in`. + +Run: + +```powershell +py -m pip install uv +py -m uv pip compile pyproject.toml --generate-hashes -o requirements.txt +py -m uv pip compile pyproject.toml --extra dev --generate-hashes -o requirements-dev.txt +py -m uv pip compile requirements-test.in --generate-hashes -o requirements-test.txt +py -m pip install pypdf==6.14.2 +``` + +Expected: all four dependency declarations contain pypdf 6.14.2 and hashes. + +- [ ] **Step 4: Implement extraction and chunking** + +`extract_document_text()` decodes text/Markdown/CSV/JSON/XML, reads XML elements whose local name is `t` from DOCX/PPTX/XLSX ZIP members, and uses `PdfReader(BytesIO(data)).pages[*].extract_text()` for PDF. It returns `status="needs_ocr"` when supported input yields no text and `status="unsupported_format"` for other suffixes. `chunk_text()` normalizes whitespace, uses the exact global limits, records global offsets and SHA-256, and returns no empty chunk. + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k 'openxml or chunks or pdf' -q` + +Expected: PASS. + +- [ ] **Step 5: Commit** + +```bash +git add pyproject.toml requirements*.txt src/sdp/document_semantics.py tests/test_file_knowledge.py +git commit -m "feat: extract text from pilot document formats" +``` + +### Task 4: Evidence-bound OpenAI semantic extraction + +**Files:** +- Modify: `src/sdp/document_semantics.py` +- Modify: `tests/test_file_knowledge.py` + +**Interfaces:** +- Produces: `CredentialRegistry`, `EphemeralCredentialRegistry`, `OpenAISemanticExtractor.extract()`. +- Consumes: `TextChunk`, injected credential registry, injected JSON HTTP transport. + +- [ ] **Step 1: Write a failing fake-transport test** + +```python +def test_openai_extractor_uses_strict_schema_and_persists_only_evidence_reference(): + captured = {} + extractor = OpenAISemanticExtractor( + EphemeralCredentialRegistry({"OPENAI_API_KEY": "test-key"}), + transport=fake_openai_transport(captured, evidence_quote="C-Cube"), + ) + assertions = extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) + assert captured["store"] is False + assert captured["text"]["format"]["type"] == "json_schema" + assert assertions[0].evidence_chunk_sha256 + assert "C-Cube" not in assertions[0].model_dump_json() + assert "test-key" not in repr(captured) +``` + +- [ ] **Step 2: Run the test and confirm failure** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k openai -q` + +Expected: FAIL because the extractor does not exist. + +- [ ] **Step 3: Implement the minimal Responses API client** + +Use `urllib.request.Request("https://api.openai.com/v1/responses", data=encoded_payload, headers=headers, method="POST")`, an injected `transport(request, timeout) -> dict`, model `gpt-5-mini-2025-08-07`, `store: false`, and strict JSON Schema fields `relation`, `target_kind`, `target_label`, `confidence`, `evidence_quote`. Reject empty/missing credential, non-completed responses, malformed JSON, unknown relations, and quotations not found verbatim in the input chunk. Convert a verified quote to chunk hash plus global offsets, then discard the quote. Merge duplicate `(relation, target_kind, casefold(target_label))` candidates by highest confidence. + +- [ ] **Step 4: Run OpenAI tests** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k openai -q` + +Expected: PASS with no network call. + +- [ ] **Step 5: Commit** + +```bash +git add src/sdp/document_semantics.py tests/test_file_knowledge.py +git commit -m "feat: extract evidence-bound semantics with OpenAI" +``` + +### Task 5: Graph projection and policy-protected API + +**Files:** +- Modify: `src/sdp/file_ontology.py` +- Modify: `src/sdp/api.py` +- Modify: `tests/test_file_knowledge.py` +- Modify: `tests/test_graph_engine.py` + +**Interfaces:** +- Produces: `upsert_file_asset()`, `get_file_asset()`, four `/file-assets` API routes. +- Consumes: existing `GraphStore`, `policy.evaluate`, graph traversal, semantic search. + +- [ ] **Step 1: Write failing graph/API tests** + +```python +def test_upsert_file_asset_merges_same_content_distributions_and_projects_assertions(): + store = InMemoryGraphStore() + first = sample_asset(locator="file:///a.xlsx") + second = sample_asset(locator="file:///copy.xlsx") + upsert_file_asset(store, first) + merged = upsert_file_asset(store, second) + assert len(merged.distributions) == 2 + graph = store.traverse(first.asset_id, direction="out", max_depth=1) + assert {edge["edge_type"] for edge in graph["edges"]} >= {"DISTRIBUTION", "HAS_TOPIC"} + +def test_file_asset_api_requires_policy_and_redacts_jsonld_locator(client): + assert client.post("/file-assets", json=sample_request(actor="analyst")).status_code == 403 + created = client.post("/file-assets", json=sample_request(actor="admin")) + assert created.status_code == 200 + exported = client.get(f"/file-assets/{created.json()['asset_id']}/jsonld", params={"actor": "analyst"}) + assert "dcat:accessURL" not in exported.text +``` + +- [ ] **Step 2: Run tests and confirm route/function failures** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py tests/test_graph_engine.py -k file_asset -q` + +Expected: FAIL because projection and routes do not exist. + +- [ ] **Step 3: Implement graph projection and retrieval** + +`upsert_file_asset(store, asset)` validates first, merges existing distributions by id, writes the file node with serializable metadata only, writes distribution and concept/reference nodes, and creates allowlisted uppercase edge labels with `confidence`, `review_status`, and evidence-reference properties. Embedding text is only title plus target labels. `get_file_asset()` reconstructs the Pydantic model from the file node. + +- [ ] **Step 4: Add routes with existing authorization** + +`POST /file-assets` calls `_authorize_graph_write`; GET/detail/JSON-LD/validate call `_authorize_graph_read`. Locator inclusion requires `include_locations=true` and an admin `create` policy check. Convert `KeyError` to 404, validation `ValueError` to 400, and authorization failure to 403. + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py tests/test_graph_engine.py -q` + +Expected: PASS. + +- [ ] **Step 5: Commit** + +```bash +git add src/sdp/file_ontology.py src/sdp/api.py tests/test_file_knowledge.py tests/test_graph_engine.py +git commit -m "feat: project file knowledge into graph API" +``` + +### Task 6: Local pilot runner and Hyosung VOC preflight + +**Files:** +- Create: `src/sdp/file_pilot.py` +- Modify: `tests/test_file_knowledge.py` + +**Interfaces:** +- Produces: `run_local_pilot()`, `python -m sdp.file_pilot`. +- Consumes: filesystem reader, document extraction, optional OpenAI extractor, graph projection, JSON-LD export. + +- [ ] **Step 1: Write a failing end-to-end local runner test** + +```python +def test_local_pilot_deduplicates_content_and_writes_no_raw_text(tmp_path): + root = tmp_path / "input" + root.mkdir() + (root / "효성중공업 VOC.txt").write_text("C-Cube PoC", encoding="utf-8") + (root / "중공업VOC copy.txt").write_text("C-Cube PoC", encoding="utf-8") + output = tmp_path / "manifest.json" + summary = run_local_pilot(root, output, name_pattern=r"효성중공업|중공업VOC", extractor=FakeExtractor()) + assert summary["files"] == 2 and summary["assets"] == 1 and summary["distributions"] == 2 + assert "C-Cube PoC" not in output.read_text(encoding="utf-8") +``` + +- [ ] **Step 2: Run the test and confirm failure** + +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k local_pilot -q` + +Expected: FAIL because the runner does not exist. + +- [ ] **Step 3: Implement the runner and CLI** + +`run_local_pilot()` groups by SHA-256, merges distributions, extracts/chunks each unique asset once, applies the optional extractor, validates and projects into an `InMemoryGraphStore`, and writes a UTF-8 JSON manifest containing summary, asset JSON-LD with locators, per-file status, and no raw text/quote/API response. CLI arguments are `--root`, `--output`, `--name-regex`, `--model`, `--max-files`, and `--no-llm`; with LLM enabled it obtains the API key via `getpass.getpass()` and an `EphemeralCredentialRegistry`, never a CLI argument or environment variable. + +- [ ] **Step 4: Run tests and the 12-file no-LLM preflight** + +Run: + +```powershell +$env:PYTHONPATH='src' +py -m pytest tests/test_file_knowledge.py -k local_pilot -q +py -m sdp.file_pilot --root 'D:\SynologyDrive\업무자료\Download_정리_2026-07-15\01_문서' --output 'C:\Users\Seongho Bae\Documents\Codex\2026-07-15\plugin-computer-use-openai-bundled-ponytail\outputs\hyosung-voc-file-index.json' --name-regex '효성중공업|중공업VOC' --max-files 12 --no-llm +``` + +Expected: 12 files, 10 assets, 12 distributions; DOCX/XLSX/PPTX/PDF extraction statuses are `extracted` or explicitly `needs_ocr`. + +- [ ] **Step 5: Commit** + +```bash +git add src/sdp/file_pilot.py tests/test_file_knowledge.py +git commit -m "feat: add safe local semantic file pilot" +``` + +### Task 7: Documentation, Codegraph, full verification, and live pilot + +**Files:** +- Modify: `README.md` +- Modify: `docs/implementation-compliance.md` +- Modify: `docs/superpowers/plans/2026-07-21-semantic-file-ontology.md` +- Local-only output: `C:/Users/Seongho Bae/Documents/Codex/2026-07-15/plugin-computer-use-openai-bundled-ponytail/outputs/hyosung-voc-file-index.json` + +**Interfaces:** +- Produces: documented API/CLI, verified local manifest, current Codegraph index. +- Consumes: all earlier tasks. + +- [ ] **Step 1: Document exact operation and security boundary** + +Add the four API endpoints, standards profile paths, supported providers/formats, injected-client pattern, interactive key prompt, local pilot command, output privacy warning, and exclusions to README. Add requirement-to-proof rows and test names to the implementation compliance matrix. + +- [ ] **Step 2: Run full tests and smoke** + +Run: + +```powershell +$env:PYTHONPATH='src' +py -m pytest -q +py -m sdp.demo_smoke +git diff --check +``` + +Expected: pytest passes with only declared integration skips; demo smoke exits 0; diff check is clean. + +- [ ] **Step 3: Sync Codegraph and inspect impact** + +Run: + +```powershell +codegraph sync +codegraph status +codegraph affected src/sdp/file_ontology.py src/sdp/storage_readers.py src/sdp/document_semantics.py src/sdp/file_pilot.py src/sdp/api.py +``` + +Expected: index is current and affected tests include the new file knowledge tests plus API/graph tests. + +- [ ] **Step 4: Run the live OpenAI pilot after secure key entry** + +Run the Task 6 pilot command without `--no-llm`, enter the project-scoped key only at the hidden prompt, and verify the output reports proposed assertions with evidence hashes/offsets and no raw chunks. If no key is available, stop at the verified no-LLM manifest and request secure local entry; do not weaken credential handling. + +- [ ] **Step 5: Commit final docs and verification evidence** + +```bash +git add README.md docs/implementation-compliance.md docs/superpowers/plans/2026-07-21-semantic-file-ontology.md +git commit -m "docs: document semantic file catalog pilot" +git status --short +``` + +Expected: clean worktree. + +## Self-Review + +- Spec coverage: standards profile, provider-neutral identity/location split, four providers, safe content extraction, OpenAI evidence binding, SHACL-compatible validation, graph/API reuse, privacy, and 12-file pilot each have a task. +- Placeholder scan: no TBD/TODO/future implementation placeholder is used; exclusions are explicit scope decisions. +- Type consistency: Tasks 2–7 consume the exact `StorageDistribution`, `SemanticAssertion`, `FileAsset`, `ObjectRef`, `TextChunk`, and extractor names produced in earlier tasks. +- Dependency scope: only pypdf is added because the approved pilot includes two PDFs; cloud SDKs and the OpenAI SDK remain injected/stdlib. From 4b911a2a4bfbd5470b7366d5d110b279d851298c Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 10:44:37 +0900 Subject: [PATCH 03/15] feat: define standards-aligned file ontology --- ontology/cwl-file-profile.ttl | 31 +++++ ontology/cwl-file-shapes.ttl | 24 ++++ src/sdp/file_ontology.py | 227 ++++++++++++++++++++++++++++++++++ tests/test_file_knowledge.py | 115 +++++++++++++++++ 4 files changed, 397 insertions(+) create mode 100644 ontology/cwl-file-profile.ttl create mode 100644 ontology/cwl-file-shapes.ttl create mode 100644 src/sdp/file_ontology.py create mode 100644 tests/test_file_knowledge.py diff --git a/ontology/cwl-file-profile.ttl b/ontology/cwl-file-profile.ttl new file mode 100644 index 0000000..85f6c86 --- /dev/null +++ b/ontology/cwl-file-profile.ttl @@ -0,0 +1,31 @@ +@prefix cwl: . +@prefix dcat: . +@prefix owl: . +@prefix prov: . +@prefix skos: . + +cwl: a owl:Ontology ; + owl:versionInfo "0.1.0" . + +cwl:FileAsset a owl:Class ; + owl:equivalentClass [ a owl:Class ; owl:intersectionOf ( dcat:Resource prov:Entity ) ] . + +cwl:BusinessProject a owl:Class ; owl:subClassOf skos:Concept . +cwl:System a owl:Class ; owl:subClassOf skos:Concept . +cwl:WorkPhase a owl:Class ; owl:subClassOf skos:Concept . +cwl:ArtifactType a owl:Class ; owl:subClassOf skos:Concept . +cwl:SemanticAssertion a owl:Class . + +cwl:belongsToProject a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:BusinessProject . +cwl:usesSystem a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:System . +cwl:hasWorkPhase a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:WorkPhase . +cwl:hasArtifactType a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:ArtifactType . +cwl:hasTopic a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range skos:Concept . +cwl:wasDerivedFrom a owl:ObjectProperty ; + owl:subPropertyOf prov:wasDerivedFrom . +cwl:previousVersion a owl:ObjectProperty . diff --git a/ontology/cwl-file-shapes.ttl b/ontology/cwl-file-shapes.ttl new file mode 100644 index 0000000..bd6d4ed --- /dev/null +++ b/ontology/cwl-file-shapes.ttl @@ -0,0 +1,24 @@ +@prefix cwl: . +@prefix dcat: . +@prefix dcterms: . +@prefix sh: . +@prefix spdx: . + +cwl:FileAssetShape a sh:NodeShape ; + sh:targetClass cwl:FileAsset ; + sh:property [ sh:path dcterms:title ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path spdx:checksum ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path dcat:distribution ; sh:minCount 1 ] . + +cwl:DistributionShape a sh:NodeShape ; + sh:targetClass dcat:Distribution ; + sh:property [ sh:path cwl:storageProvider ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path cwl:endpointId ; sh:minCount 1 ; sh:maxCount 1 ] . + +cwl:SemanticAssertionShape a sh:NodeShape ; + sh:targetClass cwl:SemanticAssertion ; + sh:property [ sh:path cwl:confidence ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path cwl:reviewStatus ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path cwl:evidenceChunkSha256 ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path cwl:evidenceStart ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path cwl:evidenceEnd ; sh:minCount 1 ; sh:maxCount 1 ] . diff --git a/src/sdp/file_ontology.py b/src/sdp/file_ontology.py new file mode 100644 index 0000000..5dc8386 --- /dev/null +++ b/src/sdp/file_ontology.py @@ -0,0 +1,227 @@ +"""Standards-aligned contracts for content-addressed file knowledge.""" + +from __future__ import annotations + +import hashlib +import re +from datetime import datetime +from typing import Any, Literal +from urllib.parse import urlsplit + +from pydantic import BaseModel, Field, field_validator + + +StorageProvider = Literal["filesystem", "s3", "s3_compatible", "azure_blob"] +AssertionRelation = Literal[ + "belongsToProject", + "usesSystem", + "hasWorkPhase", + "hasArtifactType", + "hasTopic", + "wasDerivedFrom", + "previousVersion", +] +TargetKind = Literal[ + "business_project", + "system", + "work_phase", + "artifact_type", + "topic", + "file_asset", +] + +_SHA256_RE = re.compile(r"^[0-9a-f]{64}$") +_RELATION_TARGETS: dict[str, str] = { + "belongsToProject": "business_project", + "usesSystem": "system", + "hasWorkPhase": "work_phase", + "hasArtifactType": "artifact_type", + "hasTopic": "topic", + "wasDerivedFrom": "file_asset", + "previousVersion": "file_asset", +} +_CONTEXT = { + "cwl": "https://contextualwisdomlab.github.io/semantic-data-portal/ontology/file#", + "dcat": "http://www.w3.org/ns/dcat#", + "dcterms": "http://purl.org/dc/terms/", + "prov": "http://www.w3.org/ns/prov#", + "skos": "http://www.w3.org/2004/02/skos/core#", + "spdx": "http://spdx.org/rdf/terms#", +} + + +class StorageDistribution(BaseModel): + id: str = Field(min_length=1) + provider: StorageProvider + locator: str = Field(min_length=1) + endpoint_id: str = Field(min_length=1) + available: bool = True + bucket: str | None = None + container: str | None = None + object_key: str | None = None + version_id: str | None = None + etag: str | None = None + + @field_validator("locator") + @classmethod + def stable_locator_without_credentials(cls, value: str) -> str: + parsed = urlsplit(value) + if not parsed.scheme: + raise ValueError("locator must be an absolute IRI") + if parsed.query or parsed.fragment: + raise ValueError("locator must not contain a query or fragment") + return value + + +class SemanticAssertion(BaseModel): + relation: AssertionRelation + target_kind: TargetKind + target_label: str = Field(min_length=1) + confidence: float = Field(ge=0, le=1) + evidence_chunk_sha256: str + evidence_start: int + evidence_end: int + method: str = "openai" + review_status: Literal["proposed", "approved", "rejected"] = "proposed" + + +class FileAsset(BaseModel): + sha256: str + title: str = Field(min_length=1) + media_type: str = Field(min_length=1) + byte_size: int = Field(ge=0) + modified_at: datetime | None = None + distributions: list[StorageDistribution] + assertions: list[SemanticAssertion] = Field(default_factory=list) + + @property + def asset_id(self) -> str: + return f"urn:sha256:{self.sha256}" + + +def _violation(path: str, message: str) -> dict[str, str]: + return { + "shape": "CWLFileAssetShape", + "path": path, + "severity": "violation", + "message": message, + } + + +def validate_file_asset(asset: FileAsset) -> dict[str, Any]: + """Return the SHACL-compatible structural validation report for an asset.""" + + violations: list[dict[str, str]] = [] + if not _SHA256_RE.fullmatch(asset.sha256): + violations.append(_violation("sha256", "must be a lowercase SHA-256 hex digest")) + if not asset.distributions: + violations.append(_violation("distributions", "at least one distribution is required")) + + for distribution in asset.distributions: + if distribution.provider in {"s3", "s3_compatible"} and not distribution.bucket: + violations.append(_violation("distributions.bucket", "S3 distributions require a bucket")) + if distribution.provider == "azure_blob" and not distribution.container: + violations.append( + _violation("distributions.container", "Azure Blob distributions require a container") + ) + if distribution.provider != "filesystem" and not distribution.object_key: + violations.append( + _violation("distributions.object_key", "object storage distributions require an object key") + ) + + for assertion in asset.assertions: + expected_kind = _RELATION_TARGETS[assertion.relation] + if assertion.target_kind != expected_kind: + violations.append( + _violation( + "assertions.target_kind", + f"{assertion.relation} requires target_kind={expected_kind}", + ) + ) + if not _SHA256_RE.fullmatch(assertion.evidence_chunk_sha256): + violations.append( + _violation("assertions.evidence_chunk_sha256", "must be a lowercase SHA-256 hex digest") + ) + if assertion.evidence_start < 0: + violations.append(_violation("assertions.evidence_start", "must be non-negative")) + if assertion.evidence_end <= assertion.evidence_start: + violations.append( + _violation("assertions.evidence_end", "must be greater than evidence_start") + ) + + return { + "asset_id": asset.asset_id, + "shacl_compatible": True, + "shape": "CWLFileAssetShape", + "conforms": not violations, + "violations": violations, + } + + +def concept_id(kind: str, label: str) -> str: + digest = hashlib.sha256(label.strip().casefold().encode("utf-8")).hexdigest()[:24] + return f"urn:cwl:{kind}:{digest}" + + +def file_asset_jsonld(asset: FileAsset, *, include_locations: bool = False) -> dict[str, Any]: + """Export one file asset using the CWL application profile.""" + + distributions: list[dict[str, Any]] = [] + for distribution in asset.distributions: + item: dict[str, Any] = { + "@id": f"urn:cwl:distribution:{distribution.id}", + "@type": "dcat:Distribution", + "cwl:storageProvider": distribution.provider, + "cwl:endpointId": distribution.endpoint_id, + "cwl:available": distribution.available, + } + if include_locations: + item["dcat:accessURL"] = distribution.locator + if distribution.version_id: + item["dcat:version"] = distribution.version_id + if distribution.etag: + item["cwl:etag"] = distribution.etag + distributions.append(item) + + subjects = [ + { + "@id": concept_id(assertion.target_kind, assertion.target_label), + "@type": "skos:Concept", + "skos:prefLabel": assertion.target_label, + } + for assertion in asset.assertions + ] + assertions = [ + { + "@type": "cwl:SemanticAssertion", + "cwl:relation": f"cwl:{assertion.relation}", + "cwl:target": {"@id": concept_id(assertion.target_kind, assertion.target_label)}, + "cwl:confidence": assertion.confidence, + "cwl:evidenceChunkSha256": assertion.evidence_chunk_sha256, + "cwl:evidenceStart": assertion.evidence_start, + "cwl:evidenceEnd": assertion.evidence_end, + "cwl:extractionMethod": assertion.method, + "cwl:reviewStatus": assertion.review_status, + } + for assertion in asset.assertions + ] + payload: dict[str, Any] = { + "@context": _CONTEXT, + "@id": asset.asset_id, + "@type": ["dcat:Resource", "prov:Entity", "cwl:FileAsset"], + "dcterms:identifier": asset.asset_id, + "dcterms:title": asset.title, + "dcterms:format": asset.media_type, + "dcat:byteSize": asset.byte_size, + "spdx:checksum": { + "@type": "spdx:Checksum", + "spdx:algorithm": "spdx:checksumAlgorithm_sha256", + "spdx:checksumValue": asset.sha256, + }, + "dcat:distribution": distributions, + "dcterms:subject": subjects, + "cwl:assertion": assertions, + } + if asset.modified_at: + payload["dcterms:modified"] = asset.modified_at.isoformat() + return payload diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py new file mode 100644 index 0000000..b9f4f69 --- /dev/null +++ b/tests/test_file_knowledge.py @@ -0,0 +1,115 @@ +from __future__ import annotations + +from copy import deepcopy + +import pytest + +from sdp.file_ontology import ( + FileAsset, + SemanticAssertion, + StorageDistribution, + file_asset_jsonld, + validate_file_asset, +) + + +ASSET_SHA256 = "a" * 64 +CHUNK_SHA256 = "b" * 64 + + +def sample_distribution(**overrides: object) -> StorageDistribution: + values: dict[str, object] = { + "id": "dist-local", + "provider": "filesystem", + "locator": "file:///D:/Documents/report.docx", + "endpoint_id": "windows-local", + } + values.update(overrides) + return StorageDistribution(**values) + + +def sample_assertion(**overrides: object) -> SemanticAssertion: + values: dict[str, object] = { + "relation": "usesSystem", + "target_kind": "system", + "target_label": "C-Cube", + "confidence": 0.91, + "evidence_chunk_sha256": CHUNK_SHA256, + "evidence_start": 3, + "evidence_end": 9, + } + values.update(overrides) + return SemanticAssertion(**values) + + +def sample_asset(**overrides: object) -> FileAsset: + values: dict[str, object] = { + "sha256": ASSET_SHA256, + "title": "효성중공업 VOC 종료보고서", + "media_type": "application/vnd.openxmlformats-officedocument.wordprocessingml.document", + "byte_size": 1024, + "distributions": [sample_distribution()], + "assertions": [sample_assertion()], + } + values.update(overrides) + return FileAsset(**values) + + +def test_file_asset_jsonld_uses_standards_and_redacts_locator(): + asset = sample_asset() + + payload = file_asset_jsonld(asset) + + assert payload["@type"] == ["dcat:Resource", "prov:Entity", "cwl:FileAsset"] + assert "dcat:accessURL" not in payload["dcat:distribution"][0] + assert payload["spdx:checksum"]["spdx:checksumValue"] == asset.sha256 + assert payload["dcterms:subject"][0]["skos:prefLabel"] == "C-Cube" + + +def test_file_asset_jsonld_can_include_authorized_locations(): + payload = file_asset_jsonld(sample_asset(), include_locations=True) + + assert payload["dcat:distribution"][0]["dcat:accessURL"] == "file:///D:/Documents/report.docx" + + +def test_file_asset_validation_rejects_unverifiable_evidence(): + raw = sample_assertion().model_dump() + raw["evidence_chunk_sha256"] = "" + raw["evidence_start"] = 4 + raw["evidence_end"] = 2 + asset = sample_asset(assertions=[raw]) + + report = validate_file_asset(asset) + + assert report["conforms"] is False + paths = {violation["path"] for violation in report["violations"]} + assert {"assertions.evidence_chunk_sha256", "assertions.evidence_end"} <= paths + + +@pytest.mark.parametrize( + ("provider", "overrides", "expected_path"), + [ + ("s3", {"bucket": None, "object_key": "report.pdf"}, "distributions.bucket"), + ("s3_compatible", {"bucket": "voc", "object_key": None}, "distributions.object_key"), + ("azure_blob", {"container": None, "object_key": "report.pdf"}, "distributions.container"), + ], +) +def test_file_asset_validation_checks_provider_coordinates(provider, overrides, expected_path): + values = deepcopy(sample_distribution().model_dump()) + values.update( + provider=provider, + locator="https://objects.example/voc/report.pdf", + endpoint_id="objects", + **overrides, + ) + asset = sample_asset(distributions=[values]) + + report = validate_file_asset(asset) + + assert report["conforms"] is False + assert expected_path in {violation["path"] for violation in report["violations"]} + + +def test_distribution_rejects_locator_query_to_avoid_secret_leak(): + with pytest.raises(ValueError, match="query or fragment"): + sample_distribution(locator="https://objects.example/report.pdf?sig=secret") From e3ab22fd0398a45962ac100a958c432b6c9e089b Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 10:47:04 +0900 Subject: [PATCH 04/15] feat: add provider-neutral read-only storage readers --- src/sdp/storage_readers.py | 231 +++++++++++++++++++++++++++++++++++ tests/test_file_knowledge.py | 98 +++++++++++++++ 2 files changed, 329 insertions(+) create mode 100644 src/sdp/storage_readers.py diff --git a/src/sdp/storage_readers.py b/src/sdp/storage_readers.py new file mode 100644 index 0000000..1d95ab4 --- /dev/null +++ b/src/sdp/storage_readers.py @@ -0,0 +1,231 @@ +"""Read-only object readers for local and injected cloud storage clients.""" + +from __future__ import annotations + +import hashlib +import os +import re +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Iterable, Protocol +from urllib.parse import quote, urlsplit + +from .file_ontology import StorageDistribution, StorageProvider + + +def _stable_id(provider: str, locator: str) -> str: + digest = hashlib.sha256(f"{provider}\0{locator}".encode("utf-8")).hexdigest()[:24] + return f"{provider}-{digest}" + + +def _is_link_or_junction(path: Path) -> bool: + is_junction = getattr(os.path, "isjunction", lambda _: False) + return path.is_symlink() or bool(is_junction(path)) + + +@dataclass(frozen=True) +class ObjectRef: + name: str + object_key: str + size: int + modified_at: datetime | None + distribution: StorageDistribution + + +class ObjectReader(Protocol): + def list( + self, prefix: str = "", *, name_pattern: str | None = None + ) -> Iterable[ObjectRef]: ... + + def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: ... + + +class FilesystemReader: + def __init__(self, root: str | Path) -> None: + self.root = Path(root).resolve(strict=True) + if not self.root.is_dir(): + raise ValueError("filesystem root must be a directory") + self.endpoint_id = f"filesystem-{hashlib.sha256(str(self.root).casefold().encode()).hexdigest()[:16]}" + + def _inside_root(self, path: Path) -> Path: + resolved = path.resolve(strict=True) + try: + resolved.relative_to(self.root) + except ValueError as exc: + raise ValueError("path escapes configured filesystem root") from exc + return resolved + + def list( + self, prefix: str = "", *, name_pattern: str | None = None + ) -> Iterable[ObjectRef]: + start = self._inside_root(self.root / prefix) + pattern = re.compile(name_pattern, re.IGNORECASE) if name_pattern else None + for current, directories, filenames in os.walk(start, followlinks=False): + current_path = Path(current) + directories[:] = [ + name + for name in directories + if not _is_link_or_junction(current_path / name) + ] + for name in filenames: + if name.startswith("~$") or (pattern and not pattern.search(name)): + continue + source = current_path / name + if _is_link_or_junction(source): + continue + resolved = self._inside_root(source) + stat = resolved.stat() + locator = resolved.as_uri() + relative = resolved.relative_to(self.root).as_posix() + yield ObjectRef( + name=name, + object_key=relative, + size=stat.st_size, + modified_at=datetime.fromtimestamp(stat.st_mtime, timezone.utc), + distribution=StorageDistribution( + id=_stable_id("filesystem", locator), + provider="filesystem", + locator=locator, + endpoint_id=self.endpoint_id, + ), + ) + + def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: + if ref.distribution.provider != "filesystem": + raise ValueError("object reference does not belong to the filesystem reader") + source = self._inside_root(self.root / Path(ref.object_key)) + if _is_link_or_junction(source): + raise ValueError("symbolic links and junctions are not readable") + if source.stat().st_size > max_bytes: + raise ValueError("object exceeds maximum size") + with source.open("rb") as handle: + data = handle.read(max_bytes + 1) + if len(data) > max_bytes: + raise ValueError("object exceeds maximum size") + return data + + +class S3Reader: + def __init__( + self, + client: Any, + bucket: str, + *, + provider: StorageProvider = "s3", + endpoint_url: str | None = None, + ) -> None: + if provider not in {"s3", "s3_compatible"}: + raise ValueError("S3Reader provider must be s3 or s3_compatible") + if provider == "s3_compatible" and not endpoint_url: + raise ValueError("S3-compatible storage requires endpoint_url") + self.client = client + self.bucket = bucket + self.provider = provider + self.endpoint_url = endpoint_url.rstrip("/") if endpoint_url else None + host = urlsplit(self.endpoint_url).netloc if self.endpoint_url else "aws" + self.endpoint_id = f"{provider}-{host}" + + def _locator(self, key: str) -> str: + encoded_key = quote(key, safe="/") + if self.provider == "s3": + return f"s3://{quote(self.bucket, safe='')}/{encoded_key}" + return f"{self.endpoint_url}/{quote(self.bucket, safe='')}/{encoded_key}" + + def list( + self, prefix: str = "", *, name_pattern: str | None = None + ) -> Iterable[ObjectRef]: + pattern = re.compile(name_pattern, re.IGNORECASE) if name_pattern else None + paginator = self.client.get_paginator("list_objects_v2") + for page in paginator.paginate(Bucket=self.bucket, Prefix=prefix): + for item in page.get("Contents", []): + key = str(item["Key"]) + name = key.rsplit("/", 1)[-1] + if not name or name.startswith("~$") or (pattern and not pattern.search(name)): + continue + locator = self._locator(key) + etag = str(item.get("ETag", "")).strip('"') or None + yield ObjectRef( + name=name, + object_key=key, + size=int(item.get("Size", 0)), + modified_at=item.get("LastModified"), + distribution=StorageDistribution( + id=_stable_id(self.provider, locator), + provider=self.provider, + locator=locator, + endpoint_id=self.endpoint_id, + bucket=self.bucket, + object_key=key, + etag=etag, + ), + ) + + def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: + if ref.distribution.provider != self.provider: + raise ValueError("object reference does not belong to this S3 reader") + if ref.size > max_bytes: + raise ValueError("object exceeds maximum size") + request: dict[str, str] = {"Bucket": self.bucket, "Key": ref.object_key} + if ref.distribution.version_id: + request["VersionId"] = ref.distribution.version_id + body = self.client.get_object(**request)["Body"] + try: + data = body.read(max_bytes + 1) + finally: + close = getattr(body, "close", None) + if close: + close() + if len(data) > max_bytes: + raise ValueError("object exceeds maximum size") + return data + + +class AzureBlobReader: + def __init__(self, container_client: Any) -> None: + self.client = container_client + self.container = str(container_client.container_name) + self.container_url = str(container_client.url).rstrip("/") + self.endpoint_id = f"azure-blob-{urlsplit(self.container_url).netloc}" + + def list( + self, prefix: str = "", *, name_pattern: str | None = None + ) -> Iterable[ObjectRef]: + pattern = re.compile(name_pattern, re.IGNORECASE) if name_pattern else None + for item in self.client.list_blobs(name_starts_with=prefix): + key = str(item.name) + name = key.rsplit("/", 1)[-1] + if not name or name.startswith("~$") or (pattern and not pattern.search(name)): + continue + locator = f"{self.container_url}/{quote(key, safe='/')}" + yield ObjectRef( + name=name, + object_key=key, + size=int(item.size), + modified_at=item.last_modified, + distribution=StorageDistribution( + id=_stable_id("azure_blob", locator), + provider="azure_blob", + locator=locator, + endpoint_id=self.endpoint_id, + container=self.container, + object_key=key, + version_id=getattr(item, "version_id", None), + etag=str(item.etag).strip('"') if item.etag else None, + ), + ) + + def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: + if ref.distribution.provider != "azure_blob": + raise ValueError("object reference does not belong to the Azure Blob reader") + if ref.size > max_bytes: + raise ValueError("object exceeds maximum size") + kwargs = ( + {"version_id": ref.distribution.version_id} + if ref.distribution.version_id + else {} + ) + data = self.client.get_blob_client(ref.object_key, **kwargs).download_blob().readall() + if len(data) > max_bytes: + raise ValueError("object exceeds maximum size") + return data diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index b9f4f69..1940a84 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -1,6 +1,8 @@ from __future__ import annotations from copy import deepcopy +from io import BytesIO +from types import SimpleNamespace import pytest @@ -11,6 +13,7 @@ file_asset_jsonld, validate_file_asset, ) +from sdp.storage_readers import AzureBlobReader, FilesystemReader, S3Reader ASSET_SHA256 = "a" * 64 @@ -113,3 +116,98 @@ def test_file_asset_validation_checks_provider_coordinates(provider, overrides, def test_distribution_rejects_locator_query_to_avoid_secret_leak(): with pytest.raises(ValueError, match="query or fragment"): sample_distribution(locator="https://objects.example/report.pdf?sig=secret") + + +def test_filesystem_reader_stays_inside_root_and_reads_bytes(tmp_path): + path = tmp_path / "효성중공업 VOC.txt" + path.write_text("C-Cube PoC", encoding="utf-8") + (tmp_path / "~$효성중공업 VOC.docx").write_bytes(b"temporary") + reader = FilesystemReader(tmp_path) + + refs = list(reader.list(name_pattern=r"효성중공업|중공업VOC")) + + assert len(refs) == 1 + assert reader.read(refs[0], max_bytes=1024) == "C-Cube PoC".encode() + assert refs[0].distribution.provider == "filesystem" + assert refs[0].distribution.locator == path.resolve().as_uri() + + +def test_filesystem_reader_rejects_oversized_object(tmp_path): + path = tmp_path / "중공업VOC.txt" + path.write_bytes(b"12345") + reader = FilesystemReader(tmp_path) + ref = next(reader.list()) + + with pytest.raises(ValueError, match="maximum size"): + reader.read(ref, max_bytes=4) + + +class FakeS3: + def get_paginator(self, operation): + assert operation == "list_objects_v2" + return SimpleNamespace( + paginate=lambda **kwargs: [ + { + "Contents": [ + { + "Key": f"{kwargs['Prefix']}a.docx", + "Size": 7, + "ETag": '"etag-a"', + } + ] + } + ] + ) + + def get_object(self, **kwargs): + assert kwargs == {"Bucket": "voc", "Key": "reports/a.docx"} + return {"Body": BytesIO(b"content")} + + +class FakeContainer: + url = "https://account.blob.core.windows.net/voc" + container_name = "voc" + + def list_blobs(self, *, name_starts_with): + return [ + SimpleNamespace( + name=f"{name_starts_with}a.docx", + size=7, + etag="etag-a", + version_id="version-a", + last_modified=None, + ) + ] + + def get_blob_client(self, name, **kwargs): + assert name == "reports/a.docx" + assert kwargs == {"version_id": "version-a"} + return SimpleNamespace(download_blob=lambda: SimpleNamespace(readall=lambda: b"content")) + + +def test_s3_and_azure_readers_use_injected_clients(): + s3 = S3Reader(FakeS3(), "voc") + azure = AzureBlobReader(FakeContainer()) + + s3_ref = next(s3.list("reports/")) + azure_ref = next(azure.list("reports/")) + + assert s3_ref.object_key == azure_ref.object_key == "reports/a.docx" + assert s3.read(s3_ref, max_bytes=20) == b"content" + assert azure.read(azure_ref, max_bytes=20) == b"content" + assert s3_ref.distribution.locator == "s3://voc/reports/a.docx" + assert azure_ref.distribution.container == "voc" + + +def test_s3_compatible_reader_uses_stable_endpoint_without_credentials(): + reader = S3Reader( + FakeS3(), + "voc", + provider="s3_compatible", + endpoint_url="https://objects.example", + ) + + ref = next(reader.list("reports/")) + + assert ref.distribution.provider == "s3_compatible" + assert ref.distribution.locator == "https://objects.example/voc/reports/a.docx" From 5f0a4fe843ac4f1199b678564164d97727b28fbe Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 10:50:02 +0900 Subject: [PATCH 05/15] feat: extract text from pilot document formats --- pyproject.toml | 1 + requirements-dev.txt | 210 ++++++++++++++++++++++++++++++++++ requirements-test.in | 1 + requirements-test.txt | 21 ++-- requirements.txt | 4 + src/sdp/document_semantics.py | 141 +++++++++++++++++++++++ tests/test_file_knowledge.py | 67 +++++++++++ 7 files changed, 435 insertions(+), 10 deletions(-) create mode 100644 src/sdp/document_semantics.py diff --git a/pyproject.toml b/pyproject.toml index a5f9a30..b5280cd 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -13,6 +13,7 @@ dependencies = [ "psycopg[binary]==3.3.4", "PyJWT[crypto]==2.13.0", "pydantic==2.13.4", + "pypdf==6.14.2", "uvicorn==0.51.0", ] diff --git a/requirements-dev.txt b/requirements-dev.txt index 26d2308..fbbc51d 100644 --- a/requirements-dev.txt +++ b/requirements-dev.txt @@ -184,6 +184,87 @@ fastapi==0.139.0 \ --hash=sha256:99ab7b2d92223c76d6cf10757ab3f89d45b38267fc20b2a136cf02f6beac3145 \ --hash=sha256:cf15e1e9e667ddb0ad63811e60bd11390d1aac838ca4a7a23f421807b2308189 # via semantic-data-portal (pyproject.toml) +greenlet==3.5.3 \ + --hash=sha256:0909f9355a9f24845d3299f3112e266a06afb68302041989fd26bd68894933db \ + --hash=sha256:0f41e4a05a3c0cb31b17023eff28dd111e1d16bf7d7d00406cd7df23f31398a7 \ + --hash=sha256:0f6ff50ff8dbd51fae9b37f4101648b04ea0df19b3f50ab2beb5061e7716a5c8 \ + --hash=sha256:0f71be4920368fe1fabeeaa53d1e3548337e2b223d9565f8ad5e392a75ba23fc \ + --hash=sha256:12a248ba75f6a9a236375f52296c498c89ff1d8badf32deb9eca7abd5853f7da \ + --hash=sha256:1540dd8e5fc2a5aec40fbb98ef8e149fa47c89a4b4a1cf2575a14d3d1869d7a8 \ + --hash=sha256:16d192579ed281051396dddd7f7754dac6259e6b1fb26378c87b66622f8e3f91 \ + --hash=sha256:176bc16a721fa5fc294d70b87b4dfa5fbdd251b3da5d5372735ecef9bd7d6d0c \ + --hash=sha256:19131729ae0ddc3c2e1ef85e650169b5e37ee32e400f215f78b94d7b0d567310 \ + --hash=sha256:1c514a468149bf8fbbab874188a3535cd8a48a3e353eb53a3d424296f8dbacd3 \ + --hash=sha256:1dae6e0091eae084317e411f047f0b7cb241c6db570f7c45fd6b900a274914ce \ + --hash=sha256:215275b1b49320987352e6c1b054acca0064f965a2c66992bed9a6f7d913f149 \ + --hash=sha256:232fec92e823addaf02d9472cf7381e24a1d046a6ced1103c5caa4c21b9dfc1d \ + --hash=sha256:2421c3564da9429d5586d46ca31ebb26516b5498a802cf65c041a8e8a8980d34 \ + --hash=sha256:271a8ea7c1024e8a0d7dd2be66dd66dda8a07193f41a17b9e924f7600f5b62be \ + --hash=sha256:2b2e857ae16f5f72142edf75f9f176fe7526ba19a2841df1420516f83831c9f2 \ + --hash=sha256:2ecda9ec22edf38fa389369eaed8c3d37c05f3c54e69f69438dbb2cc1de1458b \ + --hash=sha256:3236754d423955ea08e9bb5f6c04a7895f9e22c290b66aa7653fcb922d839eb0 \ + --hash=sha256:37bf9c538f5ae6e63d643f88dec37c0c83bdf0e2ebc62961dedcf458822f7b71 \ + --hash=sha256:4399eb8d041f20b68d943918bc55502a93d6fdc0a37c14da7881c04139acee9d \ + --hash=sha256:483d08c11181c83a6ce1a7a61df0f624a208ec40817a3bb2302714592eee4f04 \ + --hash=sha256:499fef2acede88c1864a57bb586b4bf533c81e1b82df7ab93451cdb47dfec227 \ + --hash=sha256:4b9d501b40e80b70e32323c799dd9b420a5577a9601469d362ae1ffb690f3a7c \ + --hash=sha256:4d77e67f65f98449e3fb83f795b5d0a8437aead2f874ca89c96576caf4be3af6 \ + --hash=sha256:5121af01cf911e70056c00d4b46d5e9b5d1415550038573d744138bacb59e6b8 \ + --hash=sha256:55cf4d777485d43110e47133cbba6d74a8885a87ec1227ef0267f9ee80c5aa21 \ + --hash=sha256:5795cd1101371140551c645f2d408b8d3c01a5a29cf8a9bce6e759c983682d23 \ + --hash=sha256:5b4807c4082c9d1b6d9eed56fcd041863e37f2228106eef24c30ca096e238605 \ + --hash=sha256:6219b6d04dbf6ba6084d77dc609e8473060dc55f759cbf626d512122781fa128 \ + --hash=sha256:629b614d2b786e89c50440e246f33eea78f58a962d0bdbbcc809e6d13605903f \ + --hash=sha256:6b1b0eed82364b0e32c4ea0f221452d33e6bb17ae094d9f72aed9851812747ea \ + --hash=sha256:6f73857adb8fee13fa56c172bd11262f888c0c648f9fea113e777bb2c7904a81 \ + --hash=sha256:719757059f5a53fd0dde23f78cffeafcdd97b21c850ddb7ca684a3c1a1f122e2 \ + --hash=sha256:73f152c895e09907e0dbe24f6c2db37beb085cd63db91c3825a0fcd0064124a8 \ + --hash=sha256:7669aa24cf2a1041d6f7899575b494a3ab4cf68bfcc8609b1dc0be7272db835e \ + --hash=sha256:766cfd421c13e450feb340cd472a3ed9957d438727b7b4593ad7c76c5d2b0deb \ + --hash=sha256:78dbef602fda6d97d957eb7937f70c9ce9e9527330347f8f6b6f9e554a9e7a47 \ + --hash=sha256:7ef56fe650f50575bf843acde967b9c567687f3c22340941a899b7bc56e956a8 \ + --hash=sha256:7faba15ac005376e02a0384504e0243be3370ce010296a44a820feb342b505ab \ + --hash=sha256:8540f1e6205bd13ca0ce685581037219ca54a1b41a0a15d228c6c9b8ad5903d7 \ + --hash=sha256:87142215824be6ac05e2e8e2786eec307ccbc27c36723c3881959df654af6861 \ + --hash=sha256:8bdb43e1a1d1873721acab2be99c5befd4d2044ddfd52e4d610801019880a702 \ + --hash=sha256:8d19fe6c39ebff9259f07bcc685d3290f8fa4ea2278e51dd0008e4d6b0f2d814 \ + --hash=sha256:8ff8bed3e3baa20a3ea261ce00526f1898ad4801d4886fd2220580ee0ad8fadf \ + --hash=sha256:915f887cf2682b66419b879423a2e072634aa7b7dce6f3ada4957cfced3f1e9a \ + --hash=sha256:962c5df2db8cb446da51edf1ca5296c389d93b99c9d8aa2ee4c7d0d8f1218260 \ + --hash=sha256:9ad04dd75458c6300b047c61b8639092433d205a25a14e310d6582a480efcca1 \ + --hash=sha256:9bcd2d72ccd70a1ec68ba6ef93e7fbb4420ef9997dabc7010d893bd4015e0bec \ + --hash=sha256:a1fad1d11e7d6aab184107baa8e4ece11ccba3ec9599cd7efa5ff4d70d43256a \ + --hash=sha256:a2d185dd1621757e70c3861cceffd5317ab4e7ed7eb09c82994828468527ade5 \ + --hash=sha256:a61efc018fd3eb317eeca31aba90ee9e7f26f22884a79b6c6ec715bf71bb62f1 \ + --hash=sha256:aca9b4ce85b152b5524ef7d88170efdff80dc0032aa8b75f9aaf7f3479ea95b4 \ + --hash=sha256:af4923b3096e26a36d7e9cf24ab88083a20f97d191e3b97f253731ce9b41b28c \ + --hash=sha256:afaabdd554cd7ae9bbb3ca070b0d7fdfd207dbf1d16865f7233837709d354bda \ + --hash=sha256:b363d46ed1ea431825fdb01471bb024fc08399bad1572a616e853c7684415adb \ + --hash=sha256:b7068bd09f761f3f5b4d214c2bed063186b2a86148c740b3873e3f56d79bac31 \ + --hash=sha256:b897d97759425953f69a9c0fac67f8fe333ec0ce7377ef186fb2b0c3ad5e354d \ + --hash=sha256:c180d22d325fb613956b443c3c6f4406eb70e6defc70d3974da2a7b59e06f48c \ + --hash=sha256:c4e7b79d83805475f0102008843f6eb45fd3bb0b2e88c774adab5fbaab27117d \ + --hash=sha256:c82304750f057167ff60d188df1d0cc1764ce9567eadf03e6a7443bcedd0b30b \ + --hash=sha256:c8d87c2134d871df96ecdea9cec7cbaab286dadab0f56476e57aaf9e8ac11550 \ + --hash=sha256:cde8adafa2365676f74a979744629589999093bc86e2484214f58e61df08902c \ + --hash=sha256:cefa9cef4b371f9844c6053db71f1138bc6807bab1578b0dae5149c1f1141357 \ + --hash=sha256:d27c0c653a60d9535f690226474a5cc1036a8b0d7b57504d1c4f89c44a07a80c \ + --hash=sha256:dc133a1569ee667b2a6ef56ce551084aeefd87a5acbc4736d336d1e2edc6cfc4 \ + --hash=sha256:dd99329bbc15ca78dcc583dba05d0b1b0bae01ab6c2174989f5aaee3e41ac930 \ + --hash=sha256:df0a0628d1597eb0897b62f55d1343f772405fd25f3b2a796c76874b0c2e22e8 \ + --hash=sha256:e0f0d160f0b2e558e6c75f7930967183255dc9735e5f5b8cae58ee09c9576d8b \ + --hash=sha256:e18619ba655ac05d78d80fc83cac4ba892bd6927b99e3b8237aee861aaacc8bb \ + --hash=sha256:e44da2f5bbdaabaf7d80b73dbb430c7035771e9f244e3c8b769715c9d8fa0a16 \ + --hash=sha256:e515757e2e36bcbf1fad09a46e1557e8b1ae1797d4b44d09da7deed88ad28608 \ + --hash=sha256:e81fa194a1d20967877bdf9c7794db2bc99063e5be36aee710c08f04c5bb087f \ + --hash=sha256:ea03f2f04367845d6b58eeed276e1e56e51f0b97d8ad5a88a7d20a91dc9056cc \ + --hash=sha256:ebd933a6adabc298bab47731a130fe6bfb888bd934eee37810f151159544540d \ + --hash=sha256:ec6f1af59f6b5f3fc9678e2ea062d8377d22ac644f7844cb7a292910cf12ff44 \ + --hash=sha256:efa9f765dd09f9d0cdac651ffdf631ee59ec5dc6ee7a73e0c012ba9c52fbdf5b \ + --hash=sha256:efc6bd60ea02e085862c74a3ef64b147ffc6f1a5ea7d9f26e7a939943f68c1e3 \ + --hash=sha256:fad5aec764399f1b5cc347ad250a59660f20c8f8888ea6bae1f93b769cce1154 \ + --hash=sha256:fd2e02fa07485778536a036222d616ab957b1d533f36b3ed98ce725d9c9d3117 + # via sqlalchemy h11==0.16.0 \ --hash=sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1 \ --hash=sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86 @@ -198,6 +279,66 @@ httpx==0.28.1 \ --hash=sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc \ --hash=sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad # via semantic-data-portal (pyproject.toml) +hypothesis==6.157.2 \ + --hash=sha256:0017db9eeb655ee959efc5ae1cc5cb9fa00189534194255e7f33095229ad907a \ + --hash=sha256:01e610327f824aeedf3fa0d418e4f3f3f1266367e71d92e719755410ee525575 \ + --hash=sha256:030631cfb8d464f87e33ee2bf15b8ac98902593204b569daad589ff8e8948930 \ + --hash=sha256:0cb1289aaa2f7f7abd5817000ceae1f71c7eb40887de95c028738ade40902e66 \ + --hash=sha256:0e24261a5d65caed070936e7cb92cde0b9e8af8911cfc5a69f6a52632d7b5258 \ + --hash=sha256:12ed6a34e2ed5f4fbc935e340396f772cbff15d253a3157487b2638bdc101a9c \ + --hash=sha256:19c103185eae022feb46ff8bedec3e9d60fff7522ca26cc0cb802fa74e9693b3 \ + --hash=sha256:1c269ed44b9691c78165398d4ec5ce97cdb8f2b48da0469d319b41d286f1c6d5 \ + --hash=sha256:1f546d0f3f3f102d97c329019457c9f593c6f6145308824930a245d652360d10 \ + --hash=sha256:221c04ffa0abbbea43922a24e7d25df60501684eeedf61c332cceef70f10adb3 \ + --hash=sha256:2ba9758b5d7198e1bb9583e5eb9904ff2e28634c1befa8c3d562fe9911f31ea1 \ + --hash=sha256:34a98ab0e88199ca3003e7026a2f16cf2fd8758fab95351db172cbb7cace5e92 \ + --hash=sha256:3575e55919eee4965a40548de77f4ba4f5086bb43afe21d9ad81525f24bf0a36 \ + --hash=sha256:38e898048ced93aa822e8d0884899cd8f39323127c1d413f4d3c5c2ae5c1ca8e \ + --hash=sha256:39428a65a1b7293fb73974193ad414799d71a1779d88ed9a25f265f337abc69f \ + --hash=sha256:3af62f871733e3d1a77659d0bd8e3790a072a9f321bc667df741481bcc417dda \ + --hash=sha256:3de04c45cbcd9810d5b9dc29c934339d5c5efbc16e3baa97ede6183eb17c2898 \ + --hash=sha256:3e2b06ce8f6578067d5437638813a372f876c885016eb910fdd3e98297c63440 \ + --hash=sha256:3f17dc55aff75e6ee357a11442df0035888ba88c2de7de282d669a2655019ade \ + --hash=sha256:42d8cfdb54e199da4346811d2a1ddf7f0dabb9651900b0c4ca8faeb1329bca84 \ + --hash=sha256:472df2754d4f276571726d3eb8a546dc76a20f08c5d5237813d515de76016d08 \ + --hash=sha256:5640c634ae4b7a538475a2167d7f79a9ebca8b55e31522e03bf3841fa19dfaec \ + --hash=sha256:5a9700c3c25cb6896041282d968b3361fb8ad7eac4b0e5a199df5e256bb7cfcb \ + --hash=sha256:62c578ca66b4eaf7e998c8ba2fe38d316e324b53ce4712fa3f4207b65eb9991c \ + --hash=sha256:64519036144e5ffecf2d40348daad4399170103c780650d20052ca2f8565d2fd \ + --hash=sha256:69d2d1642c838e2f3e3525f015f2ffea62d5facd8a4815a3350a9975c8e808f8 \ + --hash=sha256:6ee92572d20739d441791c9107996080e8767bde2604e901d81519f427a1648e \ + --hash=sha256:730cf8837a21a7470d97e45912d21a623857c2c6998cb47acdc59f2d54ccd743 \ + --hash=sha256:762d2b082bdc19e254ff5a09b5ac23e9918fda26795e9175f85db9d0a22b5742 \ + --hash=sha256:78940650153233820c85327861258dbe2ac21ac6700157406748893e53758515 \ + --hash=sha256:7b60fd6fb250532ed2d871c4df2b2fe23f2485a01769e106c993308841e29f5c \ + --hash=sha256:853565b4e981023d9a7cc76aba4bf39192ac77ab80e4e8dc2a5d770fbf5119d2 \ + --hash=sha256:988b458e1b8e309afcf6bfe09e283ebdb436da421108c4fba6236ecc9fbce4f0 \ + --hash=sha256:9cfd9a6e3d6c4960eb8f1269bf66e0d07340c2d3c9ba46ca5925cdeb9fe8237c \ + --hash=sha256:a245188a31ec392ac602bd2f8f6a7a7dff6670eef4ceef3a27c86e6a60bf54b2 \ + --hash=sha256:a6463e9415cd324da9663316a5182a7db7bf41bbba843510b3d588fd6a6acfc2 \ + --hash=sha256:a96a718f832a39b6ff9569f4a86434366b0f01d397a63eb64196250d10a38041 \ + --hash=sha256:ab4c4f463576135cf763e5dec371a87e3f5d60393f3b36757c1e9e5a44b7b4c6 \ + --hash=sha256:abade0c00bee6b64b225250bae0ba6aa2ef811bbf6cd22017f65efbe8db9c6e1 \ + --hash=sha256:acc645dc0b5a46c679560b0f18b91d7648c7abc6fc5fdf917975faba91068d9a \ + --hash=sha256:b2d4ffd9314879059f5bddf71f4d5868f23391aedf88f2ae2e436135b87bf5df \ + --hash=sha256:b5c35e59b53d10546547f73cfbd7ca17cd1db0610ae685904ea5b879ee969a66 \ + --hash=sha256:b6559b0d2ff71164a6fea1e4e230aefeb8e4df600fe73b1f61613c968b2ddf73 \ + --hash=sha256:b6a1f6ea8845b4accf1d5c1c57e2908b01fdf19d3157ab1c20b3b629c921573f \ + --hash=sha256:b7bfc426564e9ee11133d8c3eb5ff03cbb38cf1fb1105d473a4f1e9312d61c75 \ + --hash=sha256:b7da2a8c34879d832e12173bc9a55bc15e16ca2a065761e055b92d064602ded3 \ + --hash=sha256:b99e32ddec41656eeeb8605fbc0c1d180f8eb7121abe1843e601762838bb6628 \ + --hash=sha256:ba9a3f382493aea444e28f079f445c7b972ba894f3a6a4913583c5a7ea1ff90e \ + --hash=sha256:bd1414b8d4c5c6f304290bcdb3b1089ce13163cf02be6234858b05f639da0e06 \ + --hash=sha256:c08fdfe05b90f78e2d9c5d0f9af91f1ebca484eca97646b183884a2d695c4651 \ + --hash=sha256:c31e33dcc775e21f3a0839272d54b80a791eba7a121c10bc51a8486baf3e9962 \ + --hash=sha256:cd4f523e55182136bf25481bb10a1493bffb71816583ec46453fadaf12af450f \ + --hash=sha256:da584101f6ed68139b096acb8a4f42e26e39487da1b83695749501932b8590a6 \ + --hash=sha256:e59075251b3128bb7ac934a8901d5045f4c2e77d5a9da439f0a102b1d23e3727 \ + --hash=sha256:e6e1ddf0457559a5d749573590e43c413ef5cb175a1e510286e0110360b40cce \ + --hash=sha256:ea0e4792eec32560a058772c4763c7b33154b5c2b8638f7b215cb5fcb1086f4f \ + --hash=sha256:fc5bcfa864b727c73375d1989ca976dc886bb2e8b3fbb61a3d649d0096df4150 \ + --hash=sha256:ff9dcdeada707306fa6bff29bc3c573cf4dd54575534d3f582085b73f29680d6 + # via semantic-data-portal (pyproject.toml) idna==3.18 \ --hash=sha256:7f952cbe720b688055e3f87de14f5c3e5fdaa8bc3928985c4077ca689de849a2 \ --hash=sha256:ffb385a7e039654cef1ab9ef32c6fafe283c0c0467bba1d9029738ce4a14a848 @@ -417,10 +558,78 @@ pyjwt==2.13.0 \ --hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \ --hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728 # via semantic-data-portal (pyproject.toml) +pypdf==6.14.2 \ + --hash=sha256:3f07891af76dc002657e04993ab9b4de81de29f9013b9761d0b7968bff12e946 \ + --hash=sha256:7873f502fe4385e79539b21d872392dc0c4e3714327c15881cbc7fbfd1f95b25 + # via semantic-data-portal (pyproject.toml) pytest==9.1.1 \ --hash=sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313 \ --hash=sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c # via semantic-data-portal (pyproject.toml) +sortedcontainers==2.4.0 \ + --hash=sha256:25caa5a06cc30b6b83d11423433f65d1f9d76c4c6a0c90e3379eaa43b9bfdb88 \ + --hash=sha256:a163dcaede0f1c021485e957a39245190e74249897e2ae4b2aa38595db237ee0 + # via hypothesis +sqlalchemy==2.0.51 \ + --hash=sha256:0378d055e9e8cd6ce4d8dff683bdd3d7d413533c4ee51d67a2b1e0f9eacc0f23 \ + --hash=sha256:0592bdadf86ddcabfd72d9ab66ea8a5d8d2cc6be1cc51fa7e66c03868ac5eac1 \ + --hash=sha256:08a204d8b5638717c26a24df18fcf40af45a6b22e35b70b1d62f0113c2e278e8 \ + --hash=sha256:0c2c62877097e1a0db401fba5cb4debee33265e5b2a55c4ccb489c02c53b4f72 \ + --hash=sha256:0e8203d2fbd5c6254692ef0a72c740d75b2f3c7ca345404f4c1a4604813c77c0 \ + --hash=sha256:0f053118c30e53161857a953e4de667d90e274980dccbe5dd3829bbbeece72a5 \ + --hash=sha256:0f6bcad487aee1c638d707235682fc96f741de00663619881ab235400d03289e \ + --hash=sha256:111604e637da87031255ddc26c7d7bc22bc6af6f5d459ccff3af1b4660233a85 \ + --hash=sha256:1181256e0f16479691b5616d36375dc2620ad8332b25978763c3d206ad3f3f1d \ + --hash=sha256:159bb6ba32059f57ad7375a8f50d844dd2f19d14954ecf820cd33e20debd46b2 \ + --hash=sha256:1aa10c0daee6705294d181daadaa793221e1a59ed55000a3fab1d42b088ce4ba \ + --hash=sha256:1af05726b3d0cdba1c55284bf408fd3b792e690fe2399bfb8304565551cda652 \ + --hash=sha256:1bed1ee8b01da6088210aa9412023326fb98a599ba502e6118308601dcbef77f \ + --hash=sha256:1d21ce524ab86c23046e992a5b81cb54c21079c6df6e78b8fc77d77cac70a6b9 \ + --hash=sha256:1e47b1199c2e832e325eacabc8d32d2487f58c9358f97e9a00f5eb93c5680d84 \ + --hash=sha256:247acaa29ccef6250dfd6a3eedf8f94ddf23564180a39fe362e32ae9dbdbde46 \ + --hash=sha256:2a97eaad21c84b4ef8010b11eeba9fe6153eb0b3df3ff8b6abc309df1b978ef7 \ + --hash=sha256:2cf39aabdf48e87c1c2c2ed6d20d33ffa0733b3071ce9c5f66357947dd009080 \ + --hash=sha256:2e54ff2dd657f2e3e0fbf2b097db1182f7bfea263eca4353f00065bae2a67c3d \ + --hash=sha256:39a76529db6305693d8d4affa58ad5b5e2e18edd62daea628b29b97930b3513d \ + --hash=sha256:4004ada0aafe8ae1991b2cd1d99c6d9146126e123bd6f883c260d974aa012e54 \ + --hash=sha256:436728ce18a80f6951a1e11cc6112c2ede9faf20766f1a26195a7c441ca12dbd \ + --hash=sha256:483b11bd46bf35fc14c52faf338b04300c9e6ce554bce9b11be85bfec3bc3195 \ + --hash=sha256:4a011ea4510683319ce4ed274b56ee05194b39b6da9d09ca7a39388f0fa84dcc \ + --hash=sha256:581921d849d6e6f994d560389192955e80e2950e18fcdfe2ccea863e01158e6e \ + --hash=sha256:59cab3686b1bc039dd9cded2f8d0c08a246e84e76bd4ab5b4f18c7cdae293825 \ + --hash=sha256:6b588fd681ddf0c196b8df1ea49a8913514894b2b8f945a9511b4b48871f99c8 \ + --hash=sha256:6e46fc36029eff666391e0531e5387b62ce6c4f1d8e50b3fb3099eaca1b42522 \ + --hash=sha256:6ea306caaae6bd5afd0a46050003c88f6bf33227377a49298c498c3cb88ff491 \ + --hash=sha256:72ca54c952107ba5cd58854b67a5a6268631289d21651a1235396f3b98b47400 \ + --hash=sha256:740cf6f35351b1ac3d82369152acf1d51d37e3dcf85d4dc0a22ca01410eabe2a \ + --hash=sha256:7c2056838b6685b72fdb36c99996cf862753461a62f2e84f4196371d3b2d6a07 \ + --hash=sha256:7c6b36ed71f41942bdcd2ad2522be46bfce09d5705be5640ecf19bbc7660e4b7 \ + --hash=sha256:7d78702b26ba1c18b2d0fb2ea940ba7f17a9581b42e8361ff93920ebbee1235a \ + --hash=sha256:804dccd8a4a6242c4e30ad961e540e18a588f6527202f2d6791b01845d59fdc9 \ + --hash=sha256:9161cfc9efce70d1715f47d6ff40f79c6778c00d53be4fbc09d70301e4b83ba7 \ + --hash=sha256:96747bfbadb055466e5b46d572618170046b45ce5a4879167f50d70a5319a499 \ + --hash=sha256:9f380393be5abeb6815f68fd39271b95127173511b6706b0a630a9995d53f8f5 \ + --hash=sha256:a42ad6afcbaaa777241e347aa2e29155993045a0d6b7db74da61053ffe875fe0 \ + --hash=sha256:a5b2ed6d828f1f09bd812861f4f59ca3bc3803f9df871f4555187f0faf018604 \ + --hash=sha256:a6d26094615306d116dd5e4a51b0304c99dd2356fc569eed6922a80a6bd3b265 \ + --hash=sha256:aa18ae738b5170e253ad0bb6c4b0f07585081e8a6e50893e4d911d47b39a0904 \ + --hash=sha256:ad30ae663711786303fbcd46a47516302d201ee49a877cb3fac61f672895110a \ + --hash=sha256:b21f0e7efc7a5c509e953784e9d1575ebb8b4318960e7e7d7a93bb803626cf64 \ + --hash=sha256:b3e693d15533a45cd5906f0589f9c35090bef6ef45bf1e8195c424aa0ae06a8d \ + --hash=sha256:b7f08588854bbb724041d9ae9d980d40040c922382e1d9a2ecb390edc4fd5032 \ + --hash=sha256:b93ab07b5292dbe7e6b8da89475275e7042744283921344b56105f3eeb0f828b \ + --hash=sha256:bb024d8b621d0be75f4f44ecc7c950450026e76d66dc8f791bb5331d7fed59d5 \ + --hash=sha256:bb1f5062f98b0b3290e72b707747fdd7e0f22d6956b236ba7ca7f5c9971d2da2 \ + --hash=sha256:c45a496d6bc05dec41dcd4c3a2b183723f47473255c159cd80b503c8f246424d \ + --hash=sha256:c5d98a2709840027f5a347c3af0a7c3d5f6c1ff93af2ca1c54494e23cba8f389 \ + --hash=sha256:c68568f3facf8f66fa76c60e0ced69b67666ffa9941d1d0a3756fda196049080 \ + --hash=sha256:c95ef01f53233a305a874a44a63fbfb1d81cd79b49de0f8529b3548cde437e37 \ + --hash=sha256:ca216e8af5c05e326efc7e28716ac2381a7cf9791749f5ee1849dccdc99c9b00 \ + --hash=sha256:ca8435d13829b92f4a97362d91975154a4015db3a2634154e1754e9a915e6b86 \ + --hash=sha256:dc261707bf5739aea8a541593f3cc1d463c2701fb05fbcbba0ce031b69a21260 \ + --hash=sha256:e5ea1a213be1fcd5e49d9904c3b9939211ded90bc2a64e93f4c01963474285de \ + --hash=sha256:fa268106c8987639a17a18514cfe0cd9bf17420ab887e1e1bf486da8836135b1 + # via semantic-data-portal (pyproject.toml) starlette==1.3.1 \ --hash=sha256:05d0213193f2fbaae60e2ecb593b4add4262ad4e46536b54abe36f11a71724e0 \ --hash=sha256:c7372aae11c3c3f26a42df7bd626cec2f47d03483d261d369516a615a53714c6 @@ -432,6 +641,7 @@ typing-extensions==4.16.0 \ # fastapi # pydantic # pydantic-core + # sqlalchemy # typing-inspection typing-inspection==0.4.2 \ --hash=sha256:4ed1cacbdc298c220f1bd249ed5287caa16f34d44ef4e9c3d0cbad5b521545e7 \ diff --git a/requirements-test.in b/requirements-test.in index 2bd73ed..d216e97 100644 --- a/requirements-test.in +++ b/requirements-test.in @@ -4,6 +4,7 @@ fastapi==0.139.0 psycopg[binary]==3.3.4 PyJWT[crypto]==2.13.0 pydantic==2.13.4 +pypdf==6.14.2 uvicorn==0.51.0 pytest==9.1.1 httpx==0.28.1 diff --git a/requirements-test.txt b/requirements-test.txt index ca0ee14..dadda76 100644 --- a/requirements-test.txt +++ b/requirements-test.txt @@ -1,5 +1,5 @@ # This file was autogenerated by uv via the following command: -# uv pip compile --generate-hashes --universal --python-version 3.12 requirements-test.in -o requirements-test.txt +# uv pip compile requirements-test.in --generate-hashes -o requirements-test.txt annotated-doc==0.0.4 \ --hash=sha256:571ac1dc6991c450b25a9c2d84a3705e2ae7a53467b5d111c24fa8baabbed320 \ --hash=sha256:fbcda96e87e9c92ad167c2e53839e57503ecfda18804ea28102353485033faa4 @@ -20,7 +20,7 @@ certifi==2026.6.17 \ # via # httpcore # httpx -cffi==2.1.0 ; platform_python_implementation != 'PyPy' \ +cffi==2.1.0 \ --hash=sha256:02cb7ff33ded4f1532476731f89ede53e2e488a8e6205515a82144246ffa7dcc \ --hash=sha256:03e9810d18c646077e501f661b682fbf5dee4676048527ca3cffe66faa9960dd \ --hash=sha256:0520e1f4c35f44e209cbbb421b67eec42e6a157f59444dfb6058874ff3610e5d \ @@ -126,7 +126,7 @@ click==8.4.2 \ --hash=sha256:9a6cea6e60b17ebe0a44c5cc636d94f09bd66142c1cd7d8b4cd731c4917a15f6 \ --hash=sha256:e6f9f66136c816745b9d65817da91d61d957fb16e02e4dcd0552553c5a197b76 # via uvicorn -colorama==0.4.6 ; sys_platform == 'win32' \ +colorama==0.4.6 \ --hash=sha256:08695f5cb7ed6e0531a20572697297273c47b8cae5a63ffc6d6ed5c201be6e44 \ --hash=sha256:4f1d9991f5acc0ca119f9d443620b77f9d6b33703e51011c16baf57afb285fc6 # via @@ -184,7 +184,7 @@ fastapi==0.139.0 \ --hash=sha256:99ab7b2d92223c76d6cf10757ab3f89d45b38267fc20b2a136cf02f6beac3145 \ --hash=sha256:cf15e1e9e667ddb0ad63811e60bd11390d1aac838ca4a7a23f421807b2308189 # via -r requirements-test.in -greenlet==3.5.3 ; platform_machine == 'AMD64' or platform_machine == 'WIN32' or platform_machine == 'aarch64' or platform_machine == 'amd64' or platform_machine == 'ppc64le' or platform_machine == 'win32' or platform_machine == 'x86_64' \ +greenlet==3.5.3 \ --hash=sha256:0909f9355a9f24845d3299f3112e266a06afb68302041989fd26bd68894933db \ --hash=sha256:0f41e4a05a3c0cb31b17023eff28dd111e1d16bf7d7d00406cd7df23f31398a7 \ --hash=sha256:0f6ff50ff8dbd51fae9b37f4101648b04ea0df19b3f50ab2beb5061e7716a5c8 \ @@ -361,7 +361,7 @@ psycopg==3.3.4 \ --hash=sha256:b6bbc25ccf05c8fad3b061d9db2ef0909a555171b84b07f29458a447253d679a \ --hash=sha256:e21207764952cff81b6b8bdacad9a3939f2793367fdac2987b3aac36a651b5bc # via -r requirements-test.in -psycopg-binary==3.3.4 ; implementation_name != 'pypy' \ +psycopg-binary==3.3.4 \ --hash=sha256:018fbed325936da502feb546642c982dcc4b9ffdea32dfef78dbf3b7f7ad4070 \ --hash=sha256:0579252a1202cd73e4da137a1426e2dae993ae44e757605344282af3a082848c \ --hash=sha256:136f199a407b5348b9b857c504aff60c77622a28482e7195839ce1b51238c4cc \ @@ -418,7 +418,7 @@ psycopg-binary==3.3.4 ; implementation_name != 'pypy' \ --hash=sha256:fa1cbc10768a796c96d3243656016bf4e337c81c71097270bb7b0ad6210d9765 \ --hash=sha256:fbd1d4ed566895ad2d3bf4ddfd8bae90026930ddf29df3b9d91d32c8c47866a7 # via psycopg -pycparser==3.0 ; implementation_name != 'PyPy' and platform_python_implementation != 'PyPy' \ +pycparser==3.0 \ --hash=sha256:600f49d217304a5902ac3c37e1281c9fe94e4d0489de643a9504c5cdfdfc6b29 \ --hash=sha256:b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992 # via cffi @@ -558,6 +558,10 @@ pyjwt==2.13.0 \ --hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \ --hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728 # via -r requirements-test.in +pypdf==6.14.2 \ + --hash=sha256:3f07891af76dc002657e04993ab9b4de81de29f9013b9761d0b7968bff12e946 \ + --hash=sha256:7873f502fe4385e79539b21d872392dc0c4e3714327c15881cbc7fbfd1f95b25 + # via -r requirements-test.in pytest==9.1.1 \ --hash=sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313 \ --hash=sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c @@ -634,13 +638,10 @@ typing-extensions==4.16.0 \ --hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \ --hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5 # via - # anyio # fastapi - # psycopg # pydantic # pydantic-core # sqlalchemy - # starlette # typing-inspection typing-inspection==0.4.2 \ --hash=sha256:4ed1cacbdc298c220f1bd249ed5287caa16f34d44ef4e9c3d0cbad5b521545e7 \ @@ -648,7 +649,7 @@ typing-inspection==0.4.2 \ # via # fastapi # pydantic -tzdata==2026.3 ; sys_platform == 'win32' \ +tzdata==2026.3 \ --hash=sha256:4a1518b8993086a7982523e071643f3c0e5f213e75b21318e78bcabfff9d1415 \ --hash=sha256:dc096730c87af6cab1b171c9d532be840741ff5d459015e7f6947bd7d7e54931 # via psycopg diff --git a/requirements.txt b/requirements.txt index 8e618b9..e95359b 100644 --- a/requirements.txt +++ b/requirements.txt @@ -379,6 +379,10 @@ pyjwt==2.13.0 \ --hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \ --hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728 # via semantic-data-portal (pyproject.toml) +pypdf==6.14.2 \ + --hash=sha256:3f07891af76dc002657e04993ab9b4de81de29f9013b9761d0b7968bff12e946 \ + --hash=sha256:7873f502fe4385e79539b21d872392dc0c4e3714327c15881cbc7fbfd1f95b25 + # via semantic-data-portal (pyproject.toml) starlette==1.3.1 \ --hash=sha256:05d0213193f2fbaae60e2ecb593b4add4262ad4e46536b54abe36f11a71724e0 \ --hash=sha256:c7372aae11c3c3f26a42df7bd626cec2f47d03483d261d369516a615a53714c6 diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py new file mode 100644 index 0000000..1afcdd6 --- /dev/null +++ b/src/sdp/document_semantics.py @@ -0,0 +1,141 @@ +"""Ephemeral document text extraction and semantic chunking.""" + +from __future__ import annotations + +import hashlib +import xml.etree.ElementTree as ET +from dataclasses import dataclass +from io import BytesIO +from pathlib import PurePath +from typing import Literal +from zipfile import BadZipFile, ZipFile + +from pypdf import PdfReader +from pypdf.errors import PdfReadError + + +ExtractionStatus = Literal["extracted", "needs_ocr", "unsupported_format", "extraction_failed"] + +_PLAIN_SUFFIXES = {".txt", ".md", ".csv", ".json", ".xml"} +_OPENXML_ROOTS = { + ".docx": ("word/",), + ".pptx": ("ppt/slides/",), + ".xlsx": ("xl/sharedStrings.xml", "xl/worksheets/"), +} +_MEDIA_TYPES = { + ".txt": "text/plain", + ".md": "text/markdown", + ".csv": "text/csv", + ".json": "application/json", + ".xml": "application/xml", + ".docx": "application/vnd.openxmlformats-officedocument.wordprocessingml.document", + ".pptx": "application/vnd.openxmlformats-officedocument.presentationml.presentation", + ".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet", + ".pdf": "application/pdf", +} + + +@dataclass(frozen=True) +class DocumentText: + status: ExtractionStatus + text: str + media_type: str + error: str | None = None + + +@dataclass(frozen=True) +class TextChunk: + text: str + start: int + end: int + sha256: str + + +def _decode_text(data: bytes) -> str: + for encoding in ("utf-8-sig", "utf-16", "cp949"): + try: + return data.decode(encoding) + except UnicodeDecodeError: + continue + return data.decode("utf-8", errors="replace") + + +def _xml_text(xml: bytes) -> list[str]: + root = ET.fromstring(xml) + return [ + element.text.strip() + for element in root.iter() + if element.tag.rsplit("}", 1)[-1] == "t" and element.text and element.text.strip() + ] + + +def _openxml_text(suffix: str, data: bytes) -> str: + roots = _OPENXML_ROOTS[suffix] + parts: list[str] = [] + with ZipFile(BytesIO(data)) as archive: + for name in sorted(archive.namelist()): + if not name.endswith(".xml") or not any( + name == root or name.startswith(root) for root in roots + ): + continue + parts.extend(_xml_text(archive.read(name))) + return "\n".join(parts) + + +def _pdf_text(data: bytes) -> str: + reader = PdfReader(BytesIO(data), strict=False) + return "\n".join(filter(None, (page.extract_text() for page in reader.pages))) + + +def extract_document_text(filename: str, data: bytes) -> DocumentText: + """Extract text without retaining the source bytes or writing temporary files.""" + + suffix = PurePath(filename).suffix.lower() + media_type = _MEDIA_TYPES.get(suffix, "application/octet-stream") + if suffix not in _PLAIN_SUFFIXES | set(_OPENXML_ROOTS) | {".pdf"}: + return DocumentText("unsupported_format", "", media_type) + try: + if suffix in _PLAIN_SUFFIXES: + text = _decode_text(data) + elif suffix in _OPENXML_ROOTS: + text = _openxml_text(suffix, data) + else: + text = _pdf_text(data) + except (BadZipFile, ET.ParseError, PdfReadError, ValueError, TypeError) as exc: + return DocumentText("extraction_failed", "", media_type, type(exc).__name__) + + normalized = "\n".join(line.strip() for line in text.splitlines() if line.strip()) + return DocumentText("extracted" if normalized else "needs_ocr", normalized, media_type) + + +def chunk_text( + text: str, + *, + max_chars: int = 6_000, + overlap: int = 300, + max_total_chars: int = 24_000, +) -> list[TextChunk]: + """Create bounded, deterministic chunks whose hashes can serve as evidence refs.""" + + if max_chars <= 0 or max_total_chars <= 0: + raise ValueError("chunk limits must be positive") + if overlap < 0 or overlap >= max_chars: + raise ValueError("overlap must be non-negative and smaller than max_chars") + normalized = " ".join(text.split())[:max_total_chars] + chunks: list[TextChunk] = [] + start = 0 + while start < len(normalized): + end = min(start + max_chars, len(normalized)) + value = normalized[start:end] + chunks.append( + TextChunk( + text=value, + start=start, + end=end, + sha256=hashlib.sha256(value.encode("utf-8")).hexdigest(), + ) + ) + if end == len(normalized): + break + start = end - overlap + return chunks diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index 1940a84..d973eeb 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -3,9 +3,14 @@ from copy import deepcopy from io import BytesIO from types import SimpleNamespace +from zipfile import ZIP_DEFLATED, ZipFile import pytest +from pypdf import PdfWriter + +from sdp.document_semantics import chunk_text, extract_document_text + from sdp.file_ontology import ( FileAsset, SemanticAssertion, @@ -211,3 +216,65 @@ def test_s3_compatible_reader_uses_stable_endpoint_without_credentials(): assert ref.distribution.provider == "s3_compatible" assert ref.distribution.locator == "https://objects.example/voc/reports/a.docx" + + +def make_openxml(member: str, text: str) -> bytes: + payload = BytesIO() + with ZipFile(payload, "w", ZIP_DEFLATED) as archive: + archive.writestr( + member, + f'{text}', + ) + return payload.getvalue() + + +@pytest.mark.parametrize( + ("filename", "member"), + [ + ("meeting.docx", "word/document.xml"), + ("briefing.pptx", "ppt/slides/slide1.xml"), + ("feedback.xlsx", "xl/sharedStrings.xml"), + ], +) +def test_openxml_text_is_extracted_without_office_dependency(filename, member): + payload = make_openxml(member, "효성중공업 VOC C-Cube") + + document = extract_document_text(filename, payload) + + assert document.status == "extracted" + assert "효성중공업 VOC C-Cube" in document.text + + +def test_chunks_are_bounded_overlapped_and_content_addressed(): + chunks = chunk_text("가" * 13_000, max_chars=6_000, overlap=300, max_total_chars=24_000) + + assert len(chunks) == 3 + assert max(len(chunk.text) for chunk in chunks) <= 6_000 + assert all(len(chunk.sha256) == 64 for chunk in chunks) + assert chunks[1].start == chunks[0].end - 300 + + +def test_pdf_without_extractable_text_is_marked_for_ocr(): + payload = BytesIO() + writer = PdfWriter() + writer.add_blank_page(width=100, height=100) + writer.write(payload) + + document = extract_document_text("scan.pdf", payload.getvalue()) + + assert document.status == "needs_ocr" + assert document.text == "" + + +def test_corrupt_pdf_is_reported_without_crashing_the_batch(): + document = extract_document_text("broken.pdf", b"not a pdf") + + assert document.status == "extraction_failed" + assert document.text == "" + + +def test_unsupported_legacy_format_is_reported_without_guessing(): + document = extract_document_text("legacy.hwp", b"not parsed") + + assert document.status == "unsupported_format" + assert document.text == "" From 0a970a9921a3f60ea438dc1bd96b18caa01dc087 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 10:52:27 +0900 Subject: [PATCH 06/15] feat: extract evidence-bound semantics with OpenAI --- src/sdp/document_semantics.py | 203 +++++++++++++++++++++++++++++++++- tests/test_file_knowledge.py | 82 +++++++++++++- 2 files changed, 283 insertions(+), 2 deletions(-) diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py index 1afcdd6..a863df6 100644 --- a/src/sdp/document_semantics.py +++ b/src/sdp/document_semantics.py @@ -3,16 +3,21 @@ from __future__ import annotations import hashlib +import json import xml.etree.ElementTree as ET from dataclasses import dataclass from io import BytesIO from pathlib import PurePath -from typing import Literal +from typing import Any, Callable, Literal, Mapping, Protocol +from urllib.error import HTTPError, URLError +from urllib.request import Request, urlopen from zipfile import BadZipFile, ZipFile from pypdf import PdfReader from pypdf.errors import PdfReadError +from .file_ontology import SemanticAssertion + ExtractionStatus = Literal["extracted", "needs_ocr", "unsupported_format", "extraction_failed"] @@ -51,6 +56,23 @@ class TextChunk: sha256: str +class CredentialRegistry(Protocol): + def get_credential(self, name: str) -> str | None: ... + + +class EphemeralCredentialRegistry: + """Process-local credentials supplied by a trusted bootstrap caller.""" + + def __init__(self, values: Mapping[str, str]) -> None: + self._values = dict(values) + + def get_credential(self, name: str) -> str | None: + return self._values.get(name) + + def __repr__(self) -> str: + return f"{type(self).__name__}(names={sorted(self._values)})" + + def _decode_text(data: bytes) -> str: for encoding in ("utf-8-sig", "utf-16", "cp949"): try: @@ -139,3 +161,182 @@ def chunk_text( break start = end - overlap return chunks + + +OpenAITransport = Callable[[Request, int], dict[str, Any]] + +_SEMANTIC_SCHEMA = { + "type": "object", + "additionalProperties": False, + "properties": { + "assertions": { + "type": "array", + "maxItems": 20, + "items": { + "type": "object", + "additionalProperties": False, + "properties": { + "relation": { + "type": "string", + "enum": [ + "belongsToProject", + "usesSystem", + "hasWorkPhase", + "hasArtifactType", + "hasTopic", + "wasDerivedFrom", + "previousVersion", + ], + }, + "target_kind": { + "type": "string", + "enum": [ + "business_project", + "system", + "work_phase", + "artifact_type", + "topic", + "file_asset", + ], + }, + "target_label": {"type": "string", "minLength": 1, "maxLength": 200}, + "confidence": {"type": "number", "minimum": 0, "maximum": 1}, + "evidence_quote": {"type": "string", "minLength": 1, "maxLength": 240}, + }, + "required": [ + "relation", + "target_kind", + "target_label", + "confidence", + "evidence_quote", + ], + }, + } + }, + "required": ["assertions"], +} + + +def _openai_http_transport(request: Request, timeout: int) -> dict[str, Any]: + try: + with urlopen(request, timeout=timeout) as response: + return json.loads(response.read().decode("utf-8")) + except HTTPError as exc: + raise RuntimeError(f"OpenAI request failed with HTTP {exc.code}") from None + except URLError: + raise RuntimeError("OpenAI request could not reach the service") from None + + +def _response_output_text(response: dict[str, Any]) -> str: + if response.get("status") != "completed": + raise ValueError("OpenAI response did not complete") + for output in response.get("output", []): + if output.get("type") != "message": + continue + for content in output.get("content", []): + if content.get("type") == "output_text" and isinstance(content.get("text"), str): + return content["text"] + raise ValueError("OpenAI response contains no output text") + + +class OpenAISemanticExtractor: + def __init__( + self, + credentials: CredentialRegistry, + *, + transport: OpenAITransport = _openai_http_transport, + model: str = "gpt-5-mini-2025-08-07", + timeout: int = 60, + ) -> None: + self.credentials = credentials + self.transport = transport + self.model = model + self.timeout = timeout + + def _payload(self, filename: str, chunk: TextChunk) -> dict[str, Any]: + return { + "model": self.model, + "store": False, + "instructions": ( + "Extract only explicitly evidenced business semantics from the document chunk. " + "Return Korean labels when the source uses Korean. Do not infer facts absent from the text." + ), + "input": [ + { + "role": "user", + "content": [ + { + "type": "input_text", + "text": f"Filename: {filename}\nDocument chunk:\n{chunk.text}", + } + ], + } + ], + "text": { + "format": { + "type": "json_schema", + "name": "cwl_file_semantics", + "strict": True, + "schema": _SEMANTIC_SCHEMA, + } + }, + "max_output_tokens": 1_500, + } + + def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: + api_key = self.credentials.get_credential("OPENAI_API_KEY") + if not api_key: + raise ValueError("OpenAI credential is unavailable") + + merged: dict[tuple[str, str, str], SemanticAssertion] = {} + for chunk in chunks: + encoded_payload = json.dumps( + self._payload(filename, chunk), ensure_ascii=False + ).encode("utf-8") + request = Request( + "https://api.openai.com/v1/responses", + data=encoded_payload, + headers={ + "Authorization": f"Bearer {api_key}", + "Content-Type": "application/json", + }, + method="POST", + ) + response = self.transport(request, self.timeout) + try: + result = json.loads(_response_output_text(response)) + candidates = result["assertions"] + except (KeyError, TypeError, json.JSONDecodeError) as exc: + raise ValueError("OpenAI semantic output is malformed") from exc + if not isinstance(candidates, list): + raise ValueError("OpenAI semantic output assertions must be a list") + + for candidate in candidates: + if not isinstance(candidate, dict): + raise ValueError("OpenAI semantic assertion must be an object") + quote_text = candidate.get("evidence_quote") + if not isinstance(quote_text, str): + raise ValueError("OpenAI semantic assertion has no evidence quote") + local_start = chunk.text.find(quote_text) + if local_start < 0: + continue + try: + assertion = SemanticAssertion( + relation=candidate["relation"], + target_kind=candidate["target_kind"], + target_label=candidate["target_label"], + confidence=candidate["confidence"], + evidence_chunk_sha256=chunk.sha256, + evidence_start=chunk.start + local_start, + evidence_end=chunk.start + local_start + len(quote_text), + ) + except (KeyError, TypeError, ValueError) as exc: + raise ValueError("OpenAI semantic assertion is invalid") from exc + key = ( + assertion.relation, + assertion.target_kind, + assertion.target_label.strip().casefold(), + ) + if key not in merged or assertion.confidence > merged[key].confidence: + merged[key] = assertion + return list(merged.values()) diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index d973eeb..75f8945 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -1,5 +1,6 @@ from __future__ import annotations +import json from copy import deepcopy from io import BytesIO from types import SimpleNamespace @@ -9,7 +10,12 @@ from pypdf import PdfWriter -from sdp.document_semantics import chunk_text, extract_document_text +from sdp.document_semantics import ( + EphemeralCredentialRegistry, + OpenAISemanticExtractor, + chunk_text, + extract_document_text, +) from sdp.file_ontology import ( FileAsset, @@ -278,3 +284,77 @@ def test_unsupported_legacy_format_is_reported_without_guessing(): assert document.status == "unsupported_format" assert document.text == "" + + +def fake_openai_transport(captured, *, evidence_quote="효성중공업"): + def transport(request, timeout): + payload = json.loads(request.data.decode("utf-8")) + captured.update(payload) + captured["timeout"] = timeout + assert request.get_header("Authorization") == "Bearer test-key" + return { + "status": "completed", + "output": [ + { + "type": "message", + "content": [ + { + "type": "output_text", + "text": json.dumps( + { + "assertions": [ + { + "relation": "usesSystem", + "target_kind": "system", + "target_label": "C-Cube", + "confidence": 0.93, + "evidence_quote": evidence_quote, + } + ] + }, + ensure_ascii=False, + ), + } + ], + } + ], + } + + return transport + + +def test_openai_extractor_uses_strict_schema_and_persists_only_evidence_reference(): + captured = {} + extractor = OpenAISemanticExtractor( + EphemeralCredentialRegistry({"OPENAI_API_KEY": "test-key"}), + transport=fake_openai_transport(captured), + ) + + assertions = extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) + + assert captured["store"] is False + assert captured["model"] == "gpt-5-mini-2025-08-07" + assert captured["text"]["format"]["type"] == "json_schema" + assert captured["text"]["format"]["strict"] is True + assert captured["timeout"] == 60 + assert assertions[0].relation == "usesSystem" + assert assertions[0].evidence_chunk_sha256 + assert assertions[0].evidence_start < assertions[0].evidence_end + assert "효성중공업" not in assertions[0].model_dump_json() + assert "test-key" not in repr(captured) + + +def test_openai_extractor_rejects_quote_not_present_in_input(): + extractor = OpenAISemanticExtractor( + EphemeralCredentialRegistry({"OPENAI_API_KEY": "test-key"}), + transport=fake_openai_transport({}, evidence_quote="invented evidence"), + ) + + assert extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) == [] + + +def test_openai_extractor_fails_closed_without_credential(): + extractor = OpenAISemanticExtractor(EphemeralCredentialRegistry({}), transport=lambda *_: {}) + + with pytest.raises(ValueError, match="credential is unavailable"): + extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) From 7106dcb2a73eb835e95d96ad5c562b493ffd629e Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 11:21:38 +0900 Subject: [PATCH 07/15] feat: route file knowledge through contextual orchestrator --- ontology/cwl-file-profile.ttl | 8 + ontology/cwl-file-shapes.ttl | 3 +- src/sdp/api.py | 86 +++++++++- src/sdp/config.py | 9 ++ src/sdp/document_semantics.py | 158 +++++++++++------- src/sdp/file_ontology.py | 116 +++++++++++++- src/sdp/file_pilot.py | 180 +++++++++++++++++++++ src/sdp/graph_store.py | 46 +++++- tests/test_file_knowledge.py | 292 +++++++++++++++++++++++++++++----- tests/test_graph_engine.py | 6 + 10 files changed, 798 insertions(+), 106 deletions(-) create mode 100644 src/sdp/file_pilot.py diff --git a/ontology/cwl-file-profile.ttl b/ontology/cwl-file-profile.ttl index 85f6c86..a44989a 100644 --- a/ontology/cwl-file-profile.ttl +++ b/ontology/cwl-file-profile.ttl @@ -3,6 +3,7 @@ @prefix owl: . @prefix prov: . @prefix skos: . +@prefix xsd: . cwl: a owl:Ontology ; owl:versionInfo "0.1.0" . @@ -29,3 +30,10 @@ cwl:hasTopic a owl:ObjectProperty ; cwl:wasDerivedFrom a owl:ObjectProperty ; owl:subPropertyOf prov:wasDerivedFrom . cwl:previousVersion a owl:ObjectProperty . + +cwl:confidence a owl:DatatypeProperty ; owl:range xsd:decimal . +cwl:reviewStatus a owl:DatatypeProperty ; owl:range xsd:string . +cwl:evidenceChunkSha256 a owl:DatatypeProperty ; owl:range xsd:string . +cwl:evidenceStart a owl:DatatypeProperty ; owl:range xsd:nonNegativeInteger . +cwl:evidenceEnd a owl:DatatypeProperty ; owl:range xsd:positiveInteger . +cwl:extractionMethod a owl:DatatypeProperty ; owl:range xsd:string . diff --git a/ontology/cwl-file-shapes.ttl b/ontology/cwl-file-shapes.ttl index bd6d4ed..3b81d6c 100644 --- a/ontology/cwl-file-shapes.ttl +++ b/ontology/cwl-file-shapes.ttl @@ -21,4 +21,5 @@ cwl:SemanticAssertionShape a sh:NodeShape ; sh:property [ sh:path cwl:reviewStatus ; sh:minCount 1 ; sh:maxCount 1 ] ; sh:property [ sh:path cwl:evidenceChunkSha256 ; sh:minCount 1 ; sh:maxCount 1 ] ; sh:property [ sh:path cwl:evidenceStart ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:evidenceEnd ; sh:minCount 1 ; sh:maxCount 1 ] . + sh:property [ sh:path cwl:evidenceEnd ; sh:minCount 1 ; sh:maxCount 1 ] ; + sh:property [ sh:path cwl:extractionMethod ; sh:minCount 1 ; sh:maxCount 1 ] . diff --git a/src/sdp/api.py b/src/sdp/api.py index 66795bc..bb57fef 100644 --- a/src/sdp/api.py +++ b/src/sdp/api.py @@ -28,6 +28,14 @@ SemanticSearchRequest, ) from .graph_store import get_store +from .file_ontology import ( + FileAsset, + FileAssetIngestRequest, + file_asset_jsonld, + get_file_asset, + upsert_file_asset, + validate_file_asset, +) from .seed import seed_store from .catalog import ( deprecate_dataset, @@ -754,7 +762,83 @@ def get_graph_node(node_id: str) -> dict[str, Any]: node = get_store().get_node(node_id) if node is None: raise HTTPException(status_code=404, detail="node not found") - return node.as_dict() + return _redact_storage_coordinates(node.as_dict()) + + +# --- Standards-aligned file assets ------------------------------------------ + + +_STORAGE_COORDINATE_KEYS = {"locator", "bucket", "container", "object_key"} + + +def _redact_storage_coordinates(value: Any) -> Any: + if isinstance(value, dict): + return { + key: _redact_storage_coordinates(item) + for key, item in value.items() + if key not in _STORAGE_COORDINATE_KEYS + } + if isinstance(value, list): + return [_redact_storage_coordinates(item) for item in value] + return value + + +@app.post("/file-assets") +def ingest_file_asset(payload: FileAssetIngestRequest) -> dict[str, Any]: + _authorize_graph_write(payload.actor, payload.asset_id) + try: + asset = upsert_file_asset( + get_store(), + FileAsset.model_validate(payload.model_dump(exclude={"actor"})), + ) + except ValueError as exc: + raise HTTPException(status_code=400, detail=str(exc)) + return { + "status": "upserted", + "asset_id": asset.asset_id, + "asset": asset.model_dump(mode="json"), + } + + +def _read_file_asset(asset_id: str, actor: str): + _authorize_graph_read(actor, asset_id) + asset = get_file_asset(get_store(), asset_id) + if asset is None: + raise HTTPException(status_code=404, detail="file asset not found") + return asset + + +@app.get("/file-assets/{asset_id}") +def file_asset_detail( + asset_id: str, + actor: str = Query(default="anonymous"), + include_locations: bool = Query(default=False), +) -> dict[str, Any]: + asset = _read_file_asset(asset_id, actor) + if include_locations: + _authorize_graph_write(actor, asset_id) + payload = {"asset_id": asset.asset_id, **asset.model_dump(mode="json")} + return payload if include_locations else _redact_storage_coordinates(payload) + + +@app.get("/file-assets/{asset_id}/jsonld") +def file_asset_jsonld_export( + asset_id: str, + actor: str = Query(default="anonymous"), + include_locations: bool = Query(default=False), +) -> dict[str, Any]: + asset = _read_file_asset(asset_id, actor) + if include_locations: + _authorize_graph_write(actor, asset_id) + return file_asset_jsonld(asset, include_locations=include_locations) + + +@app.get("/file-assets/{asset_id}/validate") +def file_asset_validation( + asset_id: str, + actor: str = Query(default="anonymous"), +) -> dict[str, Any]: + return validate_file_asset(_read_file_asset(asset_id, actor)) # --- Graph traversal + semantic retrieval ------------------------------------ diff --git a/src/sdp/config.py b/src/sdp/config.py index fc12b25..5465092 100644 --- a/src/sdp/config.py +++ b/src/sdp/config.py @@ -34,6 +34,9 @@ "cors_allow_methods": ["GET", "POST", "PATCH", "OPTIONS"], "cors_allow_headers": ["*"], "embedding_dimension": 128, + "orchestrator_base_url": "", + "semantic_model": "gpt-5-mini-2025-08-07", + "embedding_model": "text-embedding-3-small", "graph_name": "semantic_graph", "semantic_search_default_limit": 5, "traversal_max_depth": 4, @@ -94,6 +97,9 @@ class AppConfig: cors_allow_methods: List[str] = field(default_factory=list) cors_allow_headers: List[str] = field(default_factory=list) embedding_dimension: int = 128 + orchestrator_base_url: str = "" + semantic_model: str = "gpt-5-mini-2025-08-07" + embedding_model: str = "text-embedding-3-small" graph_name: str = "semantic_graph" semantic_search_default_limit: int = 5 traversal_max_depth: int = 4 @@ -112,6 +118,9 @@ def from_mapping(cls, values: Dict[str, Any], *, source: str) -> "AppConfig": cors_allow_methods=list(merged["cors_allow_methods"]), cors_allow_headers=list(merged["cors_allow_headers"]), embedding_dimension=int(merged["embedding_dimension"]), + orchestrator_base_url=str(merged["orchestrator_base_url"]).rstrip("/"), + semantic_model=str(merged["semantic_model"]), + embedding_model=str(merged["embedding_model"]), graph_name=str(merged["graph_name"]), semantic_search_default_limit=int(merged["semantic_search_default_limit"]), traversal_max_depth=int(merged["traversal_max_depth"]), diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py index a863df6..c57e29f 100644 --- a/src/sdp/document_semantics.py +++ b/src/sdp/document_semantics.py @@ -10,6 +10,7 @@ from pathlib import PurePath from typing import Any, Callable, Literal, Mapping, Protocol from urllib.error import HTTPError, URLError +from urllib.parse import urlsplit from urllib.request import Request, urlopen from zipfile import BadZipFile, ZipFile @@ -163,7 +164,7 @@ def chunk_text( return chunks -OpenAITransport = Callable[[Request, int], dict[str, Any]] +OrchestratorTransport = Callable[[Request, int], dict[str, Any]] _SEMANTIC_SCHEMA = { "type": "object", @@ -217,106 +218,146 @@ def chunk_text( } -def _openai_http_transport(request: Request, timeout: int) -> dict[str, Any]: +def _orchestrator_http_transport(request: Request, timeout: int) -> dict[str, Any]: try: with urlopen(request, timeout=timeout) as response: return json.loads(response.read().decode("utf-8")) except HTTPError as exc: - raise RuntimeError(f"OpenAI request failed with HTTP {exc.code}") from None + raise RuntimeError(f"orchestrator request failed with HTTP {exc.code}") from None except URLError: - raise RuntimeError("OpenAI request could not reach the service") from None + raise RuntimeError("orchestrator request could not reach the service") from None -def _response_output_text(response: dict[str, Any]) -> str: - if response.get("status") != "completed": - raise ValueError("OpenAI response did not complete") - for output in response.get("output", []): - if output.get("type") != "message": - continue - for content in output.get("content", []): - if content.get("type") == "output_text" and isinstance(content.get("text"), str): - return content["text"] - raise ValueError("OpenAI response contains no output text") +def _chat_completion_content(response: dict[str, Any]) -> str: + try: + content = response["choices"][0]["message"]["content"] + except (KeyError, IndexError, TypeError) as exc: + raise ValueError("orchestrator response contains no completion") from exc + if not isinstance(content, str): + raise ValueError("orchestrator completion content must be text") + return content -class OpenAISemanticExtractor: +class ContextualOrchestratorClient: + """Authenticated LLM and embedding client for contextual-orchestrator.""" + def __init__( self, credentials: CredentialRegistry, *, - transport: OpenAITransport = _openai_http_transport, - model: str = "gpt-5-mini-2025-08-07", + base_url: str, + transport: OrchestratorTransport = _orchestrator_http_transport, + semantic_model: str = "gpt-5-mini-2025-08-07", + embedding_model: str = "text-embedding-3-small", timeout: int = 60, ) -> None: + parsed = urlsplit(base_url) + if parsed.scheme not in {"http", "https"} or not parsed.netloc: + raise ValueError("orchestrator base URL must be absolute HTTP(S)") + if parsed.query or parsed.fragment: + raise ValueError("orchestrator base URL must not contain query or fragment") self.credentials = credentials + self.base_url = base_url.rstrip("/") self.transport = transport - self.model = model + self.semantic_model = semantic_model + self.embedding_model = embedding_model self.timeout = timeout def _payload(self, filename: str, chunk: TextChunk) -> dict[str, Any]: return { - "model": self.model, + "model": self.semantic_model, "store": False, - "instructions": ( - "Extract only explicitly evidenced business semantics from the document chunk. " - "Return Korean labels when the source uses Korean. Do not infer facts absent from the text." - ), - "input": [ + "messages": [ + { + "role": "system", + "content": ( + "Extract only explicitly evidenced business semantics from the document chunk. " + "Return Korean labels when the source uses Korean. Do not infer facts absent from the text." + ), + }, { "role": "user", - "content": [ - { - "type": "input_text", - "text": f"Filename: {filename}\nDocument chunk:\n{chunk.text}", - } - ], - } + "content": f"Filename: {filename}\nDocument chunk:\n{chunk.text}", + }, ], - "text": { - "format": { - "type": "json_schema", + "response_format": { + "type": "json_schema", + "json_schema": { "name": "cwl_file_semantics", "strict": True, "schema": _SEMANTIC_SCHEMA, - } + }, }, - "max_output_tokens": 1_500, + "max_completion_tokens": 1_500, } - def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: - api_key = self.credentials.get_credential("OPENAI_API_KEY") - if not api_key: - raise ValueError("OpenAI credential is unavailable") + def _post(self, path: str, payload: dict[str, Any]) -> dict[str, Any]: + token = self.credentials.get_credential("CONTEXTUAL_ORCHESTRATOR_TOKEN") + if not token: + raise ValueError("orchestrator credential is unavailable") + request = Request( + f"{self.base_url}{path}", + data=json.dumps(payload, ensure_ascii=False).encode("utf-8"), + headers={ + "Authorization": f"Bearer {token}", + "Content-Type": "application/json", + }, + method="POST", + ) + return self.transport(request, self.timeout) + + def embed(self, texts: list[str]) -> list[list[float]]: + if not texts or any(not isinstance(text, str) for text in texts): + raise ValueError("embedding input must be a non-empty list of strings") + response = self._post( + "/v1/embeddings", + { + "model": self.embedding_model, + "input": texts, + "metadata": {"service": "semantic-data-portal"}, + }, + ) + data = response.get("data") + if not isinstance(data, list) or len(data) != len(texts): + raise ValueError("orchestrator embedding response is malformed") + indexed: dict[int, list[float]] = {} + for item in data: + if not isinstance(item, dict) or not isinstance(item.get("index"), int): + raise ValueError("orchestrator embedding item is malformed") + vector = item.get("embedding") + if not isinstance(vector, list) or not vector: + raise ValueError("orchestrator embedding vector is missing") + try: + indexed[item["index"]] = [float(component) for component in vector] + except (TypeError, ValueError) as exc: + raise ValueError("orchestrator embedding vector is malformed") from exc + if set(indexed) != set(range(len(texts))): + raise ValueError("orchestrator embedding indices are malformed") + return [indexed[index] for index in range(len(texts))] + + def embed_one(self, text: str) -> list[float]: + return self.embed([text])[0] + def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: merged: dict[tuple[str, str, str], SemanticAssertion] = {} for chunk in chunks: - encoded_payload = json.dumps( - self._payload(filename, chunk), ensure_ascii=False - ).encode("utf-8") - request = Request( - "https://api.openai.com/v1/responses", - data=encoded_payload, - headers={ - "Authorization": f"Bearer {api_key}", - "Content-Type": "application/json", - }, - method="POST", + response = self._post( + "/v1/chat/completions", self._payload(filename, chunk) ) - response = self.transport(request, self.timeout) try: - result = json.loads(_response_output_text(response)) + result = json.loads(_chat_completion_content(response)) candidates = result["assertions"] except (KeyError, TypeError, json.JSONDecodeError) as exc: - raise ValueError("OpenAI semantic output is malformed") from exc + raise ValueError("orchestrator semantic output is malformed") from exc if not isinstance(candidates, list): - raise ValueError("OpenAI semantic output assertions must be a list") + raise ValueError("orchestrator semantic output assertions must be a list") for candidate in candidates: if not isinstance(candidate, dict): - raise ValueError("OpenAI semantic assertion must be an object") + raise ValueError("orchestrator semantic assertion must be an object") quote_text = candidate.get("evidence_quote") if not isinstance(quote_text, str): - raise ValueError("OpenAI semantic assertion has no evidence quote") + raise ValueError("orchestrator semantic assertion has no evidence quote") local_start = chunk.text.find(quote_text) if local_start < 0: continue @@ -329,9 +370,10 @@ def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssert evidence_chunk_sha256=chunk.sha256, evidence_start=chunk.start + local_start, evidence_end=chunk.start + local_start + len(quote_text), + method="contextual-orchestrator", ) except (KeyError, TypeError, ValueError) as exc: - raise ValueError("OpenAI semantic assertion is invalid") from exc + raise ValueError("orchestrator semantic assertion is invalid") from exc key = ( assertion.relation, assertion.target_kind, diff --git a/src/sdp/file_ontology.py b/src/sdp/file_ontology.py index 5dc8386..66a735d 100644 --- a/src/sdp/file_ontology.py +++ b/src/sdp/file_ontology.py @@ -10,6 +10,8 @@ from pydantic import BaseModel, Field, field_validator +from .graph_store import GraphStore + StorageProvider = Literal["filesystem", "s3", "s3_compatible", "azure_blob"] AssertionRelation = Literal[ @@ -81,7 +83,7 @@ class SemanticAssertion(BaseModel): evidence_chunk_sha256: str evidence_start: int evidence_end: int - method: str = "openai" + method: str = Field(min_length=1) review_status: Literal["proposed", "approved", "rejected"] = "proposed" @@ -99,6 +101,10 @@ def asset_id(self) -> str: return f"urn:sha256:{self.sha256}" +class FileAssetIngestRequest(FileAsset): + actor: str = Field(default="anonymous", min_length=1) + + def _violation(path: str, message: str) -> dict[str, str]: return { "shape": "CWLFileAssetShape", @@ -225,3 +231,111 @@ def file_asset_jsonld(asset: FileAsset, *, include_locations: bool = False) -> d if asset.modified_at: payload["dcterms:modified"] = asset.modified_at.isoformat() return payload + + +_EDGE_TYPES = { + "belongsToProject": "BELONGS_TO_PROJECT", + "usesSystem": "USES_SYSTEM", + "hasWorkPhase": "HAS_WORK_PHASE", + "hasArtifactType": "HAS_ARTIFACT_TYPE", + "hasTopic": "HAS_TOPIC", + "wasDerivedFrom": "WAS_DERIVED_FROM", + "previousVersion": "PREVIOUS_VERSION", +} + + +def get_file_asset(store: GraphStore, asset_id: str) -> FileAsset | None: + node = store.get_node(asset_id) + if node is None or node.kind != "file_asset": + return None + payload = node.properties.get("file_asset") + return FileAsset.model_validate(payload) if isinstance(payload, dict) else None + + +def _merge_asset(existing: FileAsset | None, incoming: FileAsset) -> FileAsset: + if existing is None: + return incoming + distributions = {item.id: item for item in existing.distributions} + distributions.update({item.id: item for item in incoming.distributions}) + assertions = { + (item.relation, item.target_kind, item.target_label.strip().casefold()): item + for item in existing.assertions + } + for item in incoming.assertions: + key = (item.relation, item.target_kind, item.target_label.strip().casefold()) + if key not in assertions or item.confidence > assertions[key].confidence: + assertions[key] = item + return existing.model_copy( + update={ + "distributions": list(distributions.values()), + "assertions": list(assertions.values()), + "modified_at": max( + filter(None, (existing.modified_at, incoming.modified_at)), + default=None, + ), + } + ) + + +def upsert_file_asset(store: GraphStore, asset: FileAsset) -> FileAsset: + """Validate, merge and project one content-addressed asset into the graph.""" + + report = validate_file_asset(asset) + if not report["conforms"]: + messages = "; ".join(item["message"] for item in report["violations"]) + raise ValueError(f"file asset does not conform: {messages}") + merged = _merge_asset(get_file_asset(store, asset.asset_id), asset) + embedding_text = " ".join( + [merged.title] + [assertion.target_label for assertion in merged.assertions] + ) + store.upsert_node( + merged.asset_id, + "file_asset", + label=merged.title, + properties={ + "profile": "CWL File Knowledge Profile 0.1", + "file_asset": merged.model_dump(mode="json"), + }, + text=embedding_text, + ) + + for distribution in merged.distributions: + distribution_id = f"urn:cwl:distribution:{distribution.id}" + store.upsert_node( + distribution_id, + "distribution", + label=distribution.id, + properties=distribution.model_dump(mode="json"), + text=f"{distribution.provider} {distribution.endpoint_id}", + ) + store.upsert_edge("DISTRIBUTION", merged.asset_id, distribution_id) + + for assertion in merged.assertions: + target_id = concept_id(assertion.target_kind, assertion.target_label) + target_kind = ( + "file_asset_reference" if assertion.target_kind == "file_asset" else assertion.target_kind + ) + store.upsert_node( + target_id, + target_kind, + label=assertion.target_label, + properties={ + "skos_pref_label": assertion.target_label, + "review_status": assertion.review_status, + }, + text=assertion.target_label, + ) + store.upsert_edge( + _EDGE_TYPES[assertion.relation], + merged.asset_id, + target_id, + properties={ + "confidence": assertion.confidence, + "evidence_chunk_sha256": assertion.evidence_chunk_sha256, + "evidence_start": assertion.evidence_start, + "evidence_end": assertion.evidence_end, + "method": assertion.method, + "review_status": assertion.review_status, + }, + ) + return merged diff --git a/src/sdp/file_pilot.py b/src/sdp/file_pilot.py new file mode 100644 index 0000000..a236895 --- /dev/null +++ b/src/sdp/file_pilot.py @@ -0,0 +1,180 @@ +"""Read-only local runner for the semantic file ontology pilot.""" + +from __future__ import annotations + +import argparse +import getpass +import hashlib +import json +from pathlib import Path +from typing import Protocol + +from .config import get_app_config +from .document_semantics import ( + ContextualOrchestratorClient, + EphemeralCredentialRegistry, + TextChunk, + chunk_text, + extract_document_text, +) +from .file_ontology import FileAsset, SemanticAssertion, file_asset_jsonld, upsert_file_asset +from .graph_store import GraphStore, InMemoryGraphStore +from .storage_readers import FilesystemReader + + +class PilotExtractor(Protocol): + def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: ... + + def embed_one(self, text: str) -> list[float]: ... + + +def run_local_pilot( + root: str | Path, + output: str | Path, + *, + name_pattern: str, + extractor: PilotExtractor | None = None, + store: GraphStore | None = None, + max_files: int = 100, + max_bytes: int = 20 * 1024 * 1024, +) -> dict[str, int]: + """Index matching files without moving, deleting, or persisting source text.""" + + if max_files <= 0 or max_bytes <= 0: + raise ValueError("pilot limits must be positive") + reader = FilesystemReader(root) + refs = sorted( + reader.list(name_pattern=name_pattern), + key=lambda ref: ref.object_key.casefold(), + )[:max_files] + assets: dict[str, FileAsset] = {} + asset_statuses: dict[str, str] = {} + file_records: list[dict[str, str]] = [] + + for ref in refs: + try: + data = reader.read(ref, max_bytes=max_bytes) + except ValueError as exc: + status = "too_large" if "maximum size" in str(exc) else "read_failed" + file_records.append( + {"name": ref.name, "distribution_id": ref.distribution.id, "status": status} + ) + continue + except OSError: + file_records.append( + {"name": ref.name, "distribution_id": ref.distribution.id, "status": "read_failed"} + ) + continue + + digest = hashlib.sha256(data).hexdigest() + asset = assets.get(digest) + if asset is None: + document = extract_document_text(ref.name, data) + assertions: list[SemanticAssertion] = [] + status = document.status + if document.status == "extracted" and extractor is not None: + try: + assertions = extractor.extract(ref.name, chunk_text(document.text)) + except (RuntimeError, ValueError): + status = "semantic_failed" + asset = FileAsset( + sha256=digest, + title=Path(ref.name).stem, + media_type=document.media_type, + byte_size=len(data), + modified_at=ref.modified_at, + distributions=[ref.distribution], + assertions=assertions, + ) + assets[digest] = asset + asset_statuses[digest] = status + else: + distributions = {item.id: item for item in asset.distributions} + distributions[ref.distribution.id] = ref.distribution + modified = max( + filter(None, (asset.modified_at, ref.modified_at)), + default=None, + ) + asset = asset.model_copy( + update={"distributions": list(distributions.values()), "modified_at": modified} + ) + assets[digest] = asset + + file_records.append( + { + "name": ref.name, + "asset_id": asset.asset_id, + "distribution_id": ref.distribution.id, + "status": asset_statuses[digest], + } + ) + + graph = store or InMemoryGraphStore( + embedder=extractor.embed_one if extractor is not None else None + ) + indexed = [upsert_file_asset(graph, asset) for asset in assets.values()] + summary = { + "files": len(refs), + "assets": len(indexed), + "distributions": sum(len(asset.distributions) for asset in indexed), + "assertions": sum(len(asset.assertions) for asset in indexed), + } + manifest = { + "profile": "CWL File Knowledge Profile 0.1", + "summary": summary, + "assets": [ + file_asset_jsonld(asset, include_locations=True) + for asset in sorted(indexed, key=lambda item: item.asset_id) + ], + "files": file_records, + } + output_path = Path(output) + output_path.parent.mkdir(parents=True, exist_ok=True) + output_path.write_text( + json.dumps(manifest, ensure_ascii=False, indent=2), + encoding="utf-8", + ) + return summary + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--root", required=True) + parser.add_argument("--output", required=True) + parser.add_argument("--name-regex", required=True) + parser.add_argument("--max-files", type=int, default=100) + parser.add_argument("--orchestrator-url") + parser.add_argument("--semantic-model") + parser.add_argument("--embedding-model") + parser.add_argument("--no-llm", action="store_true") + args = parser.parse_args() + + extractor = None + if not args.no_llm: + config = get_app_config() + orchestrator_url = args.orchestrator_url or config.orchestrator_base_url + if not orchestrator_url: + parser.error( + "--orchestrator-url or KV orchestrator_base_url is required unless --no-llm is used" + ) + token = getpass.getpass("contextual-orchestrator inference token: ") + extractor = ContextualOrchestratorClient( + EphemeralCredentialRegistry({"CONTEXTUAL_ORCHESTRATOR_TOKEN": token}), + base_url=orchestrator_url, + semantic_model=args.semantic_model or config.semantic_model, + embedding_model=args.embedding_model or config.embedding_model, + ) + + summary = run_local_pilot( + args.root, + args.output, + name_pattern=args.name_regex, + extractor=extractor, + max_files=args.max_files, + ) + print(json.dumps(summary, ensure_ascii=False)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/sdp/graph_store.py b/src/sdp/graph_store.py index d4cd8d7..c1ffeaf 100644 --- a/src/sdp/graph_store.py +++ b/src/sdp/graph_store.py @@ -24,7 +24,7 @@ from collections import deque from dataclasses import dataclass, field from datetime import datetime, timezone -from typing import Any, Dict, Iterable, List, Optional, Tuple +from typing import Any, Callable, Dict, Iterable, List, Optional, Tuple from .config import AppConfig, get_app_config, load_bootstrap from .embeddings import cosine_similarity, embed_text @@ -192,9 +192,18 @@ def stats(self) -> Dict[str, int]: class InMemoryGraphStore(GraphStore): - def __init__(self, config: Optional[AppConfig] = None) -> None: + def __init__( + self, + config: Optional[AppConfig] = None, + *, + embedder: Optional[Callable[[str], List[float]]] = None, + ) -> None: self._config = config or get_app_config() self.dimension = self._config.embedding_dimension + self._embedder = embedder or ( + lambda text: embed_text(text, self.dimension) + ) + self._external_embedder = embedder is not None self._nodes: Dict[str, GraphNode] = {} self._edges: List[GraphEdge] = [] self._embeddings: Dict[str, List[float]] = {} @@ -220,7 +229,14 @@ def upsert_node( ) self._nodes[node_id] = node embed_source = text or label or node_id - self._embeddings[node_id] = embed_text(embed_source, self.dimension) + vector = [float(component) for component in self._embedder(embed_source)] + if not vector: + raise ValueError("embedding vector must not be empty") + if self._external_embedder and not self._embeddings: + self.dimension = len(vector) + if len(vector) != self.dimension: + raise ValueError(f"embedding dimension must be {self.dimension}") + self._embeddings[node_id] = vector return node def upsert_edge( @@ -374,7 +390,9 @@ def traverse( def semantic_search( self, query: str, *, kind: Optional[str] = None, limit: int = 5 ) -> List[Dict[str, Any]]: - query_vec = embed_text(query, self.dimension) + query_vec = [float(component) for component in self._embedder(query)] + if len(query_vec) != self.dimension: + raise ValueError(f"embedding dimension must be {self.dimension}") scored: List[Dict[str, Any]] = [] for node_id, vector in self._embeddings.items(): node = self._nodes[node_id] @@ -421,11 +439,20 @@ def _vector_literal(vector: List[float]) -> str: class PostgresGraphStore(GraphStore): """Backend on Postgres with Apache AGE (openCypher) and pgvector (KNN).""" - def __init__(self, dsn: str, config: Optional[AppConfig] = None) -> None: + def __init__( + self, + dsn: str, + config: Optional[AppConfig] = None, + *, + embedder: Optional[Callable[[str], List[float]]] = None, + ) -> None: from sqlalchemy import create_engine self._config = config or get_app_config() self.dimension = self._config.embedding_dimension + self._embedder = embedder or ( + lambda text: embed_text(text, self.dimension) + ) self.graph_name = self._config.graph_name self._engine = create_engine(dsn, pool_pre_ping=True, future=True) @@ -492,7 +519,9 @@ def upsert_node( label = label or node_id props = dict(properties or {}) embed_source = text or label or node_id - vector = embed_text(embed_source, self.dimension) + vector = [float(component) for component in self._embedder(embed_source)] + if len(vector) != self.dimension: + raise ValueError(f"embedding dimension must be {self.dimension}") with self._engine.begin() as conn: self._prepare(conn) self._cypher( @@ -751,7 +780,10 @@ def semantic_search( ) -> List[Dict[str, Any]]: from sqlalchemy import text as sql - vector = _vector_literal(embed_text(query, self.dimension)) + query_vector = [float(component) for component in self._embedder(query)] + if len(query_vector) != self.dimension: + raise ValueError(f"embedding dimension must be {self.dimension}") + vector = _vector_literal(query_vector) stmt = ( "SELECT e.node_id, e.node_kind, n.node_label, " "1 - (e.embedding <=> CAST(:vec AS vector)) AS score " diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index 75f8945..be13099 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -3,16 +3,20 @@ import json from copy import deepcopy from io import BytesIO +from pathlib import Path from types import SimpleNamespace from zipfile import ZIP_DEFLATED, ZipFile import pytest +from fastapi.testclient import TestClient from pypdf import PdfWriter +from sdp import api as api_module +from sdp.api import app from sdp.document_semantics import ( + ContextualOrchestratorClient, EphemeralCredentialRegistry, - OpenAISemanticExtractor, chunk_text, extract_document_text, ) @@ -22,13 +26,18 @@ SemanticAssertion, StorageDistribution, file_asset_jsonld, + get_file_asset, + upsert_file_asset, validate_file_asset, ) +from sdp.file_pilot import run_local_pilot +from sdp.graph_store import InMemoryGraphStore from sdp.storage_readers import AzureBlobReader, FilesystemReader, S3Reader ASSET_SHA256 = "a" * 64 CHUNK_SHA256 = "b" * 64 +client = TestClient(app) def sample_distribution(**overrides: object) -> StorageDistribution: @@ -51,6 +60,7 @@ def sample_assertion(**overrides: object) -> SemanticAssertion: "evidence_chunk_sha256": CHUNK_SHA256, "evidence_start": 3, "evidence_end": 9, + "method": "manual-test", } values.update(overrides) return SemanticAssertion(**values) @@ -100,6 +110,23 @@ def test_file_asset_validation_rejects_unverifiable_evidence(): assert {"assertions.evidence_chunk_sha256", "assertions.evidence_end"} <= paths +def test_semantic_assertion_requires_extraction_provenance(): + values = sample_assertion().model_dump() + values.pop("method") + + with pytest.raises(ValueError, match="method"): + SemanticAssertion(**values) + + +def test_machine_readable_profile_requires_assertion_provenance(): + root = Path(__file__).resolve().parents[1] + profile = (root / "ontology" / "cwl-file-profile.ttl").read_text(encoding="utf-8") + shapes = (root / "ontology" / "cwl-file-shapes.ttl").read_text(encoding="utf-8") + + assert "cwl:extractionMethod a owl:DatatypeProperty" in profile + assert "sh:path cwl:extractionMethod ; sh:minCount 1" in shapes + + @pytest.mark.parametrize( ("provider", "overrides", "expected_path"), [ @@ -286,36 +313,33 @@ def test_unsupported_legacy_format_is_reported_without_guessing(): assert document.text == "" -def fake_openai_transport(captured, *, evidence_quote="효성중공업"): +def fake_orchestrator_transport(captured, *, evidence_quote="효성중공업"): def transport(request, timeout): payload = json.loads(request.data.decode("utf-8")) captured.update(payload) + captured["url"] = request.full_url captured["timeout"] = timeout - assert request.get_header("Authorization") == "Bearer test-key" + assert request.get_header("Authorization") == "Bearer orchestrator-token" return { - "status": "completed", - "output": [ + "choices": [ { - "type": "message", - "content": [ - { - "type": "output_text", - "text": json.dumps( - { - "assertions": [ - { - "relation": "usesSystem", - "target_kind": "system", - "target_label": "C-Cube", - "confidence": 0.93, - "evidence_quote": evidence_quote, - } - ] - }, - ensure_ascii=False, - ), - } - ], + "message": { + "role": "assistant", + "content": json.dumps( + { + "assertions": [ + { + "relation": "usesSystem", + "target_kind": "system", + "target_label": "C-Cube", + "confidence": 0.93, + "evidence_quote": evidence_quote, + } + ] + }, + ensure_ascii=False, + ), + } } ], } @@ -323,38 +347,230 @@ def transport(request, timeout): return transport -def test_openai_extractor_uses_strict_schema_and_persists_only_evidence_reference(): +def test_orchestrator_extractor_uses_strict_schema_and_persists_only_evidence_reference(): captured = {} - extractor = OpenAISemanticExtractor( - EphemeralCredentialRegistry({"OPENAI_API_KEY": "test-key"}), - transport=fake_openai_transport(captured), + extractor = ContextualOrchestratorClient( + EphemeralCredentialRegistry( + {"CONTEXTUAL_ORCHESTRATOR_TOKEN": "orchestrator-token"} + ), + base_url="https://orchestrator.example", + transport=fake_orchestrator_transport(captured), ) assertions = extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) assert captured["store"] is False assert captured["model"] == "gpt-5-mini-2025-08-07" - assert captured["text"]["format"]["type"] == "json_schema" - assert captured["text"]["format"]["strict"] is True + assert captured["response_format"]["type"] == "json_schema" + assert captured["response_format"]["json_schema"]["strict"] is True + assert captured["url"] == "https://orchestrator.example/v1/chat/completions" assert captured["timeout"] == 60 assert assertions[0].relation == "usesSystem" + assert assertions[0].method == "contextual-orchestrator" assert assertions[0].evidence_chunk_sha256 assert assertions[0].evidence_start < assertions[0].evidence_end assert "효성중공업" not in assertions[0].model_dump_json() - assert "test-key" not in repr(captured) + assert "orchestrator-token" not in repr(captured) -def test_openai_extractor_rejects_quote_not_present_in_input(): - extractor = OpenAISemanticExtractor( - EphemeralCredentialRegistry({"OPENAI_API_KEY": "test-key"}), - transport=fake_openai_transport({}, evidence_quote="invented evidence"), +def test_orchestrator_extractor_rejects_quote_not_present_in_input(): + extractor = ContextualOrchestratorClient( + EphemeralCredentialRegistry( + {"CONTEXTUAL_ORCHESTRATOR_TOKEN": "orchestrator-token"} + ), + base_url="https://orchestrator.example", + transport=fake_orchestrator_transport({}, evidence_quote="invented evidence"), ) assert extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) == [] -def test_openai_extractor_fails_closed_without_credential(): - extractor = OpenAISemanticExtractor(EphemeralCredentialRegistry({}), transport=lambda *_: {}) +def test_orchestrator_extractor_fails_closed_without_credential(): + extractor = ContextualOrchestratorClient( + EphemeralCredentialRegistry({}), + base_url="https://orchestrator.example", + transport=lambda *_: {}, + ) with pytest.raises(ValueError, match="credential is unavailable"): extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) + + +def test_orchestrator_client_uses_sync_embeddings_endpoint(): + captured = {} + + def transport(request, timeout): + captured["url"] = request.full_url + captured["payload"] = json.loads(request.data.decode("utf-8")) + assert request.get_header("Authorization") == "Bearer orchestrator-token" + return { + "object": "list", + "data": [ + {"object": "embedding", "index": 1, "embedding": [0.0, 1.0]}, + {"object": "embedding", "index": 0, "embedding": [1.0, 0.0]}, + ], + "model": "text-embedding-3-small", + "usage": {"prompt_tokens": 4, "total_tokens": 4}, + } + + client = ContextualOrchestratorClient( + EphemeralCredentialRegistry( + {"CONTEXTUAL_ORCHESTRATOR_TOKEN": "orchestrator-token"} + ), + base_url="https://orchestrator.example/", + transport=transport, + ) + + assert client.embed(["alpha", "beta"]) == [[1.0, 0.0], [0.0, 1.0]] + assert captured["url"] == "https://orchestrator.example/v1/embeddings" + assert captured["payload"] == { + "model": "text-embedding-3-small", + "input": ["alpha", "beta"], + "metadata": {"service": "semantic-data-portal"}, + } + + +def test_graph_store_can_use_orchestrator_embedding_client(): + embedded = [] + + def embedder(text): + embedded.append(text) + return [1.0, 0.0] if "VOC" in text else [0.0, 1.0] + + store = InMemoryGraphStore(embedder=embedder) + store.upsert_node("voc", "file_asset", label="효성중공업 VOC") + + assert store.semantic_search("VOC query")[0]["node_id"] == "voc" + assert embedded == ["효성중공업 VOC", "VOC query"] + + +def test_local_pilot_deduplicates_content_and_writes_no_raw_text(tmp_path): + root = tmp_path / "input" + root.mkdir() + (root / "효성중공업 VOC.txt").write_text("C-Cube PoC", encoding="utf-8") + (root / "중공업VOC copy.txt").write_text("C-Cube PoC", encoding="utf-8") + output = tmp_path / "manifest.json" + + class FakeExtractor: + def __init__(self): + self.extract_calls = 0 + + def extract(self, filename, chunks): + self.extract_calls += 1 + chunk = chunks[0] + return [ + SemanticAssertion( + relation="usesSystem", + target_kind="system", + target_label="C-Cube", + confidence=0.9, + evidence_chunk_sha256=chunk.sha256, + evidence_start=chunk.start, + evidence_end=chunk.start + len("C-Cube"), + method="fake-extractor", + ) + ] + + def embed_one(self, text): + return [1.0, 0.0] + + extractor = FakeExtractor() + summary = run_local_pilot( + root, + output, + name_pattern=r"효성중공업|중공업VOC", + extractor=extractor, + ) + + manifest = output.read_text(encoding="utf-8") + assert summary["files"] == 2 + assert summary["assets"] == 1 + assert summary["distributions"] == 2 + assert extractor.extract_calls == 1 + assert "C-Cube PoC" not in manifest + assert "evidence_quote" not in manifest + assert "dcat:accessURL" in manifest + + +def test_upsert_file_asset_merges_same_content_distributions_and_projects_assertions(): + store = InMemoryGraphStore() + first = sample_asset() + second = sample_asset( + distributions=[ + sample_distribution( + id="dist-copy", + locator="file:///D:/Documents/report-copy.docx", + ) + ] + ) + + upsert_file_asset(store, first) + merged = upsert_file_asset(store, second) + + assert len(merged.distributions) == 2 + assert get_file_asset(store, first.asset_id) == merged + graph = store.traverse(first.asset_id, direction="out", max_depth=1) + assert {edge["edge_type"] for edge in graph["edges"]} >= { + "DISTRIBUTION", + "USES_SYSTEM", + } + node = store.get_node(first.asset_id) + assert node is not None + assert "C-Cube PoC" not in json.dumps(node.properties, ensure_ascii=False) + + +def sample_request(actor: str) -> dict[str, object]: + return {**sample_asset().model_dump(mode="json"), "actor": actor} + + +def test_file_asset_api_requires_policy_and_redacts_jsonld_locator(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + + denied = client.post("/file-assets", json=sample_request("analyst")) + created = client.post("/file-assets", json=sample_request("admin")) + + assert denied.status_code == 403 + assert created.status_code == 200 + asset_id = created.json()["asset_id"] + detail = client.get(f"/file-assets/{asset_id}", params={"actor": "analyst"}) + graph_node = client.get(f"/graph/nodes/{asset_id}") + exported = client.get(f"/file-assets/{asset_id}/jsonld", params={"actor": "analyst"}) + assert detail.status_code == 200 + assert "file:///" not in detail.text + assert graph_node.status_code == 200 + assert "file:///" not in graph_node.text + assert exported.status_code == 200 + assert "dcat:accessURL" not in exported.text + + location_denied = client.get( + f"/file-assets/{asset_id}/jsonld", + params={"actor": "analyst", "include_locations": True}, + ) + location_allowed = client.get( + f"/file-assets/{asset_id}/jsonld", + params={"actor": "admin", "include_locations": True}, + ) + detail_location_allowed = client.get( + f"/file-assets/{asset_id}", + params={"actor": "admin", "include_locations": True}, + ) + assert location_denied.status_code == 403 + assert location_allowed.status_code == 200 + assert location_allowed.json()["dcat:distribution"][0]["dcat:accessURL"].startswith("file:") + assert detail_location_allowed.status_code == 200 + assert detail_location_allowed.json()["distributions"][0]["locator"].startswith("file:") + + +def test_file_asset_api_exposes_validation_and_404(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + created = client.post("/file-assets", json=sample_request("admin")) + asset_id = created.json()["asset_id"] + + report = client.get(f"/file-assets/{asset_id}/validate", params={"actor": "analyst"}) + missing = client.get(f"/file-assets/urn:sha256:{'f' * 64}", params={"actor": "analyst"}) + + assert report.status_code == 200 + assert report.json()["conforms"] is True + assert missing.status_code == 404 diff --git a/tests/test_graph_engine.py b/tests/test_graph_engine.py index 6117c7b..c5150ed 100644 --- a/tests/test_graph_engine.py +++ b/tests/test_graph_engine.py @@ -200,12 +200,18 @@ def test_config_loads_from_kv_mapping(monkeypatch): override = { "cors_allow_origins": ["https://portal.example.org"], "embedding_dimension": 128, + "orchestrator_base_url": "https://orchestrator.example", + "semantic_model": "semantic-model-v1", + "embedding_model": "embedding-model-v1", } monkeypatch.setattr(config_module, "_load_from_kv_table", lambda bootstrap: override) config_module.reset_config_cache() try: cfg = config_module.get_app_config() assert cfg.cors_allow_origins == ["https://portal.example.org"] + assert cfg.orchestrator_base_url == "https://orchestrator.example" + assert cfg.semantic_model == "semantic-model-v1" + assert cfg.embedding_model == "embedding-model-v1" assert cfg.source == "config_entries" finally: config_module.reset_config_cache() From b6d5852d75c62cceafa15d44a03064d9d6b952bd Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 11:21:47 +0900 Subject: [PATCH 08/15] docs: document orchestrated file ontology pilot --- README.md | 35 ++++++++++++ docs/implementation-compliance.md | 21 +++++++- .../2026-07-21-semantic-file-ontology.md | 54 ++++++++++--------- ...026-07-21-semantic-file-ontology-design.md | 16 +++--- 4 files changed, 91 insertions(+), 35 deletions(-) diff --git a/README.md b/README.md index 9eaafc9..6b400f3 100644 --- a/README.md +++ b/README.md @@ -105,6 +105,40 @@ SDP_DATABASE_DSN='postgresql+psycopg://sdp_graph_app:@loca - `GET /ontology/term/{term}/graph` — 개념 그래프 (그래프 스토어 백엔드) - `POST /search/semantic` — pgvector KNN 시맨틱 검색 (kind 필터) +### 파일 지식 온톨로지 + +`CWL File Knowledge Profile 0.1`은 파일 내용의 SHA-256 정체성과 물리 저장 위치를 +분리합니다. 같은 bytes가 로컬/Synology 동기화 폴더, S3, S3 호환 저장소, Azure Blob에 +복제되어도 하나의 `FileAsset`과 여러 DCAT `Distribution`으로 표현됩니다. 기계 판독 +프로파일과 shape는 `ontology/cwl-file-profile.ttl`, `ontology/cwl-file-shapes.ttl`에 있습니다. + +- `POST /file-assets` — 관리자 정책을 통과한 자산·후보 주장 적재 +- `GET /file-assets/{asset_id}` — 자산과 의미 관계 조회 +- `GET /file-assets/{asset_id}/jsonld` — 기본 locator 비공개 JSON-LD +- `GET /file-assets/{asset_id}/validate` — SHACL 호환 검증 리포트 +- 지원 reader: `filesystem`(로컬/UNC/Synology 포함), `s3`, `s3_compatible`, `azure_blob` +- 지원 본문 추출: TXT/Markdown/CSV/JSON/XML, DOCX/PPTX/XLSX, PDF + +파일 의미 추출과 embedding은 OpenAI를 직접 호출하지 않고 +[`ContextualWisdomLab/contextual-orchestrator`](https://github.com/ContextualWisdomLab/contextual-orchestrator)만 +사용합니다. 의미 추출은 `/v1/chat/completions`, embedding은 orchestrator에 추가된 동기 +`/v1/embeddings`를 사용합니다. `orchestrator_base_url`, `semantic_model`, +`embedding_model`은 `config_entries` KV 설정이고, inference token은 주입된 credential +registry에서만 가져옵니다. OpenAI/provider key는 포털에 두지 않습니다. + +읽기 전용 로컬 파일럿은 다음처럼 실행합니다. `--no-llm`은 파일 이동·삭제나 네트워크 +호출 없이 중복·추출 상태만 확인합니다. + +```powershell +$env:PYTHONPATH='src' +py -m sdp.file_pilot --root '' --output '' --name-regex '효성중공업|중공업VOC' --max-files 12 --no-llm +``` + +LLM을 사용할 때는 `--orchestrator-url`을 주거나 KV의 `orchestrator_base_url`을 사용하며, +inference token은 숨김 prompt로만 입력합니다. manifest에는 원문 조각·근거 인용문·API +응답·credential을 저장하지 않습니다. 실제 파일명과 locator가 들어가므로 출력은 로컬의 +Git 제외 경로에만 보관합니다. + ### Catalog / governance / enterprise (기존) - `GET /health` @@ -154,6 +188,7 @@ PYTHONPATH=src python -m sdp.demo_smoke | Browse/Query | `src/sdp/browse.py`, `/browse/*` | | Policy Service | `src/sdp/policy.py`, `/policy/decision` | | LLM Orchestrator | `src/sdp/orchestrator.py`, `/llm/*` | +| File Knowledge Profile | `src/sdp/file_ontology.py`, `src/sdp/storage_readers.py`, `src/sdp/document_semantics.py`, `src/sdp/file_pilot.py`, `/file-assets/*` | | JSON-LD Export | `/catalog/datasets/{id}/jsonld` | | Enterprise Core Contracts | `src/sdp_core/contracts.py`, `src/sdp_core/readiness.py`, `src/sdp_core/demo_seed.py`, `src/sdp_core/enterprise.py`, `src/sdp_core/rbac.py`, `src/sdp/enterprise_evidence.py`, `src/sdp/semantic_validation.py`, `src/sdp/steward_review.py`, `src/sdp/observability.py`, `/enterprise/*` | diff --git a/docs/implementation-compliance.md b/docs/implementation-compliance.md index 7b6081c..79056d8 100644 --- a/docs/implementation-compliance.md +++ b/docs/implementation-compliance.md @@ -133,13 +133,30 @@ - `tests/test_api.py::test_enterprise_console_renders_operator_surface` - `tests/test_api.py::test_enterprise_demo_smoke_summary_is_ready` -## 7) 다음 단계 (현재 브랜치에서 미반영 권고) +## 7) File Knowledge / Hybrid Ontology + +- 표준 프로파일: RDF/OWL 2/SKOS/SHACL/DCAT 3/DCMI/PROV-O/SPDX/JSON-LD 기반 `ontology/cwl-file-profile.ttl`, `ontology/cwl-file-shapes.ttl`. +- 정체성/위치 분리: SHA-256 `FileAsset` 하나에 filesystem/S3/S3-compatible/Azure Blob `Distribution` 여러 개를 결합한다. +- 읽기 전용 수집: `src/sdp/storage_readers.py`는 list/read만 제공하고 이동·삭제·원격 mutation 기능이 없다. +- 의미 근거: `SemanticAssertion`은 confidence, proposed review status, chunk SHA-256, 문자 offset을 보존하며 원문 인용은 저장하지 않는다. +- LLM 경계: `src/sdp/document_semantics.py::ContextualOrchestratorClient`가 `/v1/chat/completions`와 `/v1/embeddings`만 호출한다. 포털에는 OpenAI/provider key가 없다. +- provider 중립성: Synology는 filesystem 배치일 뿐 필수 구성요소가 아니며, 파일 정체성은 저장소 URL과 독립적이다. +- 정책/API: `POST /file-assets`, `GET /file-assets/{asset_id}`, `/jsonld`, `/validate`가 기존 policy 및 graph store를 재사용한다. +- 파일럿: `src/sdp/file_pilot.py`가 content deduplication, 안전한 문서 추출, graph projection, 로컬 전용 manifest를 수행한다. +- 증빙 테스트: + - `tests/test_file_knowledge.py::test_orchestrator_extractor_uses_strict_schema_and_persists_only_evidence_reference` + - `tests/test_file_knowledge.py::test_orchestrator_client_uses_sync_embeddings_endpoint` + - `tests/test_file_knowledge.py::test_local_pilot_deduplicates_content_and_writes_no_raw_text` + - `tests/test_file_knowledge.py::test_file_asset_api_requires_policy_and_redacts_jsonld_locator` + - `tests/test_graph_engine.py::test_config_loads_from_kv_mapping` + +## 8) 다음 단계 (현재 브랜치에서 미반영 권고) 1. 조직 정책 기준으로 `search` 및 `list` 에 대한 사용 권한/발견성 정책을 명시적으로 강화 2. API level 감사 이벤트 보존 기간 및 위변조 방지(로그 저장소 정책) 적용 3. OpenCode/PR 리뷰 증적 저장(`PR`, `review`, `merge` 로그)과 main 병합 완료 상태 정기 기록 -## 8) 구현 완료 증적(현재 HEAD 기준) +## 9) 구현 완료 증적(현재 HEAD 기준) - 대상 브랜치: `codex/sdp-enterprise-foundation` - 기준: `origin/main` 병합 후 현재 브랜치 HEAD diff --git a/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md b/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md index f040d4f..361228b 100644 --- a/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md +++ b/docs/superpowers/plans/2026-07-21-semantic-file-ontology.md @@ -2,9 +2,9 @@ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. -**Goal:** Add a standards-aligned, provider-neutral file knowledge catalog with read-only filesystem/S3/S3-compatible/Azure Blob ingestion, evidence-bound OpenAI semantic extraction, and a Hyosung Heavy Industries VOC pilot. +**Goal:** Add a standards-aligned, provider-neutral file knowledge catalog with read-only filesystem/S3/S3-compatible/Azure Blob ingestion, evidence-bound semantic extraction and embeddings routed exclusively through `contextual-orchestrator`, and a Hyosung Heavy Industries VOC pilot. -**Architecture:** Keep Apache AGE and pgvector as the persistence/search implementation. Represent content-addressed `FileAsset` nodes separately from one-or-many DCAT `Distribution` nodes, project validated LLM assertions into the existing graph, and export JSON-LD. Read document bytes through small injected readers; raw text never enters graph properties or Git. +**Architecture:** Keep Apache AGE and pgvector as the persistence/search implementation. Represent content-addressed `FileAsset` nodes separately from one-or-many DCAT `Distribution` nodes, project validated LLM assertions into the existing graph, and export JSON-LD. Read document bytes through small injected readers; raw text never enters graph properties or Git. Route all LLM and embedding traffic through `ContextualWisdomLab/contextual-orchestrator`; the portal never holds an OpenAI key. **Tech Stack:** Python 3.10+, FastAPI, Pydantic 2, existing graph store, stdlib `pathlib`/`zipfile`/`urllib`, pypdf 6.14.2, W3C RDF/OWL/SKOS/SHACL/DCAT/PROV/JSON-LD vocabularies. @@ -13,12 +13,12 @@ - File operations are read-only; no move, delete, copy, or remote mutation. - Providers are exactly `filesystem`, `s3`, `s3_compatible`, and `azure_blob`; Synology is a filesystem deployment. - Object bytes are capped at 20 MiB; LLM text is chunked at 6,000 characters with 300-character overlap and capped at 24,000 characters per file. -- The pinned default model is `gpt-5-mini-2025-08-07`; Requests API calls set `store: false` and a 60-second timeout. -- OpenAI API keys and cloud credentials come from an injected credential registry/client, never `os.getenv()`, graph properties, logs, CLI arguments, or GitHub fixtures. +- The pinned semantic model is `gpt-5-mini-2025-08-07`; orchestrator chat-completions calls set `store: false`, strict `response_format`, and a 60-second timeout. The default embedding model is `text-embedding-3-small`. +- The portal receives only an orchestrator base URL and inference token through KV config and an injected credential registry. OpenAI/provider keys stay inside orchestrator; no credential enters `os.getenv()`, graph properties, logs, CLI arguments, or GitHub fixtures. - LLM assertions remain `proposed`; structural validation does not imply steward approval. - Raw chunks and evidence quotations are not persisted. Persist only chunk SHA-256 and character offsets after verifying the quotation exists. - Existing policy authorization, AGE/pgvector stores, semantic search, and graph traversal are reused. -- CI uses fake transports and fake cloud clients; it makes no OpenAI or cloud call. +- CI uses fake orchestrator transports and fake cloud clients; it makes no LLM or cloud call. - No OCR, HWP, DOC, XLS parser, mandatory cloud SDK, new database, or automatic ontology approval. --- @@ -86,7 +86,7 @@ class SemanticAssertion(BaseModel): evidence_chunk_sha256: str evidence_start: int = Field(ge=0) evidence_end: int = Field(gt=0) - method: str = "openai" + method: str review_status: Literal["proposed", "approved", "rejected"] = "proposed" class FileAsset(BaseModel): @@ -236,46 +236,48 @@ git add pyproject.toml requirements*.txt src/sdp/document_semantics.py tests/tes git commit -m "feat: extract text from pilot document formats" ``` -### Task 4: Evidence-bound OpenAI semantic extraction +### Task 4: Evidence-bound contextual-orchestrator semantic extraction and embeddings **Files:** - Modify: `src/sdp/document_semantics.py` - Modify: `tests/test_file_knowledge.py` **Interfaces:** -- Produces: `CredentialRegistry`, `EphemeralCredentialRegistry`, `OpenAISemanticExtractor.extract()`. +- Produces: `CredentialRegistry`, `EphemeralCredentialRegistry`, `ContextualOrchestratorClient.extract()`, `ContextualOrchestratorClient.embed()`. - Consumes: `TextChunk`, injected credential registry, injected JSON HTTP transport. - [ ] **Step 1: Write a failing fake-transport test** ```python -def test_openai_extractor_uses_strict_schema_and_persists_only_evidence_reference(): +def test_orchestrator_extractor_uses_strict_schema_and_persists_only_evidence_reference(): captured = {} - extractor = OpenAISemanticExtractor( - EphemeralCredentialRegistry({"OPENAI_API_KEY": "test-key"}), - transport=fake_openai_transport(captured, evidence_quote="C-Cube"), + extractor = ContextualOrchestratorClient( + EphemeralCredentialRegistry({"CONTEXTUAL_ORCHESTRATOR_TOKEN": "test-token"}), + base_url="https://orchestrator.example", + transport=fake_orchestrator_transport(captured, evidence_quote="C-Cube"), ) assertions = extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) assert captured["store"] is False - assert captured["text"]["format"]["type"] == "json_schema" + assert captured["response_format"]["type"] == "json_schema" + assert captured["url"].endswith("/v1/chat/completions") assert assertions[0].evidence_chunk_sha256 assert "C-Cube" not in assertions[0].model_dump_json() - assert "test-key" not in repr(captured) + assert "test-token" not in repr(captured) ``` - [ ] **Step 2: Run the test and confirm failure** -Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k openai -q` +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k orchestrator -q` Expected: FAIL because the extractor does not exist. -- [ ] **Step 3: Implement the minimal Responses API client** +- [ ] **Step 3: Add the missing orchestrator sync endpoint and implement the portal client** -Use `urllib.request.Request("https://api.openai.com/v1/responses", data=encoded_payload, headers=headers, method="POST")`, an injected `transport(request, timeout) -> dict`, model `gpt-5-mini-2025-08-07`, `store: false`, and strict JSON Schema fields `relation`, `target_kind`, `target_label`, `confidence`, `evidence_quote`. Reject empty/missing credential, non-completed responses, malformed JSON, unknown relations, and quotations not found verbatim in the input chunk. Convert a verified quote to chunk hash plus global offsets, then discard the quote. Merge duplicate `(relation, target_kind, casefold(target_label))` candidates by highest confidence. +In `ContextualWisdomLab/contextual-orchestrator`, add authenticated `POST /v1/embeddings`. Reuse the existing embedding batch core, poll asynchronous backends for at most 30 seconds, and translate a completed batch document into the OpenAI-compatible `object/data/model/usage` response. In the portal, send strict-schema semantic requests to `/v1/chat/completions` and vectors to `/v1/embeddings` through one injected `ContextualOrchestratorClient`. Reject missing orchestrator token, malformed output, unknown relations, and quotations not found verbatim in the input chunk. Convert a verified quote to chunk hash plus global offsets, then discard the quote. Merge duplicate `(relation, target_kind, casefold(target_label))` candidates by highest confidence. -- [ ] **Step 4: Run OpenAI tests** +- [ ] **Step 4: Run orchestrator contract and portal client tests** -Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k openai -q` +Run: `$env:PYTHONPATH='src'; py -m pytest tests/test_file_knowledge.py -k orchestrator -q` Expected: PASS with no network call. @@ -283,7 +285,7 @@ Expected: PASS with no network call. ```bash git add src/sdp/document_semantics.py tests/test_file_knowledge.py -git commit -m "feat: extract evidence-bound semantics with OpenAI" +git commit -m "feat: route file semantics through contextual orchestrator" ``` ### Task 5: Graph projection and policy-protected API @@ -352,7 +354,7 @@ git commit -m "feat: project file knowledge into graph API" **Interfaces:** - Produces: `run_local_pilot()`, `python -m sdp.file_pilot`. -- Consumes: filesystem reader, document extraction, optional OpenAI extractor, graph projection, JSON-LD export. +- Consumes: filesystem reader, document extraction, optional contextual-orchestrator client, graph projection, JSON-LD export. - [ ] **Step 1: Write a failing end-to-end local runner test** @@ -376,7 +378,7 @@ Expected: FAIL because the runner does not exist. - [ ] **Step 3: Implement the runner and CLI** -`run_local_pilot()` groups by SHA-256, merges distributions, extracts/chunks each unique asset once, applies the optional extractor, validates and projects into an `InMemoryGraphStore`, and writes a UTF-8 JSON manifest containing summary, asset JSON-LD with locators, per-file status, and no raw text/quote/API response. CLI arguments are `--root`, `--output`, `--name-regex`, `--model`, `--max-files`, and `--no-llm`; with LLM enabled it obtains the API key via `getpass.getpass()` and an `EphemeralCredentialRegistry`, never a CLI argument or environment variable. +`run_local_pilot()` groups by SHA-256, merges distributions, extracts/chunks each unique asset once, applies the optional orchestrator client, validates and projects into an `InMemoryGraphStore` using `client.embed_one`, and writes a UTF-8 JSON manifest containing summary, asset JSON-LD with locators, per-file status, and no raw text/quote/API response. CLI arguments are `--root`, `--output`, `--name-regex`, `--orchestrator-url`, `--semantic-model`, `--embedding-model`, `--max-files`, and `--no-llm`; with LLM enabled it obtains the orchestrator inference token via `getpass.getpass()` and an `EphemeralCredentialRegistry`, never a CLI argument or environment variable. - [ ] **Step 4: Run tests and the 12-file no-LLM preflight** @@ -438,9 +440,9 @@ codegraph affected src/sdp/file_ontology.py src/sdp/storage_readers.py src/sdp/d Expected: index is current and affected tests include the new file knowledge tests plus API/graph tests. -- [ ] **Step 4: Run the live OpenAI pilot after secure key entry** +- [ ] **Step 4: Run the live contextual-orchestrator pilot after secure token entry** -Run the Task 6 pilot command without `--no-llm`, enter the project-scoped key only at the hidden prompt, and verify the output reports proposed assertions with evidence hashes/offsets and no raw chunks. If no key is available, stop at the verified no-LLM manifest and request secure local entry; do not weaken credential handling. +Run the Task 6 pilot command without `--no-llm`, enter only the orchestrator inference token at the hidden prompt, and verify the output reports proposed assertions with evidence hashes/offsets and no raw chunks. Provider/OpenAI keys remain exclusively in orchestrator. If the URL or token is unavailable, stop at the verified no-LLM manifest and request secure local entry; do not weaken credential handling. - [ ] **Step 5: Commit final docs and verification evidence** @@ -454,7 +456,7 @@ Expected: clean worktree. ## Self-Review -- Spec coverage: standards profile, provider-neutral identity/location split, four providers, safe content extraction, OpenAI evidence binding, SHACL-compatible validation, graph/API reuse, privacy, and 12-file pilot each have a task. +- Spec coverage: standards profile, provider-neutral identity/location split, four providers, safe content extraction, orchestrator-only LLM/embedding routing, evidence binding, SHACL-compatible validation, graph/API reuse, privacy, and 12-file pilot each have a task. - Placeholder scan: no TBD/TODO/future implementation placeholder is used; exclusions are explicit scope decisions. - Type consistency: Tasks 2–7 consume the exact `StorageDistribution`, `SemanticAssertion`, `FileAsset`, `ObjectRef`, `TextChunk`, and extractor names produced in earlier tasks. -- Dependency scope: only pypdf is added because the approved pilot includes two PDFs; cloud SDKs and the OpenAI SDK remain injected/stdlib. +- Dependency scope: only pypdf is added because the approved pilot includes two PDFs; cloud and LLM SDKs remain injected/stdlib. diff --git a/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md b/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md index b5f7a44..5e174c4 100644 --- a/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md +++ b/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md @@ -9,11 +9,11 @@ ## 승인된 범위 - 기존 Apache AGE 그래프와 pgvector 의미 검색을 재사용한다. -- 파일 내용은 로컬에서 추출하며, 사용자가 승인한 최소 원문 조각만 OpenAI Responses API로 전송한다. +- 파일 내용은 로컬에서 추출하며, 사용자가 승인한 최소 원문 조각만 `contextual-orchestrator`의 LLM 호환 API로 전송한다. 포털은 OpenAI를 직접 호출하지 않는다. - 파일을 이동·삭제·복사하지 않는다. 수집은 읽기 전용이다. - Synology는 필수 기반시설이 아니라 `filesystem` 저장소의 한 배치 형태다. - 저장소 공급자는 `filesystem`, `s3`, `s3_compatible`, `azure_blob`을 지원한다. -- 공급자 자격증명과 OpenAI API key는 그래프·문서·로그·GitHub에 저장하지 않는다. +- 공급자 자격증명과 OpenAI API key는 `contextual-orchestrator`의 credential registry 경계 안에만 둔다. 포털은 orchestrator inference token만 주입받으며 어떤 자격증명도 그래프·문서·로그·GitHub에 저장하지 않는다. - LLM 결과는 근거 조각, 신뢰도, `proposed` 검토 상태를 가진 후보 주장으로 저장한다. LLM이 온톨로지의 정답을 직접 확정하지 않는다. - OCR, HWP, legacy DOC/XLS 파서는 1차 파일럿에 포함하지 않는다. 실제 12개 파일은 PDF, DOCX, PPTX, XLSX로 구성되어 있다. @@ -91,13 +91,13 @@ SDK를 포털의 필수 의존성으로 추가하지 않는다. 실제 클라우 3. DOCX/PPTX/XLSX는 Python 표준 `zipfile`·XML parser로 텍스트를 추출한다. 4. PDF는 `pypdf`로 텍스트를 추출한다. 이미지 전용 PDF는 `needs_ocr`로 보고하고 전송하지 않는다. 5. 텍스트를 고정 크기 조각으로 나누고 파일당 최대 전송 문자 수를 적용한다. -6. OpenAI Responses API의 strict JSON Schema 출력으로 의미 후보와 짧은 근거 인용을 받는다. +6. `contextual-orchestrator`의 `POST /v1/chat/completions`에 `response_format=json_schema`, `store=false`를 보내 의미 후보와 짧은 근거 인용을 받는다. orchestrator가 provider와 OpenAI 자격증명을 소유한다. 7. 근거 인용이 입력 조각에 실제로 존재하는지 확인한 뒤 인용문 대신 chunk SHA-256과 문자 offset만 남긴다. 8. 여러 조각의 후보를 `(관계, 정규화 label)` 기준으로 합치고 가장 높은 신뢰도와 그 근거 참조를 유지한다. 9. SHACL 호환 검증을 통과한 후보만 그래프에 `proposed` assertion으로 기록한다. -10. 파일 node의 embedding text는 제목과 제안된 개념 label만 사용한다. 원문과 근거 인용문은 저장하지 않는다. +10. 파일 node의 embedding text는 제목과 제안된 개념 label만 사용하고 `contextual-orchestrator`의 동기 `POST /v1/embeddings`로 벡터화한다. 원문과 근거 인용문은 저장하지 않는다. -OpenAI API key는 extractor에 credential registry lookup으로 주입한다. 애플리케이션 코드가 `os.getenv()`로 key를 읽지 않는다. HTTP transport는 Python 표준 라이브러리를 사용하고 요청에 `store: false`와 pinned model id를 포함한다. CI에서는 HTTP 호출을 fake transport로 대체한다. +포털에는 orchestrator base URL과 model id를 KV application config로, inference token을 주입된 credential registry로 공급한다. 포털 코드가 OpenAI key나 `os.getenv()`를 읽지 않는다. HTTP transport는 Python 표준 라이브러리를 사용하고 LLM 요청에 `store: false`와 pinned model id를 포함한다. orchestrator의 동기 embedding endpoint는 기존 `/v1/batch/embeddings` 코어의 분할·provider backend·비용 원장을 재사용하고 최대 30초 동안 완료를 기다린 뒤 OpenAI 호환 응답을 반환한다. CI에서는 양쪽 HTTP 호출을 fake transport로 대체한다. ## API와 정책 @@ -132,7 +132,7 @@ OpenAI API key는 extractor에 credential registry lookup으로 주입한다. - 파일 크기 제한을 넘으면 `too_large`로 기록하고 읽지 않는다. - 지원하지 않는 형식은 `unsupported_format`, 추출 문자가 없으면 `needs_ocr`로 기록한다. - API timeout, rate limit, 불완전 응답은 해당 파일을 `extraction_failed`로 남기며 기존 그래프를 덮어쓰지 않는다. -- storage credential과 OpenAI key가 없으면 fail closed 한다. +- storage credential 또는 orchestrator inference token이 없으면 fail closed 한다. OpenAI key는 포털에 존재하지 않는다. - 쓰기·삭제·이동·원격 object mutation 기능은 구현하지 않는다. ## 검증 기준 @@ -140,7 +140,8 @@ OpenAI API key는 extractor에 credential registry lookup으로 주입한다. - 같은 bytes의 서로 다른 locator가 한 `FileAsset`으로 합쳐진다. - 각 공급자 reader가 list/stat/read 계약을 만족한다. - 지원 문서에서 텍스트가 추출되고 원문은 graph property에 남지 않는다. -- OpenAI 요청은 strict schema, `store: false`, 최소 조각만 포함한다. +- orchestrator LLM 요청은 strict schema, `store: false`, 최소 조각만 포함하고 OpenAI 직접 URL을 사용하지 않는다. +- embedding 요청은 orchestrator의 동기 `/v1/embeddings`만 사용하며 벡터 순서·차원을 검증한다. - 허용되지 않은 관계나 검증 가능한 근거 참조가 없는 후보는 validation에서 거부된다. - JSON-LD가 DCAT/DCTERMS/PROV/SKOS/SPDX/CWL context와 locator redaction을 지킨다. - 기존 전체 pytest suite와 새 단일 파일 지식 테스트가 통과한다. @@ -154,3 +155,4 @@ OpenAI API key는 extractor에 credential registry lookup으로 주입한다. - cloud object 쓰기 또는 lifecycle 관리 - 자동 온톨로지 승인 - 원문·embedding의 GitHub 업로드 +- `semantic-data-portal`에서 OpenAI 또는 다른 LLM provider를 직접 호출하는 경로 From 5dbdbd3e869cb4a198dcf892810fc772ccd02814 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 12:31:26 +0900 Subject: [PATCH 09/15] fix: harden standards file knowledge integration --- README.md | 27 +- docs/implementation-compliance.md | 7 +- ...026-07-21-semantic-file-ontology-design.md | 8 +- ontology/cwl-file-profile.ttl | 5 +- ontology/cwl-file-shapes.ttl | 70 +++- pyproject.toml | 4 + requirements-dev.txt | 34 +- requirements-test.in | 1 + requirements-test.txt | 34 +- requirements.txt | 34 ++ src/sdp/api.py | 252 +++++++++--- src/sdp/document_semantics.py | 98 ++++- src/sdp/file_ontology.py | 191 +++++++-- src/sdp/file_pilot.py | 1 + src/sdp/graph_models.py | 21 +- src/sdp/graph_store.py | 92 ++++- src/sdp/policy.py | 80 ++++ src/sdp/resources/cwl-file-profile.ttl | 42 ++ src/sdp/resources/cwl-file-shapes.ttl | 73 ++++ src/sdp/storage_readers.py | 6 +- tests/test_file_knowledge.py | 383 +++++++++++++++++- tests/test_graph_engine.py | 46 ++- tests/test_graph_security.py | 80 +++- 23 files changed, 1402 insertions(+), 187 deletions(-) create mode 100644 src/sdp/resources/cwl-file-profile.ttl create mode 100644 src/sdp/resources/cwl-file-shapes.ttl diff --git a/README.md b/README.md index 6b400f3..fbccf1e 100644 --- a/README.md +++ b/README.md @@ -110,12 +110,13 @@ SDP_DATABASE_DSN='postgresql+psycopg://sdp_graph_app:@loca `CWL File Knowledge Profile 0.1`은 파일 내용의 SHA-256 정체성과 물리 저장 위치를 분리합니다. 같은 bytes가 로컬/Synology 동기화 폴더, S3, S3 호환 저장소, Azure Blob에 복제되어도 하나의 `FileAsset`과 여러 DCAT `Distribution`으로 표현됩니다. 기계 판독 -프로파일과 shape는 `ontology/cwl-file-profile.ttl`, `ontology/cwl-file-shapes.ttl`에 있습니다. +프로파일과 shape는 `ontology/cwl-file-profile.ttl`, `ontology/cwl-file-shapes.ttl`에 있으며, +ingest/validate 때 pySHACL로 실제 실행됩니다. - `POST /file-assets` — 관리자 정책을 통과한 자산·후보 주장 적재 - `GET /file-assets/{asset_id}` — 자산과 의미 관계 조회 - `GET /file-assets/{asset_id}/jsonld` — 기본 locator 비공개 JSON-LD -- `GET /file-assets/{asset_id}/validate` — SHACL 호환 검증 리포트 +- `GET /file-assets/{asset_id}/validate` — pySHACL 검증 리포트 - 지원 reader: `filesystem`(로컬/UNC/Synology 포함), `s3`, `s3_compatible`, `azure_blob` - 지원 본문 추출: TXT/Markdown/CSV/JSON/XML, DOCX/PPTX/XLSX, PDF @@ -124,7 +125,27 @@ SDP_DATABASE_DSN='postgresql+psycopg://sdp_graph_app:@loca 사용합니다. 의미 추출은 `/v1/chat/completions`, embedding은 orchestrator에 추가된 동기 `/v1/embeddings`를 사용합니다. `orchestrator_base_url`, `semantic_model`, `embedding_model`은 `config_entries` KV 설정이고, inference token은 주입된 credential -registry에서만 가져옵니다. OpenAI/provider key는 포털에 두지 않습니다. +registry에서만 가져옵니다. `embedding_dimension`은 `/v1/embeddings`의 `dimensions`로 +전달되어 pgvector 차원과 일치해야 합니다. 운영 graph store도 같은 orchestrator client의 +`embed_one`을 ingest와 검색에 주입합니다. OpenAI/provider key는 포털에 두지 않습니다. + +`/file-assets/*`는 요청 본문이나 query의 `actor`를 신뢰하지 않습니다. 검증된 OIDC Bearer +토큰에서 subject/role/tenant를 도출합니다. `FileAsset.tenant_id`와 중앙 policy decision을 +대조해 tenant 경계를 적용하며, locator 포함 응답은 같은 tenant의 `admin` 또는 +`platform-admin`만 요청할 수 있습니다. 같은 SHA-256이나 파일 관계 대상이 다른 tenant에 +이미 속하면 ingest를 거부해 distribution/assertion을 섞지 않습니다. `/graph/nodes`, +`/graph/edges`, `/graph/query`, `/search/semantic`, `/ontology/concepts`도 동일한 Bearer +context를 요구하고 body `actor`를 거부합니다. 일반 graph API는 governed file +node/edge를 수정할 수 없으며, traversal/search 결과는 tenant로 필터링되고 locator는 +항상 redaction됩니다. + +GitHub Secret에 값을 저장하는 것만으로는 런타임 주입이 되지 않습니다. 배포 호스트는 +secret manager에서 token을 읽는 `CredentialRegistry` 구현을 만든 뒤 +`sdp.api.create_app(registry)`로 ASGI 앱을 구성해야 합니다. KV에 +`orchestrator_base_url`이 있는데 `CONTEXTUAL_ORCHESTRATOR_TOKEN`이 주입되지 않으면 +lifespan/startup이 fail-closed하며, registry 교체 시 credential을 캡처한 graph store도 +폐기·재생성됩니다. TTL profile/shape는 `sdp/resources/*.ttl` package data로 배포되므로 +wheel과 공식 컨테이너에서도 pySHACL 검증이 동일하게 동작합니다. 읽기 전용 로컬 파일럿은 다음처럼 실행합니다. `--no-llm`은 파일 이동·삭제나 네트워크 호출 없이 중복·추출 상태만 확인합니다. diff --git a/docs/implementation-compliance.md b/docs/implementation-compliance.md index 79056d8..4425f61 100644 --- a/docs/implementation-compliance.md +++ b/docs/implementation-compliance.md @@ -135,13 +135,16 @@ ## 7) File Knowledge / Hybrid Ontology -- 표준 프로파일: RDF/OWL 2/SKOS/SHACL/DCAT 3/DCMI/PROV-O/SPDX/JSON-LD 기반 `ontology/cwl-file-profile.ttl`, `ontology/cwl-file-shapes.ttl`. +- 표준 프로파일: RDF/OWL 2/SKOS/SHACL/DCAT 3/DCMI/PROV-O/SPDX/JSON-LD 기반 `ontology/cwl-file-profile.ttl`, `ontology/cwl-file-shapes.ttl`; `validate_file_asset`가 pySHACL로 shape를 실행한다. - 정체성/위치 분리: SHA-256 `FileAsset` 하나에 filesystem/S3/S3-compatible/Azure Blob `Distribution` 여러 개를 결합한다. - 읽기 전용 수집: `src/sdp/storage_readers.py`는 list/read만 제공하고 이동·삭제·원격 mutation 기능이 없다. - 의미 근거: `SemanticAssertion`은 confidence, proposed review status, chunk SHA-256, 문자 offset을 보존하며 원문 인용은 저장하지 않는다. - LLM 경계: `src/sdp/document_semantics.py::ContextualOrchestratorClient`가 `/v1/chat/completions`와 `/v1/embeddings`만 호출한다. 포털에는 OpenAI/provider key가 없다. - provider 중립성: Synology는 filesystem 배치일 뿐 필수 구성요소가 아니며, 파일 정체성은 저장소 URL과 독립적이다. -- 정책/API: `POST /file-assets`, `GET /file-assets/{asset_id}`, `/jsonld`, `/validate`가 기존 policy 및 graph store를 재사용한다. +- 정책/API: `POST /file-assets`, `GET /file-assets/{asset_id}`, `/jsonld`, `/validate`는 검증된 OIDC Bearer actor context와 저장된 `FileAsset.tenant_id`를 중앙 policy decision에 전달한다. body/query `actor`는 권한 근거가 아니며 locator 공개는 동일 tenant의 admin 또는 platform-admin이 필요하다. +- 그래프 격리: 같은 SHA 및 파일 관계 대상의 cross-tenant merge를 거부한다. generic graph mutation은 file/distribution node를 다룰 수 없고 graph traversal/semantic search는 OIDC tenant policy로 file node를 필터링하며 locator를 redaction한다. +- 운영 패키징: SHACL/OWL TTL은 `sdp/resources/*.ttl` package data로 wheel/container에 포함되고, `orchestrator_base_url` 구성 시 app lifespan은 runtime `CredentialRegistry` token이 없으면 fail-closed한다. +- 벡터 일관성: KV `embedding_dimension`을 orchestrator `dimensions`에 전달하고 같은 orchestrator embedder를 memory/Postgres graph store의 ingest와 검색에 주입한다. - 파일럿: `src/sdp/file_pilot.py`가 content deduplication, 안전한 문서 추출, graph projection, 로컬 전용 manifest를 수행한다. - 증빙 테스트: - `tests/test_file_knowledge.py::test_orchestrator_extractor_uses_strict_schema_and_persists_only_evidence_reference` diff --git a/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md b/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md index 5e174c4..17474d2 100644 --- a/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md +++ b/docs/superpowers/specs/2026-07-21-semantic-file-ontology-design.md @@ -65,7 +65,7 @@ LLM 또는 규칙이 제안한 파일과 업무 개념 사이의 관계다. - 허용 관계: `belongsToProject`, `usesSystem`, `hasWorkPhase`, `hasArtifactType`, `hasTopic`, `wasDerivedFrom`, `previousVersion` - 필수 값: 대상 개념, 근거 참조(chunk SHA-256과 문자 offset), 신뢰도(0..1), 추출 방법, 검토 상태 - 모든 LLM 주장의 초기 상태는 `proposed`다. -- SHACL 호환 검증은 필수 필드·관계 allowlist·신뢰도 범위·근거 참조 존재를 검사한다. +- pySHACL 실행 검증은 필수 필드·관계 allowlist·신뢰도 범위·근거 참조·offset 순서를 검사한다. - 주장은 그래프 탐색과 검색에 사용하되 응답에 검토 상태를 항상 노출한다. ## 저장소 어댑터 @@ -94,17 +94,17 @@ SDK를 포털의 필수 의존성으로 추가하지 않는다. 실제 클라우 6. `contextual-orchestrator`의 `POST /v1/chat/completions`에 `response_format=json_schema`, `store=false`를 보내 의미 후보와 짧은 근거 인용을 받는다. orchestrator가 provider와 OpenAI 자격증명을 소유한다. 7. 근거 인용이 입력 조각에 실제로 존재하는지 확인한 뒤 인용문 대신 chunk SHA-256과 문자 offset만 남긴다. 8. 여러 조각의 후보를 `(관계, 정규화 label)` 기준으로 합치고 가장 높은 신뢰도와 그 근거 참조를 유지한다. -9. SHACL 호환 검증을 통과한 후보만 그래프에 `proposed` assertion으로 기록한다. +9. pySHACL 검증을 통과한 후보만 그래프에 `proposed` assertion으로 기록한다. 10. 파일 node의 embedding text는 제목과 제안된 개념 label만 사용하고 `contextual-orchestrator`의 동기 `POST /v1/embeddings`로 벡터화한다. 원문과 근거 인용문은 저장하지 않는다. 포털에는 orchestrator base URL과 model id를 KV application config로, inference token을 주입된 credential registry로 공급한다. 포털 코드가 OpenAI key나 `os.getenv()`를 읽지 않는다. HTTP transport는 Python 표준 라이브러리를 사용하고 LLM 요청에 `store: false`와 pinned model id를 포함한다. orchestrator의 동기 embedding endpoint는 기존 `/v1/batch/embeddings` 코어의 분할·provider backend·비용 원장을 재사용하고 최대 30초 동안 완료를 기다린 뒤 OpenAI 호환 응답을 반환한다. CI에서는 양쪽 HTTP 호출을 fake transport로 대체한다. ## API와 정책 -- `POST /file-assets`: 관리자만 검증된 파일 메타데이터와 후보 주장을 ingest한다. +- `POST /file-assets`: 검증된 OIDC Bearer actor context의 관리자만 파일 메타데이터와 후보 주장을 ingest한다. body/query actor는 신뢰하지 않는다. - `GET /file-assets/{asset_id}`: 인증된 reader가 자산과 의미 관계를 조회한다. - `GET /file-assets/{asset_id}/jsonld`: 기본적으로 저장소 locator를 제거한 JSON-LD를 반환한다. -- `GET /file-assets/{asset_id}/validate`: SHACL 호환 validation report를 반환한다. +- `GET /file-assets/{asset_id}/validate`: pySHACL validation report를 반환한다. - 기존 `POST /graph/query`와 `POST /search/semantic`을 그대로 사용해 관계 탐색과 의미 검색을 제공한다. 모든 write/read route는 기존 `policy.evaluate()` 경로를 재사용한다. 원문 조각, API key, cloud credential은 request/response, graph property, audit detail에 포함하지 않는다. diff --git a/ontology/cwl-file-profile.ttl b/ontology/cwl-file-profile.ttl index a44989a..685a063 100644 --- a/ontology/cwl-file-profile.ttl +++ b/ontology/cwl-file-profile.ttl @@ -2,6 +2,7 @@ @prefix dcat: . @prefix owl: . @prefix prov: . +@prefix rdfs: . @prefix skos: . @prefix xsd: . @@ -9,7 +10,7 @@ cwl: a owl:Ontology ; owl:versionInfo "0.1.0" . cwl:FileAsset a owl:Class ; - owl:equivalentClass [ a owl:Class ; owl:intersectionOf ( dcat:Resource prov:Entity ) ] . + rdfs:subClassOf dcat:Resource, prov:Entity . cwl:BusinessProject a owl:Class ; owl:subClassOf skos:Concept . cwl:System a owl:Class ; owl:subClassOf skos:Concept . @@ -32,6 +33,8 @@ cwl:wasDerivedFrom a owl:ObjectProperty ; cwl:previousVersion a owl:ObjectProperty . cwl:confidence a owl:DatatypeProperty ; owl:range xsd:decimal . +cwl:tenantId a owl:DatatypeProperty ; owl:range xsd:string . +cwl:targetKind a owl:DatatypeProperty ; owl:range xsd:string . cwl:reviewStatus a owl:DatatypeProperty ; owl:range xsd:string . cwl:evidenceChunkSha256 a owl:DatatypeProperty ; owl:range xsd:string . cwl:evidenceStart a owl:DatatypeProperty ; owl:range xsd:nonNegativeInteger . diff --git a/ontology/cwl-file-shapes.ttl b/ontology/cwl-file-shapes.ttl index 3b81d6c..7fee210 100644 --- a/ontology/cwl-file-shapes.ttl +++ b/ontology/cwl-file-shapes.ttl @@ -3,23 +3,71 @@ @prefix dcterms: . @prefix sh: . @prefix spdx: . +@prefix xsd: . cwl:FileAssetShape a sh:NodeShape ; sh:targetClass cwl:FileAsset ; - sh:property [ sh:path dcterms:title ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path spdx:checksum ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path dcat:distribution ; sh:minCount 1 ] . + sh:property [ + sh:path dcterms:identifier ; sh:minCount 1 ; sh:maxCount 1 ; + sh:pattern "^urn:sha256:[0-9a-f]{64}$" + ] ; + sh:property [ sh:path dcterms:title ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path cwl:tenantId ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path dcterms:format ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path dcat:byteSize ; sh:minCount 1 ; sh:maxCount 1 ; sh:minInclusive 0 ] ; + sh:property [ sh:path spdx:checksum ; sh:minCount 1 ; sh:maxCount 1 ; sh:node cwl:ChecksumShape ] ; + sh:property [ sh:path dcat:distribution ; sh:minCount 1 ; sh:node cwl:DistributionShape ] ; + sh:property [ sh:path cwl:assertion ; sh:node cwl:SemanticAssertionShape ] . + +cwl:ChecksumShape a sh:NodeShape ; + sh:property [ + sh:path spdx:checksumValue ; sh:minCount 1 ; sh:maxCount 1 ; + sh:pattern "^[0-9a-f]{64}$" + ] . cwl:DistributionShape a sh:NodeShape ; sh:targetClass dcat:Distribution ; - sh:property [ sh:path cwl:storageProvider ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:endpointId ; sh:minCount 1 ; sh:maxCount 1 ] . + sh:property [ + sh:path cwl:storageProvider ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( "filesystem" "s3" "s3_compatible" "azure_blob" ) + ] ; + sh:property [ sh:path cwl:endpointId ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path dcat:accessURL ; sh:minCount 1 ; sh:maxCount 1 ] . cwl:SemanticAssertionShape a sh:NodeShape ; sh:targetClass cwl:SemanticAssertion ; - sh:property [ sh:path cwl:confidence ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:reviewStatus ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:evidenceChunkSha256 ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:evidenceStart ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:evidenceEnd ; sh:minCount 1 ; sh:maxCount 1 ] ; - sh:property [ sh:path cwl:extractionMethod ; sh:minCount 1 ; sh:maxCount 1 ] . + sh:property [ + sh:path cwl:relation ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( + cwl:belongsToProject cwl:usesSystem cwl:hasWorkPhase + cwl:hasArtifactType cwl:hasTopic cwl:wasDerivedFrom cwl:previousVersion + ) + ] ; + sh:property [ sh:path cwl:target ; sh:minCount 1 ; sh:maxCount 1 ; sh:nodeKind sh:IRI ] ; + sh:property [ + sh:path cwl:targetKind ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( + "business_project" "system" "work_phase" + "artifact_type" "topic" "file_asset" + ) + ] ; + sh:property [ + sh:path cwl:confidence ; sh:minCount 1 ; sh:maxCount 1 ; + sh:minInclusive 0 ; sh:maxInclusive 1 + ] ; + sh:property [ + sh:path cwl:reviewStatus ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( "proposed" "approved" "rejected" ) + ] ; + sh:property [ + sh:path cwl:evidenceChunkSha256 ; sh:minCount 1 ; sh:maxCount 1 ; + sh:pattern "^[0-9a-f]{64}$" + ] ; + sh:property [ + sh:path cwl:evidenceStart ; sh:minCount 1 ; sh:maxCount 1 ; + sh:minInclusive 0 ; sh:lessThan cwl:evidenceEnd + ] ; + sh:property [ + sh:path cwl:evidenceEnd ; sh:minCount 1 ; sh:maxCount 1 ; sh:minInclusive 1 + ] ; + sh:property [ sh:path cwl:extractionMethod ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] . diff --git a/pyproject.toml b/pyproject.toml index b5280cd..5b568a0 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -14,9 +14,13 @@ dependencies = [ "PyJWT[crypto]==2.13.0", "pydantic==2.13.4", "pypdf==6.14.2", + "pyshacl==0.40.0", "uvicorn==0.51.0", ] +[tool.setuptools.package-data] +sdp = ["resources/*.ttl"] + [project.optional-dependencies] # Postgres + Apache AGE (openCypher) + pgvector backend for the graph engine. graph = [ diff --git a/requirements-dev.txt b/requirements-dev.txt index fbbc51d..b4bbb7f 100644 --- a/requirements-dev.txt +++ b/requirements-dev.txt @@ -271,6 +271,10 @@ h11==0.16.0 \ # via # httpcore # uvicorn +html5rdf==1.2.1 \ + --hash=sha256:1f519121bc366af3e485310dc8041d2e86e5173c1a320fac3dc9d2604069b83e \ + --hash=sha256:ace9b420ce52995bb4f05e7425eedf19e433c981dfe7a831ab391e2fa2e1a195 + # via rdflib httpcore==1.0.9 \ --hash=sha256:2d400746a40668fc9dec9810239072b40b4484b640a8c38fd654a024c7a1bf55 \ --hash=sha256:6e34463af53fd2ab5d807f399a9b45ea31c3dfa2276f15a2c3f00afff6e176e8 @@ -349,14 +353,24 @@ iniconfig==2.3.0 \ --hash=sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730 \ --hash=sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12 # via pytest +owlrl==7.6.2 \ + --hash=sha256:83347bf7f133979e87b2b18695d51d25510b99cec3f6919b5df05d4fbf058ae0 \ + --hash=sha256:c743f35c2d908396e77823852bb1ebbce88340cd49961493983bec42c93283a8 + # via pyshacl packaging==26.2 \ --hash=sha256:5fc45236b9446107ff2415ce77c807cee2862cb6fac22b8a73826d0693b0980e \ --hash=sha256:ff452ff5a3e828ce110190feff1178bb1f2ea2281fa2075aadb987c2fb221661 - # via pytest + # via + # pyshacl + # pytest pluggy==1.6.0 \ --hash=sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3 \ --hash=sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746 # via pytest +prettytable==3.18.0 \ + --hash=sha256:439217116152244369caf3d9f1caf2f9fe29b03bd79e88d2928c8e718c95d680 \ + --hash=sha256:b3346e0e6f79180833aebaac088ae926340586cf6d7d991b9eb125b65f72313a + # via pyshacl psycopg==3.3.4 \ --hash=sha256:b6bbc25ccf05c8fad3b061d9db2ef0909a555171b84b07f29458a447253d679a \ --hash=sha256:e21207764952cff81b6b8bdacad9a3939f2793367fdac2987b3aac36a651b5bc @@ -558,14 +572,28 @@ pyjwt==2.13.0 \ --hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \ --hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728 # via semantic-data-portal (pyproject.toml) +pyparsing==3.3.2 \ + --hash=sha256:850ba148bd908d7e2411587e247a1e4f0327839c40e2e5e6d05a007ecc69911d \ + --hash=sha256:c777f4d763f140633dcb6d8a3eda953bf7a214dc4eff598413c070bcdc117cbc + # via rdflib pypdf==6.14.2 \ --hash=sha256:3f07891af76dc002657e04993ab9b4de81de29f9013b9761d0b7968bff12e946 \ --hash=sha256:7873f502fe4385e79539b21d872392dc0c4e3714327c15881cbc7fbfd1f95b25 # via semantic-data-portal (pyproject.toml) +pyshacl==0.40.0 \ + --hash=sha256:1f814baa7fc2b877c999819fed9e4b8e4ba2461bc00ed8ed3e8bb4346258500f \ + --hash=sha256:d9cd86f108266a157162afb69a106689c868388e5d7be80c1fcf16070cf6ad08 + # via semantic-data-portal (pyproject.toml) pytest==9.1.1 \ --hash=sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313 \ --hash=sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c # via semantic-data-portal (pyproject.toml) +rdflib==7.6.0 \ + --hash=sha256:30c0a3ebf4c0e09215f066be7246794b6492e054e782d7ac2a34c9f70a15e0dd \ + --hash=sha256:6c831288d5e4a5a7ece85d0ccde9877d512a3d0f02d7c06455d00d6d0ea379df + # via + # owlrl + # pyshacl sortedcontainers==2.4.0 \ --hash=sha256:25caa5a06cc30b6b83d11423433f65d1f9d76c4c6a0c90e3379eaa43b9bfdb88 \ --hash=sha256:a163dcaede0f1c021485e957a39245190e74249897e2ae4b2aa38595db237ee0 @@ -657,3 +685,7 @@ uvicorn==0.51.0 \ --hash=sha256:5d38af6cd620f2ae3849fb44fd4879e0890aa1febe8d47eb355fb45d93fe6a5b \ --hash=sha256:f6f4b69b657c312f516dd2d268ab9ae6f254b11e4bac504f37b2ab58b24dd0b0 # via semantic-data-portal (pyproject.toml) +wcwidth==0.8.2 \ + --hash=sha256:91fbef97204b96a3d4d421609b80340b760cf33e26da123ff243d76b1fda8dda \ + --hash=sha256:d63947694a0539a1d51e01eda7caf800c291020e6cdd7e28ad7b14dd33ad4f85 + # via prettytable diff --git a/requirements-test.in b/requirements-test.in index d216e97..b1370a6 100644 --- a/requirements-test.in +++ b/requirements-test.in @@ -5,6 +5,7 @@ psycopg[binary]==3.3.4 PyJWT[crypto]==2.13.0 pydantic==2.13.4 pypdf==6.14.2 +pyshacl==0.40.0 uvicorn==0.51.0 pytest==9.1.1 httpx==0.28.1 diff --git a/requirements-test.txt b/requirements-test.txt index dadda76..8f0f709 100644 --- a/requirements-test.txt +++ b/requirements-test.txt @@ -271,6 +271,10 @@ h11==0.16.0 \ # via # httpcore # uvicorn +html5rdf==1.2.1 \ + --hash=sha256:1f519121bc366af3e485310dc8041d2e86e5173c1a320fac3dc9d2604069b83e \ + --hash=sha256:ace9b420ce52995bb4f05e7425eedf19e433c981dfe7a831ab391e2fa2e1a195 + # via rdflib httpcore==1.0.9 \ --hash=sha256:2d400746a40668fc9dec9810239072b40b4484b640a8c38fd654a024c7a1bf55 \ --hash=sha256:6e34463af53fd2ab5d807f399a9b45ea31c3dfa2276f15a2c3f00afff6e176e8 @@ -349,14 +353,24 @@ iniconfig==2.3.0 \ --hash=sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730 \ --hash=sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12 # via pytest +owlrl==7.6.2 \ + --hash=sha256:83347bf7f133979e87b2b18695d51d25510b99cec3f6919b5df05d4fbf058ae0 \ + --hash=sha256:c743f35c2d908396e77823852bb1ebbce88340cd49961493983bec42c93283a8 + # via pyshacl packaging==26.2 \ --hash=sha256:5fc45236b9446107ff2415ce77c807cee2862cb6fac22b8a73826d0693b0980e \ --hash=sha256:ff452ff5a3e828ce110190feff1178bb1f2ea2281fa2075aadb987c2fb221661 - # via pytest + # via + # pyshacl + # pytest pluggy==1.6.0 \ --hash=sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3 \ --hash=sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746 # via pytest +prettytable==3.18.0 \ + --hash=sha256:439217116152244369caf3d9f1caf2f9fe29b03bd79e88d2928c8e718c95d680 \ + --hash=sha256:b3346e0e6f79180833aebaac088ae926340586cf6d7d991b9eb125b65f72313a + # via pyshacl psycopg==3.3.4 \ --hash=sha256:b6bbc25ccf05c8fad3b061d9db2ef0909a555171b84b07f29458a447253d679a \ --hash=sha256:e21207764952cff81b6b8bdacad9a3939f2793367fdac2987b3aac36a651b5bc @@ -558,14 +572,28 @@ pyjwt==2.13.0 \ --hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \ --hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728 # via -r requirements-test.in +pyparsing==3.3.2 \ + --hash=sha256:850ba148bd908d7e2411587e247a1e4f0327839c40e2e5e6d05a007ecc69911d \ + --hash=sha256:c777f4d763f140633dcb6d8a3eda953bf7a214dc4eff598413c070bcdc117cbc + # via rdflib pypdf==6.14.2 \ --hash=sha256:3f07891af76dc002657e04993ab9b4de81de29f9013b9761d0b7968bff12e946 \ --hash=sha256:7873f502fe4385e79539b21d872392dc0c4e3714327c15881cbc7fbfd1f95b25 # via -r requirements-test.in +pyshacl==0.40.0 \ + --hash=sha256:1f814baa7fc2b877c999819fed9e4b8e4ba2461bc00ed8ed3e8bb4346258500f \ + --hash=sha256:d9cd86f108266a157162afb69a106689c868388e5d7be80c1fcf16070cf6ad08 + # via -r requirements-test.in pytest==9.1.1 \ --hash=sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313 \ --hash=sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c # via -r requirements-test.in +rdflib==7.6.0 \ + --hash=sha256:30c0a3ebf4c0e09215f066be7246794b6492e054e782d7ac2a34c9f70a15e0dd \ + --hash=sha256:6c831288d5e4a5a7ece85d0ccde9877d512a3d0f02d7c06455d00d6d0ea379df + # via + # owlrl + # pyshacl sortedcontainers==2.4.0 \ --hash=sha256:25caa5a06cc30b6b83d11423433f65d1f9d76c4c6a0c90e3379eaa43b9bfdb88 \ --hash=sha256:a163dcaede0f1c021485e957a39245190e74249897e2ae4b2aa38595db237ee0 @@ -657,3 +685,7 @@ uvicorn==0.51.0 \ --hash=sha256:5d38af6cd620f2ae3849fb44fd4879e0890aa1febe8d47eb355fb45d93fe6a5b \ --hash=sha256:f6f4b69b657c312f516dd2d268ab9ae6f254b11e4bac504f37b2ab58b24dd0b0 # via -r requirements-test.in +wcwidth==0.8.2 \ + --hash=sha256:91fbef97204b96a3d4d421609b80340b760cf33e26da123ff243d76b1fda8dda \ + --hash=sha256:d63947694a0539a1d51e01eda7caf800c291020e6cdd7e28ad7b14dd33ad4f85 + # via prettytable diff --git a/requirements.txt b/requirements.txt index e95359b..1f2875c 100644 --- a/requirements.txt +++ b/requirements.txt @@ -178,10 +178,26 @@ h11==0.16.0 \ --hash=sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1 \ --hash=sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86 # via uvicorn +html5rdf==1.2.1 \ + --hash=sha256:1f519121bc366af3e485310dc8041d2e86e5173c1a320fac3dc9d2604069b83e \ + --hash=sha256:ace9b420ce52995bb4f05e7425eedf19e433c981dfe7a831ab391e2fa2e1a195 + # via rdflib idna==3.18 \ --hash=sha256:7f952cbe720b688055e3f87de14f5c3e5fdaa8bc3928985c4077ca689de849a2 \ --hash=sha256:ffb385a7e039654cef1ab9ef32c6fafe283c0c0467bba1d9029738ce4a14a848 # via anyio +owlrl==7.6.2 \ + --hash=sha256:83347bf7f133979e87b2b18695d51d25510b99cec3f6919b5df05d4fbf058ae0 \ + --hash=sha256:c743f35c2d908396e77823852bb1ebbce88340cd49961493983bec42c93283a8 + # via pyshacl +packaging==26.2 \ + --hash=sha256:5fc45236b9446107ff2415ce77c807cee2862cb6fac22b8a73826d0693b0980e \ + --hash=sha256:ff452ff5a3e828ce110190feff1178bb1f2ea2281fa2075aadb987c2fb221661 + # via pyshacl +prettytable==3.18.0 \ + --hash=sha256:439217116152244369caf3d9f1caf2f9fe29b03bd79e88d2928c8e718c95d680 \ + --hash=sha256:b3346e0e6f79180833aebaac088ae926340586cf6d7d991b9eb125b65f72313a + # via pyshacl psycopg==3.3.4 \ --hash=sha256:b6bbc25ccf05c8fad3b061d9db2ef0909a555171b84b07f29458a447253d679a \ --hash=sha256:e21207764952cff81b6b8bdacad9a3939f2793367fdac2987b3aac36a651b5bc @@ -379,10 +395,24 @@ pyjwt==2.13.0 \ --hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \ --hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728 # via semantic-data-portal (pyproject.toml) +pyparsing==3.3.2 \ + --hash=sha256:850ba148bd908d7e2411587e247a1e4f0327839c40e2e5e6d05a007ecc69911d \ + --hash=sha256:c777f4d763f140633dcb6d8a3eda953bf7a214dc4eff598413c070bcdc117cbc + # via rdflib pypdf==6.14.2 \ --hash=sha256:3f07891af76dc002657e04993ab9b4de81de29f9013b9761d0b7968bff12e946 \ --hash=sha256:7873f502fe4385e79539b21d872392dc0c4e3714327c15881cbc7fbfd1f95b25 # via semantic-data-portal (pyproject.toml) +pyshacl==0.40.0 \ + --hash=sha256:1f814baa7fc2b877c999819fed9e4b8e4ba2461bc00ed8ed3e8bb4346258500f \ + --hash=sha256:d9cd86f108266a157162afb69a106689c868388e5d7be80c1fcf16070cf6ad08 + # via semantic-data-portal (pyproject.toml) +rdflib==7.6.0 \ + --hash=sha256:30c0a3ebf4c0e09215f066be7246794b6492e054e782d7ac2a34c9f70a15e0dd \ + --hash=sha256:6c831288d5e4a5a7ece85d0ccde9877d512a3d0f02d7c06455d00d6d0ea379df + # via + # owlrl + # pyshacl starlette==1.3.1 \ --hash=sha256:05d0213193f2fbaae60e2ecb593b4add4262ad4e46536b54abe36f11a71724e0 \ --hash=sha256:c7372aae11c3c3f26a42df7bd626cec2f47d03483d261d369516a615a53714c6 @@ -409,3 +439,7 @@ uvicorn==0.51.0 \ --hash=sha256:5d38af6cd620f2ae3849fb44fd4879e0890aa1febe8d47eb355fb45d93fe6a5b \ --hash=sha256:f6f4b69b657c312f516dd2d268ab9ae6f254b11e4bac504f37b2ab58b24dd0b0 # via semantic-data-portal (pyproject.toml) +wcwidth==0.8.2 \ + --hash=sha256:91fbef97204b96a3d4d421609b80340b760cf33e26da123ff243d76b1fda8dda \ + --hash=sha256:d63947694a0539a1d51e01eda7caf800c291020e6cdd7e28ad7b14dd33ad4f85 + # via prettytable diff --git a/src/sdp/api.py b/src/sdp/api.py index bb57fef..3a2604d 100644 --- a/src/sdp/api.py +++ b/src/sdp/api.py @@ -1,14 +1,16 @@ from __future__ import annotations import logging +from contextlib import asynccontextmanager from datetime import datetime, timezone from time import monotonic from typing import Any -from fastapi import FastAPI, HTTPException, Query, Request, Response +from fastapi import Depends, FastAPI, HTTPException, Query, Request, Response from fastapi.responses import HTMLResponse from fastapi.middleware.cors import CORSMiddleware from sdp_core import ( + ActorContext, buyer_demo_activation_plan, enterprise_controls_manifest, enterprise_kpi_framework, @@ -27,7 +29,8 @@ OntologyConceptRequest, SemanticSearchRequest, ) -from .graph_store import get_store +from .document_semantics import CredentialRegistry, set_credential_registry, validate_runtime_credentials +from .graph_store import get_store, set_store from .file_ontology import ( FileAsset, FileAssetIngestRequest, @@ -72,7 +75,7 @@ request_id_from_headers, request_id_header, ) -from .policy import evaluate +from .policy import evaluate, evaluate_file_asset_access, evaluate_graph_access from .semantic_validation import enterprise_shacl_validation_summary, validate_dataset_semantics from .steward_review import build_steward_review_summary @@ -80,10 +83,22 @@ _logger = logging.getLogger(__name__) +@asynccontextmanager +async def _runtime_lifespan(application: FastAPI): + registry = getattr(application.state, "credential_registry", None) + if registry is not None: + set_credential_registry(registry) + validate_runtime_credentials(get_app_config()) + set_store(None) + seed_store() + yield + + app = FastAPI( title="Semantic Data Portal", description="온톨로지 기반 그래프 데이터 카탈로그 및 시맨틱 검색 서비스", version="0.3.0", + lifespan=_runtime_lifespan, ) # CORS allowlist comes from config (KV table `config_entries` when a database is @@ -97,23 +112,11 @@ ) -@app.on_event("startup") -def _bootstrap_graph_engine() -> None: - """Seed the active graph store on startup (idempotent).""" - - try: - seed_store() - except Exception: # pragma: no cover - seeding must never block startup - pass +def create_app(credential_registry: CredentialRegistry | None = None) -> FastAPI: + """Return the ASGI app with trusted credentials installed before startup.""" - -# Seed at import time as well: the test client and embedded/submodule callers -# may hit endpoints without triggering ASGI startup events. Seeding is -# idempotent so running it here and on startup is safe. -try: - seed_store() -except Exception: # pragma: no cover - pass + app.state.credential_registry = credential_registry + return app @app.middleware("http") @@ -697,41 +700,60 @@ def ontology_term_graph(term: str) -> dict[str, Any]: # --- Graph ingestion (nodes / edges / concepts) ------------------------------ -def _authorize_graph_write(actor: str, resource: str) -> None: - """Graph/ontology writes require an authorized (admin) subject. +def _authenticated_actor(request: Request) -> ActorContext: + authorization = request.headers.get("Authorization", "") + scheme, _, token = authorization.partition(" ") + if scheme.lower() != "bearer" or not token: + raise HTTPException( + status_code=401, + detail="authenticated bearer token is required", + headers={"WWW-Authenticate": "Bearer"}, + ) + try: + context, _claims = authz.verify_oidc_jwks_token(token) + except ValueError: + raise HTTPException( + status_code=401, + detail="bearer token is invalid", + headers={"WWW-Authenticate": "Bearer"}, + ) from None + request.state.actor_context = context + return context - Uses the same policy engine as the catalog mutation endpoints: the ``create`` - action is allowed only for an admin subject, so unauthenticated/anonymous - writers are refused with 403. - """ - decision = evaluate(subject=actor, resource=resource, action="create", purpose="graph") +def _enforce_graph_policy(actor: ActorContext, resource: str, action: str) -> None: + decision = evaluate_graph_access(actor, resource, action) if decision.effect != "allow": raise HTTPException(status_code=403, detail=decision.reason) -def _authorize_graph_read(actor: str, resource: str) -> None: - """Graph traversal / semantic search require an authenticated reader. - - Uses the catalog discovery policy branch (``search``): allowed for any reader - role, denied for anonymous/unauthenticated subjects. - """ - - decision = evaluate(subject=actor, resource=resource, action="search", purpose="graph") - if decision.effect != "allow": - raise HTTPException(status_code=403, detail=decision.reason) +def _is_governed_file_node_id(node_id: str) -> bool: + return node_id.startswith("urn:sha256:") or node_id.startswith("urn:cwl:distribution:") @app.post("/ontology/concepts") -def ingest_concept(payload: OntologyConceptRequest) -> dict[str, Any]: - _authorize_graph_write(payload.actor, payload.concept) - record = get_store().upsert_concept(payload.model_dump(exclude={"actor"})) +def ingest_concept( + payload: OntologyConceptRequest, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + _enforce_graph_policy(actor, payload.concept, "write_graph") + try: + record = get_store().upsert_concept(payload.model_dump()) + except ValueError as exc: + raise HTTPException(status_code=400, detail=str(exc)) return {"status": "upserted", "concept": record} @app.post("/graph/nodes") -def ingest_graph_node(payload: GraphNodeRequest) -> dict[str, Any]: - _authorize_graph_write(payload.actor, payload.node_id) +def ingest_graph_node( + payload: GraphNodeRequest, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + _enforce_graph_policy(actor, payload.node_id, "write_graph") + if payload.kind in {"file_asset", "distribution"} or _is_governed_file_node_id( + payload.node_id + ): + raise HTTPException(status_code=400, detail="governed file nodes require /file-assets") node = get_store().upsert_node( payload.node_id, payload.kind, @@ -743,10 +765,22 @@ def ingest_graph_node(payload: GraphNodeRequest) -> dict[str, Any]: @app.post("/graph/edges") -def ingest_graph_edge(payload: GraphEdgeRequest) -> dict[str, Any]: - _authorize_graph_write(payload.actor, payload.source_id) +def ingest_graph_edge( + payload: GraphEdgeRequest, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + _enforce_graph_policy(actor, payload.source_id, "write_graph") + store = get_store() + source = store.get_node(payload.source_id) + target = store.get_node(payload.target_id) + if ( + _is_governed_file_node_id(payload.source_id) + or _is_governed_file_node_id(payload.target_id) + or any(node and node.kind in {"file_asset", "distribution"} for node in (source, target)) + ): + raise HTTPException(status_code=400, detail="governed file edges require /file-assets") try: - edge = get_store().upsert_edge( + edge = store.upsert_edge( payload.edge_type, payload.source_id, payload.target_id, @@ -758,10 +792,15 @@ def ingest_graph_edge(payload: GraphEdgeRequest) -> dict[str, Any]: @app.get("/graph/nodes/{node_id}") -def get_graph_node(node_id: str) -> dict[str, Any]: +def get_graph_node( + node_id: str, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + _enforce_graph_policy(actor, node_id, "read_graph") node = get_store().get_node(node_id) if node is None: raise HTTPException(status_code=404, detail="node not found") + _enforce_governed_node_read(actor, node) return _redact_storage_coordinates(node.as_dict()) @@ -783,14 +822,68 @@ def _redact_storage_coordinates(value: Any) -> Any: return value +def _enforce_file_policy( + actor: ActorContext, + resource: str, + action: str, + resource_tenant_id: str, +) -> None: + decision = evaluate_file_asset_access(actor, resource, action, resource_tenant_id) + if decision.effect != "allow": + raise HTTPException(status_code=403, detail=decision.reason) + + +def _enforce_governed_node_read(actor: ActorContext, node: Any) -> None: + if node.kind not in {"file_asset", "distribution"}: + return + properties = node.properties if isinstance(node.properties, dict) else {} + tenant_id = str( + properties.get("tenant_id") + or (properties.get("file_asset") or {}).get("tenant_id") + or "" + ) + _enforce_file_policy(actor, node.node_id, "read_file_asset", tenant_id) + + +def _filter_graph_result(actor: ActorContext, store: Any, result: dict[str, Any]) -> dict[str, Any]: + visible_nodes: list[dict[str, Any]] = [] + visible_ids: set[str] = set() + for item in result.get("nodes", []): + node_id = str(item.get("node_id", "")) + node = store.get_node(node_id) + try: + if node is not None: + _enforce_governed_node_read(actor, node) + except HTTPException as exc: + if exc.status_code == 403: + continue + raise + visible_nodes.append(_redact_storage_coordinates(item)) + visible_ids.add(node_id) + return { + **result, + "nodes": visible_nodes, + "edges": [ + _redact_storage_coordinates(edge) + for edge in result.get("edges", []) + if edge.get("source_id") in visible_ids and edge.get("target_id") in visible_ids + ], + } + + @app.post("/file-assets") -def ingest_file_asset(payload: FileAssetIngestRequest) -> dict[str, Any]: - _authorize_graph_write(payload.actor, payload.asset_id) +def ingest_file_asset( + payload: FileAssetIngestRequest, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + asset = FileAsset.model_validate(payload.model_dump()).model_copy( + update={"tenant_id": actor.tenant_id} + ) + _enforce_file_policy(actor, asset.asset_id, "create_file_asset", asset.tenant_id) try: - asset = upsert_file_asset( - get_store(), - FileAsset.model_validate(payload.model_dump(exclude={"actor"})), - ) + asset = upsert_file_asset(get_store(), asset) + except PermissionError as exc: + raise HTTPException(status_code=403, detail=str(exc)) except ValueError as exc: raise HTTPException(status_code=400, detail=str(exc)) return { @@ -800,23 +893,23 @@ def ingest_file_asset(payload: FileAssetIngestRequest) -> dict[str, Any]: } -def _read_file_asset(asset_id: str, actor: str): - _authorize_graph_read(actor, asset_id) +def _read_file_asset(asset_id: str, actor: ActorContext): asset = get_file_asset(get_store(), asset_id) if asset is None: raise HTTPException(status_code=404, detail="file asset not found") + _enforce_file_policy(actor, asset_id, "read_file_asset", asset.tenant_id) return asset @app.get("/file-assets/{asset_id}") def file_asset_detail( asset_id: str, - actor: str = Query(default="anonymous"), include_locations: bool = Query(default=False), + actor: ActorContext = Depends(_authenticated_actor), ) -> dict[str, Any]: asset = _read_file_asset(asset_id, actor) if include_locations: - _authorize_graph_write(actor, asset_id) + _enforce_file_policy(actor, asset_id, "read_file_locations", asset.tenant_id) payload = {"asset_id": asset.asset_id, **asset.model_dump(mode="json")} return payload if include_locations else _redact_storage_coordinates(payload) @@ -824,19 +917,19 @@ def file_asset_detail( @app.get("/file-assets/{asset_id}/jsonld") def file_asset_jsonld_export( asset_id: str, - actor: str = Query(default="anonymous"), include_locations: bool = Query(default=False), + actor: ActorContext = Depends(_authenticated_actor), ) -> dict[str, Any]: asset = _read_file_asset(asset_id, actor) if include_locations: - _authorize_graph_write(actor, asset_id) + _enforce_file_policy(actor, asset_id, "read_file_locations", asset.tenant_id) return file_asset_jsonld(asset, include_locations=include_locations) @app.get("/file-assets/{asset_id}/validate") def file_asset_validation( asset_id: str, - actor: str = Query(default="anonymous"), + actor: ActorContext = Depends(_authenticated_actor), ) -> dict[str, Any]: return validate_file_asset(_read_file_asset(asset_id, actor)) @@ -845,15 +938,23 @@ def file_asset_validation( @app.post("/graph/query") -def graph_query(payload: GraphTraversalRequest) -> dict[str, Any]: - _authorize_graph_read(payload.actor, payload.start_id) +def graph_query( + payload: GraphTraversalRequest, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + _enforce_graph_policy(actor, payload.start_id, "read_graph") + store = get_store() + start_node = store.get_node(payload.start_id) + if start_node is not None: + _enforce_governed_node_read(actor, start_node) try: - return get_store().traverse( + result = store.traverse( payload.start_id, edge_types=payload.edge_types, direction=payload.direction, max_depth=payload.max_depth, ) + return _filter_graph_result(actor, store, result) except KeyError as exc: raise HTTPException(status_code=404, detail=str(exc)) except ValueError as exc: @@ -861,11 +962,32 @@ def graph_query(payload: GraphTraversalRequest) -> dict[str, Any]: @app.post("/search/semantic") -def semantic_search(payload: SemanticSearchRequest) -> dict[str, Any]: - _authorize_graph_read(payload.actor, "graph") - results = get_store().semantic_search( - payload.query, kind=payload.kind, limit=payload.limit +def semantic_search( + payload: SemanticSearchRequest, + actor: ActorContext = Depends(_authenticated_actor), +) -> dict[str, Any]: + _enforce_graph_policy(actor, "graph", "read_graph") + store = get_store() + search_tenant = None if "platform-admin" in actor.roles else actor.tenant_id + candidates = store.semantic_search( + payload.query, + kind=payload.kind, + limit=payload.limit, + tenant_id=search_tenant, ) + results: list[dict[str, Any]] = [] + for item in candidates: + node = store.get_node(str(item.get("node_id", ""))) + try: + if node is not None: + _enforce_governed_node_read(actor, node) + except HTTPException as exc: + if exc.status_code == 403: + continue + raise + results.append(_redact_storage_coordinates(item)) + if len(results) >= payload.limit: + break return {"query": payload.query, "count": len(results), "results": results} diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py index c57e29f..78e7d53 100644 --- a/src/sdp/document_semantics.py +++ b/src/sdp/document_semantics.py @@ -39,6 +39,8 @@ ".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet", ".pdf": "application/pdf", } +_MAX_OPENXML_MEMBER_BYTES = 8 * 1024 * 1024 +_MAX_OPENXML_TOTAL_BYTES = 32 * 1024 * 1024 @dataclass(frozen=True) @@ -74,6 +76,37 @@ def __repr__(self) -> str: return f"{type(self).__name__}(names={sorted(self._values)})" +_CREDENTIAL_REGISTRY: CredentialRegistry = EphemeralCredentialRegistry({}) + + +def set_credential_registry(registry: CredentialRegistry) -> None: + """Install the trusted host provider and invalidate credential-bound state.""" + + global _CREDENTIAL_REGISTRY + _CREDENTIAL_REGISTRY = registry + # Imported lazily to avoid the graph_store -> document_semantics cycle. + from .graph_store import set_store + + set_store(None) + + +def get_credential_registry() -> CredentialRegistry: + return _CREDENTIAL_REGISTRY + + +def validate_runtime_credentials(config: Any) -> None: + """Fail closed when governed LLM routing is configured without its token.""" + + if ( + config.orchestrator_base_url + and not get_credential_registry().get_credential("CONTEXTUAL_ORCHESTRATOR_TOKEN") + ): + raise RuntimeError( + "CONTEXTUAL_ORCHESTRATOR_TOKEN is required in the runtime credential registry " + "when orchestrator_base_url is configured" + ) + + def _decode_text(data: bytes) -> str: for encoding in ("utf-8-sig", "utf-16", "cp949"): try: @@ -95,13 +128,24 @@ def _xml_text(xml: bytes) -> list[str]: def _openxml_text(suffix: str, data: bytes) -> str: roots = _OPENXML_ROOTS[suffix] parts: list[str] = [] + total_uncompressed = 0 with ZipFile(BytesIO(data)) as archive: - for name in sorted(archive.namelist()): + for info in sorted(archive.infolist(), key=lambda item: item.filename): + name = info.filename if not name.endswith(".xml") or not any( name == root or name.startswith(root) for root in roots ): continue - parts.extend(_xml_text(archive.read(name))) + if info.file_size > _MAX_OPENXML_MEMBER_BYTES: + raise ValueError("OpenXML member exceeds extraction limit") + total_uncompressed += info.file_size + if total_uncompressed > _MAX_OPENXML_TOTAL_BYTES: + raise ValueError("OpenXML package exceeds extraction limit") + with archive.open(info) as member: + xml = member.read(_MAX_OPENXML_MEMBER_BYTES + 1) + if len(xml) > _MAX_OPENXML_MEMBER_BYTES: + raise ValueError("OpenXML member exceeds extraction limit") + parts.extend(_xml_text(xml)) return "\n".join(parts) @@ -201,6 +245,10 @@ def chunk_text( ], }, "target_label": {"type": "string", "minLength": 1, "maxLength": 200}, + "target_asset_id": { + "type": ["string", "null"], + "pattern": "^urn:sha256:[0-9a-f]{64}$", + }, "confidence": {"type": "number", "minimum": 0, "maximum": 1}, "evidence_quote": {"type": "string", "minLength": 1, "maxLength": 240}, }, @@ -208,6 +256,7 @@ def chunk_text( "relation", "target_kind", "target_label", + "target_asset_id", "confidence", "evidence_quote", ], @@ -249,6 +298,7 @@ def __init__( transport: OrchestratorTransport = _orchestrator_http_transport, semantic_model: str = "gpt-5-mini-2025-08-07", embedding_model: str = "text-embedding-3-small", + embedding_dimensions: int | None = None, timeout: int = 60, ) -> None: parsed = urlsplit(base_url) @@ -256,11 +306,14 @@ def __init__( raise ValueError("orchestrator base URL must be absolute HTTP(S)") if parsed.query or parsed.fragment: raise ValueError("orchestrator base URL must not contain query or fragment") + if embedding_dimensions is not None and embedding_dimensions <= 0: + raise ValueError("embedding dimensions must be positive") self.credentials = credentials self.base_url = base_url.rstrip("/") self.transport = transport self.semantic_model = semantic_model self.embedding_model = embedding_model + self.embedding_dimensions = embedding_dimensions self.timeout = timeout def _payload(self, filename: str, chunk: TextChunk) -> dict[str, Any]: @@ -273,6 +326,8 @@ def _payload(self, filename: str, chunk: TextChunk) -> dict[str, Any]: "content": ( "Extract only explicitly evidenced business semantics from the document chunk. " "Return Korean labels when the source uses Korean. Do not infer facts absent from the text." + " For file-to-file relations, return the explicit urn:sha256 target_asset_id; " + "otherwise set target_asset_id to null." ), }, { @@ -309,14 +364,14 @@ def _post(self, path: str, payload: dict[str, Any]) -> dict[str, Any]: def embed(self, texts: list[str]) -> list[list[float]]: if not texts or any(not isinstance(text, str) for text in texts): raise ValueError("embedding input must be a non-empty list of strings") - response = self._post( - "/v1/embeddings", - { - "model": self.embedding_model, - "input": texts, - "metadata": {"service": "semantic-data-portal"}, - }, - ) + payload: dict[str, Any] = { + "model": self.embedding_model, + "input": texts, + "metadata": {"service": "semantic-data-portal"}, + } + if self.embedding_dimensions is not None: + payload["dimensions"] = self.embedding_dimensions + response = self._post("/v1/embeddings", payload) data = response.get("data") if not isinstance(data, list) or len(data) != len(texts): raise ValueError("orchestrator embedding response is malformed") @@ -327,6 +382,11 @@ def embed(self, texts: list[str]) -> list[list[float]]: vector = item.get("embedding") if not isinstance(vector, list) or not vector: raise ValueError("orchestrator embedding vector is missing") + if ( + self.embedding_dimensions is not None + and len(vector) != self.embedding_dimensions + ): + raise ValueError("orchestrator embedding dimension is malformed") try: indexed[item["index"]] = [float(component) for component in vector] except (TypeError, ValueError) as exc: @@ -366,6 +426,7 @@ def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssert relation=candidate["relation"], target_kind=candidate["target_kind"], target_label=candidate["target_label"], + target_asset_id=candidate.get("target_asset_id"), confidence=candidate["confidence"], evidence_chunk_sha256=chunk.sha256, evidence_start=chunk.start + local_start, @@ -377,8 +438,23 @@ def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssert key = ( assertion.relation, assertion.target_kind, - assertion.target_label.strip().casefold(), + assertion.target_asset_id or assertion.target_label.strip().casefold(), ) if key not in merged or assertion.confidence > merged[key].confidence: merged[key] = assertion return list(merged.values()) + + +def build_orchestrator_client(config: Any) -> ContextualOrchestratorClient: + """Build the governed client from KV config plus the injected secret registry.""" + + if not config.orchestrator_base_url: + raise ValueError("orchestrator base URL is not configured") + validate_runtime_credentials(config) + return ContextualOrchestratorClient( + get_credential_registry(), + base_url=config.orchestrator_base_url, + semantic_model=config.semantic_model, + embedding_model=config.embedding_model, + embedding_dimensions=config.embedding_dimension, + ) diff --git a/src/sdp/file_ontology.py b/src/sdp/file_ontology.py index 66a735d..78f1a2b 100644 --- a/src/sdp/file_ontology.py +++ b/src/sdp/file_ontology.py @@ -3,12 +3,14 @@ from __future__ import annotations import hashlib +import json import re from datetime import datetime +from importlib.resources import files from typing import Any, Literal from urllib.parse import urlsplit -from pydantic import BaseModel, Field, field_validator +from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator from .graph_store import GraphStore @@ -72,6 +74,8 @@ def stable_locator_without_credentials(cls, value: str) -> str: raise ValueError("locator must be an absolute IRI") if parsed.query or parsed.fragment: raise ValueError("locator must not contain a query or fragment") + if parsed.username is not None or parsed.password is not None: + raise ValueError("locator must not contain userinfo credentials") return value @@ -79,6 +83,7 @@ class SemanticAssertion(BaseModel): relation: AssertionRelation target_kind: TargetKind target_label: str = Field(min_length=1) + target_asset_id: str | None = None confidence: float = Field(ge=0, le=1) evidence_chunk_sha256: str evidence_start: int @@ -86,9 +91,21 @@ class SemanticAssertion(BaseModel): method: str = Field(min_length=1) review_status: Literal["proposed", "approved", "rejected"] = "proposed" + @model_validator(mode="after") + def validate_file_target(self) -> "SemanticAssertion": + if self.target_kind == "file_asset": + if not self.target_asset_id or not re.fullmatch( + r"urn:sha256:[0-9a-f]{64}", self.target_asset_id + ): + raise ValueError("file_asset assertions require a content-addressed target_asset_id") + elif self.target_asset_id is not None: + raise ValueError("target_asset_id is only valid for file_asset assertions") + return self + class FileAsset(BaseModel): sha256: str + tenant_id: str = Field(default="demo", min_length=1) title: str = Field(min_length=1) media_type: str = Field(min_length=1) byte_size: int = Field(ge=0) @@ -102,7 +119,7 @@ def asset_id(self) -> str: class FileAssetIngestRequest(FileAsset): - actor: str = Field(default="anonymous", min_length=1) + model_config = ConfigDict(extra="forbid") def _violation(path: str, message: str) -> dict[str, str]: @@ -114,6 +131,47 @@ def _violation(path: str, message: str) -> dict[str, str]: } +def _shacl_violations(asset: FileAsset) -> list[dict[str, str]]: + from pyshacl import validate as shacl_validate + from rdflib import Graph, Namespace + from rdflib.namespace import RDF + + resources = files("sdp").joinpath("resources") + data_graph = Graph().parse( + data=json.dumps(file_asset_jsonld(asset, include_locations=True), ensure_ascii=False), + format="json-ld", + ) + shapes_graph = Graph().parse( + data=resources.joinpath("cwl-file-shapes.ttl").read_text(encoding="utf-8"), + format="turtle", + ) + ontology_graph = Graph().parse( + data=resources.joinpath("cwl-file-profile.ttl").read_text(encoding="utf-8"), + format="turtle", + ) + conforms, results_graph, _text = shacl_validate( + data_graph, + shacl_graph=shapes_graph, + ont_graph=ontology_graph, + inference="rdfs", + advanced=True, + ) + if conforms: + return [] + sh = Namespace("http://www.w3.org/ns/shacl#") + violations: list[dict[str, str]] = [] + for result in results_graph.subjects(RDF.type, sh.ValidationResult): + path = results_graph.value(result, sh.resultPath) + message = results_graph.value(result, sh.resultMessage) + violations.append( + _violation( + str(path) if path is not None else "shacl", + str(message) if message is not None else "SHACL constraint violated", + ) + ) + return violations or [_violation("shacl", "SHACL graph did not conform")] + + def validate_file_asset(asset: FileAsset) -> dict[str, Any]: """Return the SHACL-compatible structural validation report for an asset.""" @@ -144,6 +202,13 @@ def validate_file_asset(asset: FileAsset) -> dict[str, Any]: f"{assertion.relation} requires target_kind={expected_kind}", ) ) + if assertion.target_kind == "file_asset" and not assertion.target_asset_id: + violations.append( + _violation( + "assertions.target_asset_id", + "file relationships require a content-addressed target asset IRI", + ) + ) if not _SHA256_RE.fullmatch(assertion.evidence_chunk_sha256): violations.append( _violation("assertions.evidence_chunk_sha256", "must be a lowercase SHA-256 hex digest") @@ -155,9 +220,12 @@ def validate_file_asset(asset: FileAsset) -> dict[str, Any]: _violation("assertions.evidence_end", "must be greater than evidence_start") ) + violations.extend(_shacl_violations(asset)) + return { "asset_id": asset.asset_id, "shacl_compatible": True, + "shacl_engine": "pyshacl", "shape": "CWLFileAssetShape", "conforms": not violations, "violations": violations, @@ -169,13 +237,26 @@ def concept_id(kind: str, label: str) -> str: return f"urn:cwl:{kind}:{digest}" +def assertion_target_id(assertion: SemanticAssertion) -> str: + if assertion.target_kind == "file_asset": + if not assertion.target_asset_id: + raise ValueError("file assertion target is missing") + return assertion.target_asset_id + return concept_id(assertion.target_kind, assertion.target_label) + + +def distribution_node_id(asset: FileAsset, distribution: StorageDistribution) -> str: + digest = hashlib.sha256(distribution.id.encode("utf-8")).hexdigest()[:24] + return f"urn:cwl:distribution:{asset.sha256}:{digest}" + + def file_asset_jsonld(asset: FileAsset, *, include_locations: bool = False) -> dict[str, Any]: """Export one file asset using the CWL application profile.""" distributions: list[dict[str, Any]] = [] for distribution in asset.distributions: item: dict[str, Any] = { - "@id": f"urn:cwl:distribution:{distribution.id}", + "@id": distribution_node_id(asset, distribution), "@type": "dcat:Distribution", "cwl:storageProvider": distribution.provider, "cwl:endpointId": distribution.endpoint_id, @@ -189,19 +270,22 @@ def file_asset_jsonld(asset: FileAsset, *, include_locations: bool = False) -> d item["cwl:etag"] = distribution.etag distributions.append(item) - subjects = [ - { - "@id": concept_id(assertion.target_kind, assertion.target_label), + subjects = [] + for assertion in asset.assertions: + if assertion.target_kind == "file_asset": + continue + subject = { + "@id": assertion_target_id(assertion), "@type": "skos:Concept", - "skos:prefLabel": assertion.target_label, } - for assertion in asset.assertions - ] + subject["skos:prefLabel"] = assertion.target_label + subjects.append(subject) assertions = [ { "@type": "cwl:SemanticAssertion", - "cwl:relation": f"cwl:{assertion.relation}", - "cwl:target": {"@id": concept_id(assertion.target_kind, assertion.target_label)}, + "cwl:relation": {"@id": f"cwl:{assertion.relation}"}, + "cwl:target": {"@id": assertion_target_id(assertion)}, + "cwl:targetKind": assertion.target_kind, "cwl:confidence": assertion.confidence, "cwl:evidenceChunkSha256": assertion.evidence_chunk_sha256, "cwl:evidenceStart": assertion.evidence_start, @@ -216,6 +300,7 @@ def file_asset_jsonld(asset: FileAsset, *, include_locations: bool = False) -> d "@id": asset.asset_id, "@type": ["dcat:Resource", "prov:Entity", "cwl:FileAsset"], "dcterms:identifier": asset.asset_id, + "cwl:tenantId": asset.tenant_id, "dcterms:title": asset.title, "dcterms:format": asset.media_type, "dcat:byteSize": asset.byte_size, @@ -258,12 +343,28 @@ def _merge_asset(existing: FileAsset | None, incoming: FileAsset) -> FileAsset: distributions = {item.id: item for item in existing.distributions} distributions.update({item.id: item for item in incoming.distributions}) assertions = { - (item.relation, item.target_kind, item.target_label.strip().casefold()): item + ( + item.relation, + item.target_kind, + item.target_asset_id or item.target_label.strip().casefold(), + ): item for item in existing.assertions } for item in incoming.assertions: - key = (item.relation, item.target_kind, item.target_label.strip().casefold()) - if key not in assertions or item.confidence > assertions[key].confidence: + key = ( + item.relation, + item.target_kind, + item.target_asset_id or item.target_label.strip().casefold(), + ) + existing_assertion = assertions.get(key) + if ( + existing_assertion is None + or item.review_status != "proposed" + or ( + existing_assertion.review_status == "proposed" + and item.confidence > existing_assertion.confidence + ) + ): assertions[key] = item return existing.model_copy( update={ @@ -284,7 +385,27 @@ def upsert_file_asset(store: GraphStore, asset: FileAsset) -> FileAsset: if not report["conforms"]: messages = "; ".join(item["message"] for item in report["violations"]) raise ValueError(f"file asset does not conform: {messages}") - merged = _merge_asset(get_file_asset(store, asset.asset_id), asset) + existing_node = store.get_node(asset.asset_id) + if existing_node is not None and existing_node.kind != "file_asset": + raise ValueError("file asset identifier is occupied by a non-file node") + existing = get_file_asset(store, asset.asset_id) + existing_tenant = ( + existing.tenant_id + if existing is not None + else str((existing_node.properties if existing_node else {}).get("tenant_id") or "") + ) + if existing_tenant and existing_tenant != asset.tenant_id: + raise PermissionError("file asset tenant boundary denied") + for assertion in asset.assertions: + if assertion.target_kind != "file_asset": + continue + target = store.get_node(assertion_target_id(assertion)) + if target is not None and target.kind != "file_asset": + raise ValueError("file relationship target is not a file asset node") + target_tenant = str((target.properties if target else {}).get("tenant_id") or "") + if target_tenant and target_tenant != asset.tenant_id: + raise PermissionError("file relationship target tenant boundary denied") + merged = _merge_asset(existing, asset) embedding_text = " ".join( [merged.title] + [assertion.target_label for assertion in merged.assertions] ) @@ -294,37 +415,45 @@ def upsert_file_asset(store: GraphStore, asset: FileAsset) -> FileAsset: label=merged.title, properties={ "profile": "CWL File Knowledge Profile 0.1", + "tenant_id": merged.tenant_id, "file_asset": merged.model_dump(mode="json"), }, text=embedding_text, ) for distribution in merged.distributions: - distribution_id = f"urn:cwl:distribution:{distribution.id}" + distribution_id = distribution_node_id(merged, distribution) store.upsert_node( distribution_id, "distribution", label=distribution.id, - properties=distribution.model_dump(mode="json"), + properties={ + **distribution.model_dump(mode="json"), + "tenant_id": merged.tenant_id, + }, text=f"{distribution.provider} {distribution.endpoint_id}", ) store.upsert_edge("DISTRIBUTION", merged.asset_id, distribution_id) for assertion in merged.assertions: - target_id = concept_id(assertion.target_kind, assertion.target_label) - target_kind = ( - "file_asset_reference" if assertion.target_kind == "file_asset" else assertion.target_kind - ) - store.upsert_node( - target_id, - target_kind, - label=assertion.target_label, - properties={ - "skos_pref_label": assertion.target_label, - "review_status": assertion.review_status, - }, - text=assertion.target_label, - ) + target_id = assertion_target_id(assertion) + target_kind = assertion.target_kind + if target_kind != "file_asset" or store.get_node(target_id) is None: + store.upsert_node( + target_id, + target_kind, + label=assertion.target_label, + properties={ + "skos_pref_label": assertion.target_label, + "review_status": assertion.review_status, + **( + {"tenant_id": merged.tenant_id} + if target_kind == "file_asset" + else {} + ), + }, + text=assertion.target_label, + ) store.upsert_edge( _EDGE_TYPES[assertion.relation], merged.asset_id, diff --git a/src/sdp/file_pilot.py b/src/sdp/file_pilot.py index a236895..c6dc65f 100644 --- a/src/sdp/file_pilot.py +++ b/src/sdp/file_pilot.py @@ -163,6 +163,7 @@ def main() -> int: base_url=orchestrator_url, semantic_model=args.semantic_model or config.semantic_model, embedding_model=args.embedding_model or config.embedding_model, + embedding_dimensions=config.embedding_dimension, ) summary = run_local_pilot( diff --git a/src/sdp/graph_models.py b/src/sdp/graph_models.py index 55fec0a..e773059 100644 --- a/src/sdp/graph_models.py +++ b/src/sdp/graph_models.py @@ -4,10 +4,14 @@ from typing import Any, Dict, List, Optional -from pydantic import BaseModel, Field +from pydantic import BaseModel, ConfigDict, Field -class GraphNodeRequest(BaseModel): +class GovernedGraphRequest(BaseModel): + model_config = ConfigDict(extra="forbid") + + +class GraphNodeRequest(GovernedGraphRequest): node_id: str = Field(min_length=1) kind: str = Field(min_length=1, description="node label, e.g. concept/dataset/column") label: Optional[str] = None @@ -16,18 +20,16 @@ class GraphNodeRequest(BaseModel): default=None, description="text to embed for semantic search; falls back to label/node_id", ) - actor: str = Field(default="anonymous", description="authenticated subject for authz") -class GraphEdgeRequest(BaseModel): +class GraphEdgeRequest(GovernedGraphRequest): edge_type: str = Field(min_length=1, description="relationship, e.g. broader/related/mapping/lineage") source_id: str = Field(min_length=1) target_id: str = Field(min_length=1) properties: Dict[str, Any] = Field(default_factory=dict) - actor: str = Field(default="anonymous", description="authenticated subject for authz") -class OntologyConceptRequest(BaseModel): +class OntologyConceptRequest(GovernedGraphRequest): concept: str = Field(min_length=1) definition: str = "" aliases: List[str] = Field(default_factory=list) @@ -35,10 +37,9 @@ class OntologyConceptRequest(BaseModel): narrower: List[str] = Field(default_factory=list) related: List[str] = Field(default_factory=list) multilingual: List[str] = Field(default_factory=list) - actor: str = Field(default="anonymous", description="authenticated subject for authz") -class GraphTraversalRequest(BaseModel): +class GraphTraversalRequest(GovernedGraphRequest): """Safe, parameterized traversal request. Raw openCypher is intentionally NOT accepted: the traversal is expressed as a @@ -52,11 +53,9 @@ class GraphTraversalRequest(BaseModel): ) direction: str = Field(default="both", pattern="^(out|in|both)$") max_depth: int = Field(default=2, ge=1, le=6) - actor: str = Field(default="anonymous", description="authenticated subject for authz") -class SemanticSearchRequest(BaseModel): +class SemanticSearchRequest(GovernedGraphRequest): query: str = Field(min_length=1) kind: Optional[str] = Field(default=None, description="restrict to a node kind") limit: int = Field(default=5, ge=1, le=50) - actor: str = Field(default="anonymous", description="authenticated subject for authz") diff --git a/src/sdp/graph_store.py b/src/sdp/graph_store.py index c1ffeaf..c4b2075 100644 --- a/src/sdp/graph_store.py +++ b/src/sdp/graph_store.py @@ -69,6 +69,24 @@ def _normalize(value: str) -> str: return value.strip().replace("_", " ").lower() +def _is_governed_file_identifier(value: Any) -> bool: + text = str(value or "") + return text.startswith("urn:sha256:") or text.startswith("urn:cwl:distribution:") + + +def _validate_concept_file_boundaries(payload: Dict[str, Any]) -> None: + values = [ + payload.get("concept"), + payload.get("broader"), + *list(payload.get("narrower", [])), + *list(payload.get("related", [])), + *list(payload.get("aliases", [])), + *list(payload.get("multilingual", [])), + ] + if any(_is_governed_file_identifier(value) for value in values if value is not None): + raise ValueError("ontology concepts cannot use governed file identifiers") + + _logger = logging.getLogger(__name__) # openCypher / AGE relationship types occupy an *identifier* position in the @@ -175,7 +193,12 @@ def traverse( @abstractmethod def semantic_search( - self, query: str, *, kind: Optional[str] = None, limit: int = 5 + self, + query: str, + *, + kind: Optional[str] = None, + limit: int = 5, + tenant_id: Optional[str] = None, ) -> List[Dict[str, Any]]: """Return the nearest graph nodes for a semantic query.""" @@ -263,6 +286,7 @@ def upsert_edge( return edge def upsert_concept(self, payload: Dict[str, Any]) -> Dict[str, Any]: + _validate_concept_file_boundaries(payload) concept = payload["concept"].strip() record = { "concept": concept, @@ -388,7 +412,12 @@ def traverse( } def semantic_search( - self, query: str, *, kind: Optional[str] = None, limit: int = 5 + self, + query: str, + *, + kind: Optional[str] = None, + limit: int = 5, + tenant_id: Optional[str] = None, ) -> List[Dict[str, Any]]: query_vec = [float(component) for component in self._embedder(query)] if len(query_vec) != self.dimension: @@ -398,6 +427,17 @@ def semantic_search( node = self._nodes[node_id] if kind and node.kind != kind: continue + node_tenant = str( + node.properties.get("tenant_id") + or (node.properties.get("file_asset") or {}).get("tenant_id") + or "" + ) + if ( + tenant_id is not None + and node.kind in {"file_asset", "distribution"} + and node_tenant != tenant_id + ): + continue score = cosine_similarity(query_vec, vector) scored.append( { @@ -592,6 +632,7 @@ def upsert_edge( def upsert_concept(self, payload: Dict[str, Any]) -> Dict[str, Any]: from sqlalchemy import text as sql + _validate_concept_file_boundaries(payload) concept = payload["concept"].strip() record = { "concept": concept, @@ -776,7 +817,12 @@ def _edges_within(self, node_ids: Iterable[str], edge_types: Optional[List[str]] return result def semantic_search( - self, query: str, *, kind: Optional[str] = None, limit: int = 5 + self, + query: str, + *, + kind: Optional[str] = None, + limit: int = 5, + tenant_id: Optional[str] = None, ) -> List[Dict[str, Any]]: from sqlalchemy import text as sql @@ -791,9 +837,19 @@ def semantic_search( "LEFT JOIN graph_nodes n ON n.node_id = e.node_id " ) params: Dict[str, Any] = {"vec": vector, "limit": limit} + conditions: List[str] = [] if kind: - stmt += "WHERE e.node_kind = :kind " + conditions.append("e.node_kind = :kind") params["kind"] = kind + if tenant_id is not None: + conditions.append( + "(e.node_kind NOT IN ('file_asset', 'distribution') OR " + "COALESCE(n.node_properties ->> 'tenant_id', " + "n.node_properties -> 'file_asset' ->> 'tenant_id') = :tenant_id)" + ) + params["tenant_id"] = tenant_id + if conditions: + stmt += "WHERE " + " AND ".join(conditions) + " " stmt += "ORDER BY e.embedding <=> CAST(:vec AS vector) LIMIT :limit" with self._engine.connect() as conn: rows = conn.execute(sql(stmt), params).fetchall() @@ -857,6 +913,15 @@ def stats(self) -> Dict[str, int]: _STORE: Optional[GraphStore] = None +def _configured_embedder(config: AppConfig) -> Optional[Callable[[str], List[float]]]: + if not config.orchestrator_base_url: + return None + # Lazy import avoids the file-ontology -> graph-store dependency cycle. + from .document_semantics import build_orchestrator_client + + return build_orchestrator_client(config).embed_one + + def build_store() -> GraphStore: """Construct the configured backend for the current bootstrap config. @@ -876,12 +941,18 @@ def build_store() -> GraphStore: bootstrap = load_bootstrap() config = get_app_config() backend = config.graph_backend + embedder = _configured_embedder(config) if backend == "memory": - return InMemoryGraphStore(config=config) + return InMemoryGraphStore(config=config, embedder=embedder) if bootstrap.has_database: - store = PostgresGraphStore(bootstrap.database_dsn, config=config) + store_kwargs = {"embedder": embedder} if embedder is not None else {} + store = PostgresGraphStore( + bootstrap.database_dsn, + config=config, + **store_kwargs, + ) readiness = store.readiness() if not readiness.get("ready"): raise RuntimeError( @@ -897,13 +968,20 @@ def build_store() -> GraphStore: "(bootstrap transport SDP_DATABASE_DSN is unset)." ) - return InMemoryGraphStore(config=config) + return InMemoryGraphStore(config=config, embedder=embedder) def get_store() -> GraphStore: + """Build and seed the active store lazily after runtime credentials exist.""" + global _STORE if _STORE is None: _STORE = build_store() + # Lazy import avoids the seed -> graph_store dependency cycle. Passing + # the concrete store prevents seed_store from recursively calling us. + from .seed import seed_store + + seed_store(_STORE) return _STORE diff --git a/src/sdp/policy.py b/src/sdp/policy.py index ae02b63..c6979b2 100644 --- a/src/sdp/policy.py +++ b/src/sdp/policy.py @@ -24,6 +24,86 @@ def _has_reader_role(subject: str) -> bool: return has_role(subject, "data-analyst", "admin", "platform-admin", "security") +def evaluate_file_asset_access( + actor_context: object, + resource: str, + action: str, + resource_tenant_id: str, +) -> PolicyDecision: + """Evaluate and record resource-scoped policy for one governed file asset.""" + + subject = str(getattr(actor_context, "subject", "anonymous")) + actor_tenant_id = str(getattr(actor_context, "tenant_id", "")) + roles = set(getattr(actor_context, "roles", [])) + action_key = action.lower() + decision_base = { + "subject": subject, + "resource": resource, + "action": action, + "decision_id": str(uuid4()), + } + obligations = { + "tenant_id": resource_tenant_id, + "actor_tenant_id": actor_tenant_id, + } + if "platform-admin" not in roles and actor_tenant_id != resource_tenant_id: + return _decision( + **decision_base, + effect="deny", + reason="tenant boundary denied", + obligations=obligations, + ) + if action_key == "create_file_asset": + allowed = bool({"admin", "platform-admin"}.intersection(roles)) + required_role = "admin" + elif action_key == "read_file_locations": + allowed = bool({"admin", "platform-admin"}.intersection(roles)) + required_role = "admin" + else: + allowed = bool( + {"data-analyst", "admin", "platform-admin", "security"}.intersection(roles) + ) + required_role = "data-analyst" + if not allowed: + return _decision( + **decision_base, + effect="deny", + reason="file asset policy denied", + obligations={**obligations, "required_role": required_role}, + ) + return _decision( + **decision_base, + effect="allow", + reason="file asset tenant and role policy satisfied", + obligations=obligations, + ) + + +def evaluate_graph_access( + actor_context: object, + resource: str, + action: str, +) -> PolicyDecision: + """Evaluate and record OIDC-derived access to the generic graph surface.""" + + subject = str(getattr(actor_context, "subject", "anonymous")) + roles = set(getattr(actor_context, "roles", [])) + write = action.lower() == "write_graph" + allowed = bool( + ({"admin", "platform-admin"} if write else {"data-analyst", "admin", "platform-admin", "security"}) + .intersection(roles) + ) + return _decision( + subject=subject, + resource=resource, + action=action, + decision_id=str(uuid4()), + effect="allow" if allowed else "deny", + reason="graph role policy satisfied" if allowed else "graph role policy denied", + obligations={"required_role": "admin" if write else "data-analyst"}, + ) + + def evaluate(subject: str, resource: str, action: str, purpose: str) -> PolicyDecision: action_key = action.lower() decision_id = str(uuid4()) diff --git a/src/sdp/resources/cwl-file-profile.ttl b/src/sdp/resources/cwl-file-profile.ttl new file mode 100644 index 0000000..685a063 --- /dev/null +++ b/src/sdp/resources/cwl-file-profile.ttl @@ -0,0 +1,42 @@ +@prefix cwl: . +@prefix dcat: . +@prefix owl: . +@prefix prov: . +@prefix rdfs: . +@prefix skos: . +@prefix xsd: . + +cwl: a owl:Ontology ; + owl:versionInfo "0.1.0" . + +cwl:FileAsset a owl:Class ; + rdfs:subClassOf dcat:Resource, prov:Entity . + +cwl:BusinessProject a owl:Class ; owl:subClassOf skos:Concept . +cwl:System a owl:Class ; owl:subClassOf skos:Concept . +cwl:WorkPhase a owl:Class ; owl:subClassOf skos:Concept . +cwl:ArtifactType a owl:Class ; owl:subClassOf skos:Concept . +cwl:SemanticAssertion a owl:Class . + +cwl:belongsToProject a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:BusinessProject . +cwl:usesSystem a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:System . +cwl:hasWorkPhase a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:WorkPhase . +cwl:hasArtifactType a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range cwl:ArtifactType . +cwl:hasTopic a owl:ObjectProperty ; + owl:domain cwl:FileAsset ; owl:range skos:Concept . +cwl:wasDerivedFrom a owl:ObjectProperty ; + owl:subPropertyOf prov:wasDerivedFrom . +cwl:previousVersion a owl:ObjectProperty . + +cwl:confidence a owl:DatatypeProperty ; owl:range xsd:decimal . +cwl:tenantId a owl:DatatypeProperty ; owl:range xsd:string . +cwl:targetKind a owl:DatatypeProperty ; owl:range xsd:string . +cwl:reviewStatus a owl:DatatypeProperty ; owl:range xsd:string . +cwl:evidenceChunkSha256 a owl:DatatypeProperty ; owl:range xsd:string . +cwl:evidenceStart a owl:DatatypeProperty ; owl:range xsd:nonNegativeInteger . +cwl:evidenceEnd a owl:DatatypeProperty ; owl:range xsd:positiveInteger . +cwl:extractionMethod a owl:DatatypeProperty ; owl:range xsd:string . diff --git a/src/sdp/resources/cwl-file-shapes.ttl b/src/sdp/resources/cwl-file-shapes.ttl new file mode 100644 index 0000000..7fee210 --- /dev/null +++ b/src/sdp/resources/cwl-file-shapes.ttl @@ -0,0 +1,73 @@ +@prefix cwl: . +@prefix dcat: . +@prefix dcterms: . +@prefix sh: . +@prefix spdx: . +@prefix xsd: . + +cwl:FileAssetShape a sh:NodeShape ; + sh:targetClass cwl:FileAsset ; + sh:property [ + sh:path dcterms:identifier ; sh:minCount 1 ; sh:maxCount 1 ; + sh:pattern "^urn:sha256:[0-9a-f]{64}$" + ] ; + sh:property [ sh:path dcterms:title ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path cwl:tenantId ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path dcterms:format ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path dcat:byteSize ; sh:minCount 1 ; sh:maxCount 1 ; sh:minInclusive 0 ] ; + sh:property [ sh:path spdx:checksum ; sh:minCount 1 ; sh:maxCount 1 ; sh:node cwl:ChecksumShape ] ; + sh:property [ sh:path dcat:distribution ; sh:minCount 1 ; sh:node cwl:DistributionShape ] ; + sh:property [ sh:path cwl:assertion ; sh:node cwl:SemanticAssertionShape ] . + +cwl:ChecksumShape a sh:NodeShape ; + sh:property [ + sh:path spdx:checksumValue ; sh:minCount 1 ; sh:maxCount 1 ; + sh:pattern "^[0-9a-f]{64}$" + ] . + +cwl:DistributionShape a sh:NodeShape ; + sh:targetClass dcat:Distribution ; + sh:property [ + sh:path cwl:storageProvider ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( "filesystem" "s3" "s3_compatible" "azure_blob" ) + ] ; + sh:property [ sh:path cwl:endpointId ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] ; + sh:property [ sh:path dcat:accessURL ; sh:minCount 1 ; sh:maxCount 1 ] . + +cwl:SemanticAssertionShape a sh:NodeShape ; + sh:targetClass cwl:SemanticAssertion ; + sh:property [ + sh:path cwl:relation ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( + cwl:belongsToProject cwl:usesSystem cwl:hasWorkPhase + cwl:hasArtifactType cwl:hasTopic cwl:wasDerivedFrom cwl:previousVersion + ) + ] ; + sh:property [ sh:path cwl:target ; sh:minCount 1 ; sh:maxCount 1 ; sh:nodeKind sh:IRI ] ; + sh:property [ + sh:path cwl:targetKind ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( + "business_project" "system" "work_phase" + "artifact_type" "topic" "file_asset" + ) + ] ; + sh:property [ + sh:path cwl:confidence ; sh:minCount 1 ; sh:maxCount 1 ; + sh:minInclusive 0 ; sh:maxInclusive 1 + ] ; + sh:property [ + sh:path cwl:reviewStatus ; sh:minCount 1 ; sh:maxCount 1 ; + sh:in ( "proposed" "approved" "rejected" ) + ] ; + sh:property [ + sh:path cwl:evidenceChunkSha256 ; sh:minCount 1 ; sh:maxCount 1 ; + sh:pattern "^[0-9a-f]{64}$" + ] ; + sh:property [ + sh:path cwl:evidenceStart ; sh:minCount 1 ; sh:maxCount 1 ; + sh:minInclusive 0 ; sh:lessThan cwl:evidenceEnd + ] ; + sh:property [ + sh:path cwl:evidenceEnd ; sh:minCount 1 ; sh:maxCount 1 ; sh:minInclusive 1 + ] ; + sh:property [ sh:path cwl:extractionMethod ; sh:minCount 1 ; sh:maxCount 1 ; sh:minLength 1 ] . diff --git a/src/sdp/storage_readers.py b/src/sdp/storage_readers.py index 1d95ab4..d20898c 100644 --- a/src/sdp/storage_readers.py +++ b/src/sdp/storage_readers.py @@ -225,7 +225,11 @@ def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: if ref.distribution.version_id else {} ) - data = self.client.get_blob_client(ref.object_key, **kwargs).download_blob().readall() + downloader = self.client.get_blob_client(ref.object_key, **kwargs).download_blob( + offset=0, + length=max_bytes + 1, + ) + data = downloader.readall() if len(data) > max_bytes: raise ValueError("object exceeds maximum size") return data diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index be13099..564cac6 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -2,6 +2,7 @@ import json from copy import deepcopy +from importlib.resources import files from io import BytesIO from pathlib import Path from types import SimpleNamespace @@ -11,6 +12,7 @@ from fastapi.testclient import TestClient from pypdf import PdfWriter +from sdp_core import ActorContext from sdp import api as api_module from sdp.api import app @@ -125,6 +127,26 @@ def test_machine_readable_profile_requires_assertion_provenance(): assert "cwl:extractionMethod a owl:DatatypeProperty" in profile assert "sh:path cwl:extractionMethod ; sh:minCount 1" in shapes + assert "owl:equivalentClass" not in profile + assert "rdfs:subClassOf dcat:Resource" in profile + assert "sh:minInclusive 0" in shapes + assert "sh:maxInclusive 1" in shapes + assert "sh:lessThan cwl:evidenceEnd" in shapes + assert "sh:in ( \"proposed\" \"approved\" \"rejected\" )" in shapes + + +def test_machine_readable_profile_is_packaged_with_the_runtime(): + resources = files("sdp").joinpath("resources") + + assert resources.joinpath("cwl-file-profile.ttl").is_file() + assert resources.joinpath("cwl-file-shapes.ttl").is_file() + + +def test_file_asset_validation_executes_shacl_engine(): + report = validate_file_asset(sample_asset()) + + assert report["conforms"] is True + assert report["shacl_engine"] == "pyshacl" @pytest.mark.parametrize( @@ -156,6 +178,11 @@ def test_distribution_rejects_locator_query_to_avoid_secret_leak(): sample_distribution(locator="https://objects.example/report.pdf?sig=secret") +def test_distribution_rejects_locator_userinfo_to_avoid_secret_leak(): + with pytest.raises(ValueError, match="userinfo"): + sample_distribution(locator="https://user:secret@objects.example/report.pdf") + + def test_filesystem_reader_stays_inside_root_and_reads_bytes(tmp_path): path = tmp_path / "효성중공업 VOC.txt" path.write_text("C-Cube PoC", encoding="utf-8") @@ -220,7 +247,12 @@ def list_blobs(self, *, name_starts_with): def get_blob_client(self, name, **kwargs): assert name == "reports/a.docx" assert kwargs == {"version_id": "version-a"} - return SimpleNamespace(download_blob=lambda: SimpleNamespace(readall=lambda: b"content")) + def download_blob(*, offset, length): + assert offset == 0 + assert length == 21 + return SimpleNamespace(readall=lambda: b"content") + + return SimpleNamespace(download_blob=download_blob) def test_s3_and_azure_readers_use_injected_clients(): @@ -278,6 +310,17 @@ def test_openxml_text_is_extracted_without_office_dependency(filename, member): assert "효성중공업 VOC C-Cube" in document.text +def test_openxml_rejects_oversized_uncompressed_member(): + payload = BytesIO() + with ZipFile(payload, "w", ZIP_DEFLATED) as archive: + archive.writestr("word/document.xml", b"x" * (8 * 1024 * 1024 + 1)) + + document = extract_document_text("bomb.docx", payload.getvalue()) + + assert document.status == "extraction_failed" + assert document.error == "ValueError" + + def test_chunks_are_bounded_overlapped_and_content_addressed(): chunks = chunk_text("가" * 13_000, max_chars=6_000, overlap=300, max_total_chars=24_000) @@ -419,6 +462,7 @@ def transport(request, timeout): ), base_url="https://orchestrator.example/", transport=transport, + embedding_dimensions=2, ) assert client.embed(["alpha", "beta"]) == [[1.0, 0.0], [0.0, 1.0]] @@ -426,10 +470,27 @@ def transport(request, timeout): assert captured["payload"] == { "model": "text-embedding-3-small", "input": ["alpha", "beta"], + "dimensions": 2, "metadata": {"service": "semantic-data-portal"}, } +def test_orchestrator_client_rejects_wrong_embedding_dimension(): + client = ContextualOrchestratorClient( + EphemeralCredentialRegistry( + {"CONTEXTUAL_ORCHESTRATOR_TOKEN": "orchestrator-token"} + ), + base_url="https://orchestrator.example", + embedding_dimensions=2, + transport=lambda *_: { + "data": [{"index": 0, "embedding": [1.0, 0.0, 0.0]}], + }, + ) + + with pytest.raises(ValueError, match="dimension"): + client.embed(["alpha"]) + + def test_graph_store_can_use_orchestrator_embedding_client(): embedded = [] @@ -519,23 +580,102 @@ def test_upsert_file_asset_merges_same_content_distributions_and_projects_assert assert "C-Cube PoC" not in json.dumps(node.properties, ensure_ascii=False) -def sample_request(actor: str) -> dict[str, object]: - return {**sample_asset().model_dump(mode="json"), "actor": actor} +def test_file_relationship_projects_content_addressed_target_id(): + store = InMemoryGraphStore() + target_id = f"urn:sha256:{'c' * 64}" + assertion = sample_assertion( + relation="previousVersion", + target_kind="file_asset", + target_label="이전 VOC 보고서", + target_asset_id=target_id, + ) + + upsert_file_asset(store, sample_asset(assertions=[assertion])) + + graph = store.traverse(sample_asset().asset_id, direction="out", max_depth=1) + edge = next(item for item in graph["edges"] if item["edge_type"] == "PREVIOUS_VERSION") + assert edge["target_id"] == target_id + + +def test_file_relationship_does_not_overwrite_existing_target_asset(): + store = InMemoryGraphStore() + target = sample_asset( + sha256="c" * 64, + title="원본 VOC 보고서", + distributions=[ + sample_distribution(id="target-dist", locator="file:///D:/Documents/original.docx") + ], + assertions=[], + ) + upsert_file_asset(store, target) + before = store.get_node(target.asset_id) + assertion = sample_assertion( + relation="previousVersion", + target_kind="file_asset", + target_label="이전 VOC 보고서", + target_asset_id=target.asset_id, + ) + + upsert_file_asset(store, sample_asset(assertions=[assertion])) + + after = store.get_node(target.asset_id) + assert before is not None and after is not None + assert after.label == before.label + assert after.properties["file_asset"] == before.properties["file_asset"] + + +def test_explicit_steward_review_supersedes_model_confidence(): + store = InMemoryGraphStore() + proposed = sample_assertion(confidence=0.99, review_status="proposed") + approved = sample_assertion(confidence=0.60, review_status="approved") + + upsert_file_asset(store, sample_asset(assertions=[proposed])) + merged = upsert_file_asset(store, sample_asset(assertions=[approved])) + + assert merged.assertions[0].review_status == "approved" + assert merged.assertions[0].confidence == 0.60 + + +def sample_request() -> dict[str, object]: + return sample_asset().model_dump(mode="json") + + +def auth_headers(role: str) -> dict[str, str]: + return {"Authorization": f"Bearer {role}-token"} + + +def install_fake_oidc(monkeypatch): + def verify(token): + identity = token.removesuffix("-token") + tenant_id, _, role = identity.partition(":") + if not role: + tenant_id, role = "demo", tenant_id + roles = { + "admin": ["admin", "data-analyst"], + "platform-admin": ["platform-admin", "admin", "data-analyst"], + "analyst": ["data-analyst"], + }.get(role, ["data-analyst"]) + return ActorContext(subject=identity, tenant_id=tenant_id, roles=roles), {} + + monkeypatch.setattr(api_module.authz, "verify_oidc_jwks_token", verify) def test_file_asset_api_requires_policy_and_redacts_jsonld_locator(monkeypatch): store = InMemoryGraphStore() monkeypatch.setattr(api_module, "get_store", lambda: store) + install_fake_oidc(monkeypatch) - denied = client.post("/file-assets", json=sample_request("analyst")) - created = client.post("/file-assets", json=sample_request("admin")) + denied = client.post("/file-assets", json=sample_request(), headers=auth_headers("analyst")) + impersonation = client.post("/file-assets", json={**sample_request(), "actor": "admin"}) + created = client.post("/file-assets", json=sample_request(), headers=auth_headers("admin")) assert denied.status_code == 403 + assert impersonation.status_code == 401 assert created.status_code == 200 asset_id = created.json()["asset_id"] - detail = client.get(f"/file-assets/{asset_id}", params={"actor": "analyst"}) - graph_node = client.get(f"/graph/nodes/{asset_id}") - exported = client.get(f"/file-assets/{asset_id}/jsonld", params={"actor": "analyst"}) + detail = client.get(f"/file-assets/{asset_id}", headers=auth_headers("analyst")) + graph_node = client.get(f"/graph/nodes/{asset_id}", headers=auth_headers("analyst")) + exported = client.get(f"/file-assets/{asset_id}/jsonld", headers=auth_headers("analyst")) assert detail.status_code == 200 assert "file:///" not in detail.text assert graph_node.status_code == 200 @@ -545,15 +685,18 @@ def test_file_asset_api_requires_policy_and_redacts_jsonld_locator(monkeypatch): location_denied = client.get( f"/file-assets/{asset_id}/jsonld", - params={"actor": "analyst", "include_locations": True}, + params={"actor": "admin", "include_locations": True}, + headers=auth_headers("analyst"), ) location_allowed = client.get( f"/file-assets/{asset_id}/jsonld", - params={"actor": "admin", "include_locations": True}, + params={"include_locations": True}, + headers=auth_headers("admin"), ) detail_location_allowed = client.get( f"/file-assets/{asset_id}", - params={"actor": "admin", "include_locations": True}, + params={"include_locations": True}, + headers=auth_headers("admin"), ) assert location_denied.status_code == 403 assert location_allowed.status_code == 200 @@ -565,12 +708,224 @@ def test_file_asset_api_requires_policy_and_redacts_jsonld_locator(monkeypatch): def test_file_asset_api_exposes_validation_and_404(monkeypatch): store = InMemoryGraphStore() monkeypatch.setattr(api_module, "get_store", lambda: store) - created = client.post("/file-assets", json=sample_request("admin")) + install_fake_oidc(monkeypatch) + created = client.post("/file-assets", json=sample_request(), headers=auth_headers("admin")) asset_id = created.json()["asset_id"] - report = client.get(f"/file-assets/{asset_id}/validate", params={"actor": "analyst"}) - missing = client.get(f"/file-assets/urn:sha256:{'f' * 64}", params={"actor": "analyst"}) + report = client.get(f"/file-assets/{asset_id}/validate", headers=auth_headers("analyst")) + missing = client.get( + f"/file-assets/urn:sha256:{'f' * 64}", headers=auth_headers("analyst") + ) assert report.status_code == 200 assert report.json()["conforms"] is True assert missing.status_code == 404 + + +def test_file_asset_api_enforces_tenant_ownership(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + install_fake_oidc(monkeypatch) + created = client.post( + "/file-assets", + json=sample_request(), + headers=auth_headers("demo:admin"), + ) + asset_id = created.json()["asset_id"] + + denied = client.get( + f"/file-assets/{asset_id}", + headers=auth_headers("external:analyst"), + ) + allowed = client.get( + f"/file-assets/{asset_id}", + headers=auth_headers("external:platform-admin"), + ) + + assert denied.status_code == 403 + assert allowed.status_code == 200 + assert get_file_asset(store, asset_id).tenant_id == "demo" + + +def test_cross_tenant_same_sha_ingest_is_rejected_without_merging(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + install_fake_oidc(monkeypatch) + first = client.post( + "/file-assets", + json=sample_request(), + headers=auth_headers("demo:admin"), + ) + second_payload = sample_request() + second_payload["distributions"] = [ + sample_distribution( + id="external-copy", + locator="file:///D:/External/report.docx", + ).model_dump(mode="json") + ] + + denied = client.post( + "/file-assets", + json=second_payload, + headers=auth_headers("external:admin"), + ) + + assert first.status_code == 200 + assert denied.status_code == 403 + stored = get_file_asset(store, first.json()["asset_id"]) + assert stored is not None + assert stored.tenant_id == "demo" + assert [distribution.id for distribution in stored.distributions] == ["dist-local"] + + +def test_cross_tenant_file_relationship_target_is_rejected(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + install_fake_oidc(monkeypatch) + target = sample_asset(sha256="c" * 64, assertions=[]) + upsert_file_asset(store, target) + relation = sample_assertion( + relation="previousVersion", + target_kind="file_asset", + target_label="다른 tenant 문서", + target_asset_id=target.asset_id, + ) + source = sample_asset(sha256="d" * 64, assertions=[relation]).model_dump(mode="json") + + denied = client.post( + "/file-assets", + json=source, + headers=auth_headers("external:admin"), + ) + + assert denied.status_code == 403 + assert get_file_asset(store, f"urn:sha256:{'d' * 64}") is None + + +def test_graph_routes_require_oidc_filter_tenant_and_redact_locator(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + install_fake_oidc(monkeypatch) + created = client.post( + "/file-assets", + json=sample_request(), + headers=auth_headers("demo:admin"), + ) + asset_id = created.json()["asset_id"] + + unauthenticated = client.post( + "/graph/query", + json={"start_id": asset_id, "actor": "analyst", "max_depth": 1}, + ) + cross_tenant = client.post( + "/graph/query", + json={"start_id": asset_id, "max_depth": 1}, + headers=auth_headers("external:analyst"), + ) + visible = client.post( + "/graph/query", + json={"start_id": asset_id, "max_depth": 1}, + headers=auth_headers("demo:analyst"), + ) + cross_tenant_search = client.post( + "/search/semantic", + json={"query": "효성중공업 VOC 종료보고서", "kind": "file_asset", "limit": 50}, + headers=auth_headers("external:analyst"), + ) + + assert unauthenticated.status_code == 401 + assert cross_tenant.status_code == 403 + assert visible.status_code == 200 + assert "file:///" not in visible.text + assert asset_id not in { + item["node_id"] for item in cross_tenant_search.json()["results"] + } + + +def test_semantic_search_filters_tenant_before_limit(): + config = SimpleNamespace( + embedding_dimension=2, + semantic_search_default_limit=5, + traversal_max_depth=4, + ) + store = InMemoryGraphStore(config=config, embedder=lambda _text: [1.0, 0.0]) + for index in range(501): + store.upsert_node( + f"external-{index}", + "file_asset", + properties={"tenant_id": "external"}, + text="same score", + ) + store.upsert_node( + "demo-file", + "file_asset", + properties={"tenant_id": "demo"}, + text="same score", + ) + + results = store.semantic_search( + "same score", + kind="file_asset", + limit=1, + tenant_id="demo", + ) + + assert [item["node_id"] for item in results] == ["demo-file"] + + +def test_generic_graph_mutation_cannot_create_governed_file_node(monkeypatch): + install_fake_oidc(monkeypatch) + + response = client.post( + "/graph/nodes", + json={"node_id": "forged", "kind": "file_asset"}, + headers=auth_headers("demo:admin"), + ) + + assert response.status_code == 400 + + reserved_id = client.post( + "/graph/nodes", + json={"node_id": f"urn:sha256:{'e' * 64}", "kind": "concept"}, + headers=auth_headers("demo:admin"), + ) + reserved_edge = client.post( + "/graph/edges", + json={ + "edge_type": "related", + "source_id": "고객", + "target_id": f"urn:sha256:{'e' * 64}", + }, + headers=auth_headers("demo:admin"), + ) + + assert reserved_id.status_code == 400 + assert reserved_edge.status_code == 400 + + +def test_ontology_concept_cannot_overwrite_or_reference_file_identifier(monkeypatch): + store = InMemoryGraphStore() + monkeypatch.setattr(api_module, "get_store", lambda: store) + install_fake_oidc(monkeypatch) + asset = sample_asset() + upsert_file_asset(store, asset) + + overwrite = client.post( + "/ontology/concepts", + json={"concept": asset.asset_id}, + headers=auth_headers("demo:admin"), + ) + reference = client.post( + "/ontology/concepts", + json={"concept": "안전한 개념", "related": [asset.asset_id]}, + headers=auth_headers("demo:admin"), + ) + + assert overwrite.status_code == 400 + assert reference.status_code == 400 + node = store.get_node(asset.asset_id) + assert node is not None and node.kind == "file_asset" + assert node.properties["tenant_id"] == "demo" + + with pytest.raises(ValueError, match="governed file identifiers"): + store.upsert_concept({"concept": asset.asset_id}) diff --git a/tests/test_graph_engine.py b/tests/test_graph_engine.py index c5150ed..ab0cf62 100644 --- a/tests/test_graph_engine.py +++ b/tests/test_graph_engine.py @@ -5,6 +5,8 @@ """ from fastapi.testclient import TestClient +import pytest +from sdp_core import ActorContext from sdp import api as api_module from sdp import config as config_module @@ -15,6 +17,18 @@ client = TestClient(app) +ADMIN_HEADERS = {"Authorization": "Bearer admin-token"} +ANALYST_HEADERS = {"Authorization": "Bearer analyst-token"} + + +@pytest.fixture(autouse=True) +def fake_graph_oidc(monkeypatch): + def verify(token): + role = token.removesuffix("-token") + roles = ["admin", "data-analyst"] if role == "admin" else ["data-analyst"] + return ActorContext(subject=role, tenant_id="demo", roles=roles), {} + + monkeypatch.setattr(api_module.authz, "verify_oidc_jwks_token", verify) # --- readiness --------------------------------------------------------------- @@ -68,18 +82,20 @@ def test_health_still_present(): def test_ingest_node_edge_and_fetch(): node = client.post( "/graph/nodes", - json={"node_id": "svc-A", "kind": "service", "label": "Billing", "text": "billing invoices", "actor": "admin"}, + json={"node_id": "svc-A", "kind": "service", "label": "Billing", "text": "billing invoices"}, + headers=ADMIN_HEADERS, ) assert node.status_code == 200 assert node.json()["node"]["node_id"] == "svc-A" - fetched = client.get("/graph/nodes/svc-A") + fetched = client.get("/graph/nodes/svc-A", headers=ANALYST_HEADERS) assert fetched.status_code == 200 assert fetched.json()["kind"] == "service" edge = client.post( "/graph/edges", - json={"edge_type": "depends_on", "source_id": "svc-A", "target_id": "고객", "actor": "admin"}, + json={"edge_type": "depends_on", "source_id": "svc-A", "target_id": "고객"}, + headers=ADMIN_HEADERS, ) assert edge.status_code == 200 assert edge.json()["edge"]["target_id"] == "고객" @@ -94,8 +110,8 @@ def test_ingest_concept_creates_traversable_node(): "aliases": ["subscription", "정기결제"], "related": ["매출"], "multilingual": ["subscription"], - "actor": "admin", }, + headers=ADMIN_HEADERS, ) assert resp.status_code == 200 graph = client.get("/ontology/term/subscription/graph") @@ -105,7 +121,7 @@ def test_ingest_concept_creates_traversable_node(): def test_missing_node_returns_404(): - assert client.get("/graph/nodes/does-not-exist").status_code == 404 + assert client.get("/graph/nodes/does-not-exist", headers=ANALYST_HEADERS).status_code == 404 # --- traversal --------------------------------------------------------------- @@ -114,7 +130,8 @@ def test_missing_node_returns_404(): def test_graph_query_traverses_concept_hierarchy(): resp = client.post( "/graph/query", - json={"start_id": "고객", "direction": "both", "max_depth": 1, "actor": "analyst"}, + json={"start_id": "고객", "direction": "both", "max_depth": 1}, + headers=ANALYST_HEADERS, ) assert resp.status_code == 200 body = resp.json() @@ -129,7 +146,8 @@ def test_graph_query_traverses_concept_hierarchy(): def test_graph_query_edge_type_filter(): resp = client.post( "/graph/query", - json={"start_id": "고객", "edge_types": ["narrower"], "direction": "out", "max_depth": 1, "actor": "analyst"}, + json={"start_id": "고객", "edge_types": ["narrower"], "direction": "out", "max_depth": 1}, + headers=ANALYST_HEADERS, ) assert resp.status_code == 200 for edge in resp.json()["edges"]: @@ -137,7 +155,11 @@ def test_graph_query_edge_type_filter(): def test_graph_query_unknown_start_is_404(): - resp = client.post("/graph/query", json={"start_id": "nope", "max_depth": 1, "actor": "analyst"}) + resp = client.post( + "/graph/query", + json={"start_id": "nope", "max_depth": 1}, + headers=ANALYST_HEADERS, + ) assert resp.status_code == 404 @@ -146,7 +168,9 @@ def test_graph_query_unknown_start_is_404(): def test_semantic_search_ranks_churn_concept_first(): resp = client.post( - "/search/semantic", json={"query": "churn 이탈한 고객", "kind": "concept", "limit": 3, "actor": "analyst"} + "/search/semantic", + json={"query": "churn 이탈한 고객", "kind": "concept", "limit": 3}, + headers=ANALYST_HEADERS, ) assert resp.status_code == 200 results = resp.json()["results"] @@ -159,7 +183,9 @@ def test_semantic_search_ranks_churn_concept_first(): def test_semantic_search_kind_filter_only_returns_datasets(): resp = client.post( - "/search/semantic", json={"query": "고객 프로필 데이터", "kind": "dataset", "limit": 5, "actor": "analyst"} + "/search/semantic", + json={"query": "고객 프로필 데이터", "kind": "dataset", "limit": 5}, + headers=ANALYST_HEADERS, ) assert resp.status_code == 200 results = resp.json()["results"] diff --git a/tests/test_graph_security.py b/tests/test_graph_security.py index 9ce2917..eb2eb10 100644 --- a/tests/test_graph_security.py +++ b/tests/test_graph_security.py @@ -20,11 +20,15 @@ from __future__ import annotations import json +from types import SimpleNamespace import pytest from fastapi.testclient import TestClient +from sdp_core import ActorContext +from sdp import api as api_module from sdp import config as config_module +from sdp import document_semantics from sdp import graph_store as gs from sdp.api import app from sdp.config import AppConfig, override_app_config @@ -36,6 +40,18 @@ ) client = TestClient(app) +ADMIN_HEADERS = {"Authorization": "Bearer admin-token"} +ANALYST_HEADERS = {"Authorization": "Bearer analyst-token"} + + +@pytest.fixture(autouse=True) +def fake_graph_oidc(monkeypatch): + def verify(token): + role = token.removesuffix("-token") + roles = ["admin", "data-analyst"] if role == "admin" else ["data-analyst"] + return ActorContext(subject=role, tenant_id="demo", roles=roles), {} + + monkeypatch.setattr(api_module.authz, "verify_oidc_jwks_token", verify) # The exact stacked-SQL breakout payload proven in the adversarial review: it # tries to close the ``$$`` dollar-quote body, terminate the cypher(), and run a @@ -262,7 +278,8 @@ def test_in_memory_traverse_rejects_bad_edge_type(): def test_graph_edge_endpoint_rejects_bad_label(): resp = client.post( "/graph/edges", - json={"edge_type": "bad-type", "source_id": "svc-A", "target_id": "고객", "actor": "admin"}, + json={"edge_type": "bad-type", "source_id": "svc-A", "target_id": "고객"}, + headers=ADMIN_HEADERS, ) assert resp.status_code == 400 @@ -276,20 +293,17 @@ def test_traversal_request_has_no_raw_cypher_field(): assert "cypher" not in GraphTraversalRequest.model_fields -def test_raw_cypher_in_body_is_ignored_not_executed(): - # An attacker-supplied ``cypher`` key is silently dropped (extra field), and - # the safe parameterized traversal runs instead. +def test_raw_cypher_in_body_is_rejected(): resp = client.post( "/graph/query", json={ "start_id": "고객", "max_depth": 1, - "actor": "analyst", "cypher": "MATCH (n) DETACH DELETE n RETURN n", }, + headers=ANALYST_HEADERS, ) - assert resp.status_code == 200 - assert resp.json()["nodes"] + assert resp.status_code == 422 # --- authz on graph endpoints ------------------------------------------------- @@ -300,13 +314,14 @@ def test_graph_node_write_refused_for_anonymous(): "/graph/nodes", json={"node_id": "unauth-node", "kind": "service"}, ) - assert resp.status_code == 403 + assert resp.status_code == 401 def test_graph_node_write_refused_for_non_admin_reader(): resp = client.post( "/graph/nodes", - json={"node_id": "unauth-node", "kind": "service", "actor": "analyst"}, + json={"node_id": "unauth-node", "kind": "service"}, + headers=ANALYST_HEADERS, ) assert resp.status_code == 403 @@ -316,28 +331,29 @@ def test_graph_edge_write_refused_for_anonymous(): "/graph/edges", json={"edge_type": "related", "source_id": "a", "target_id": "b"}, ) - assert resp.status_code == 403 + assert resp.status_code == 401 def test_concept_write_refused_for_anonymous(): resp = client.post("/ontology/concepts", json={"concept": "무단개념"}) - assert resp.status_code == 403 + assert resp.status_code == 401 def test_graph_query_refused_for_anonymous(): resp = client.post("/graph/query", json={"start_id": "고객", "max_depth": 1}) - assert resp.status_code == 403 + assert resp.status_code == 401 def test_semantic_search_refused_for_anonymous(): resp = client.post("/search/semantic", json={"query": "고객"}) - assert resp.status_code == 403 + assert resp.status_code == 401 def test_graph_write_allowed_for_admin(): resp = client.post( "/graph/nodes", - json={"node_id": "authz-node", "kind": "service", "actor": "admin"}, + json={"node_id": "authz-node", "kind": "service"}, + headers=ADMIN_HEADERS, ) assert resp.status_code == 200 @@ -358,6 +374,42 @@ def test_build_store_auto_without_dsn_uses_memory(monkeypatch): assert isinstance(build_store(), InMemoryGraphStore) +def test_build_store_injects_configured_orchestrator_embedder(monkeypatch): + embedded = [] + config = override_app_config( + graph_backend="memory", + orchestrator_base_url="https://orchestrator.example", + embedding_dimension=2, + ) + monkeypatch.setattr(gs, "get_app_config", lambda: config) + monkeypatch.setattr( + document_semantics, + "build_orchestrator_client", + lambda configured: SimpleNamespace( + embed_one=lambda text: embedded.append(text) or [1.0, 0.0] + ), + ) + + store = build_store() + store.upsert_node("voc", "file_asset", text="효성중공업 VOC") + + assert store.dimension == 2 + assert embedded == ["효성중공업 VOC"] + + +def test_configured_orchestrator_fails_closed_without_runtime_credential(monkeypatch): + config = override_app_config( + graph_backend="memory", + orchestrator_base_url="https://orchestrator.example", + ) + document_semantics.set_credential_registry( + document_semantics.EphemeralCredentialRegistry({}) + ) + + with pytest.raises(RuntimeError, match="CONTEXTUAL_ORCHESTRATOR_TOKEN"): + document_semantics.validate_runtime_credentials(config) + + def test_build_store_postgres_without_dsn_fails_loud(monkeypatch): monkeypatch.setattr(gs, "get_app_config", lambda: override_app_config(graph_backend="postgres")) monkeypatch.setattr( From 45605e46b66bd2d42df8aa5b24ed8dcbd226684f Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 14:20:32 +0900 Subject: [PATCH 10/15] fix: harden live semantic extraction --- docs/implementation-compliance.md | 1 + src/sdp/document_semantics.py | 67 +++++++++++++++++++------------ tests/test_file_knowledge.py | 26 ++++++++++++ 3 files changed, 68 insertions(+), 26 deletions(-) diff --git a/docs/implementation-compliance.md b/docs/implementation-compliance.md index 4425f61..edd2556 100644 --- a/docs/implementation-compliance.md +++ b/docs/implementation-compliance.md @@ -146,6 +146,7 @@ - 운영 패키징: SHACL/OWL TTL은 `sdp/resources/*.ttl` package data로 wheel/container에 포함되고, `orchestrator_base_url` 구성 시 app lifespan은 runtime `CredentialRegistry` token이 없으면 fail-closed한다. - 벡터 일관성: KV `embedding_dimension`을 orchestrator `dimensions`에 전달하고 같은 orchestrator embedder를 memory/Postgres graph store의 ingest와 검색에 주입한다. - 파일럿: `src/sdp/file_pilot.py`가 content deduplication, 안전한 문서 추출, graph projection, 로컬 전용 manifest를 수행한다. +- 2026-07-21 live 파일럿: `contextual-orchestrator`를 거친 OpenAI structured output으로 12개 파일을 10개 content-addressed 자산과 12개 distribution, 근거·신뢰도 포함 proposed assertion 286건으로 변환했다. 12개 모두 `extracted`, manifest의 10개 자산 모두 pySHACL conform이며 원문 조각·근거 인용문·credential marker는 0건이다. 민감한 파일명과 locator가 포함된 결과는 Git 제외 로컬 `outputs/hyosung-voc-file-index-live.json`에만 보관한다. - 증빙 테스트: - `tests/test_file_knowledge.py::test_orchestrator_extractor_uses_strict_schema_and_persists_only_evidence_reference` - `tests/test_file_knowledge.py::test_orchestrator_client_uses_sync_embeddings_endpoint` diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py index 78e7d53..6283dd4 100644 --- a/src/sdp/document_semantics.py +++ b/src/sdp/document_semantics.py @@ -320,6 +320,7 @@ def _payload(self, filename: str, chunk: TextChunk) -> dict[str, Any]: return { "model": self.semantic_model, "store": False, + "reasoning_effort": "none", "messages": [ { "role": "system", @@ -398,31 +399,29 @@ def embed(self, texts: list[str]) -> list[list[float]]: def embed_one(self, text: str) -> list[float]: return self.embed([text])[0] - def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: - merged: dict[tuple[str, str, str], SemanticAssertion] = {} - for chunk in chunks: - response = self._post( - "/v1/chat/completions", self._payload(filename, chunk) - ) + def _extract_chunk(self, filename: str, chunk: TextChunk) -> list[SemanticAssertion]: + response = self._post("/v1/chat/completions", self._payload(filename, chunk)) + try: + result = json.loads(_chat_completion_content(response)) + candidates = result["assertions"] + except (KeyError, TypeError, json.JSONDecodeError) as exc: + raise ValueError("orchestrator semantic output is malformed") from exc + if not isinstance(candidates, list): + raise ValueError("orchestrator semantic output assertions must be a list") + + assertions: list[SemanticAssertion] = [] + for candidate in candidates: + if not isinstance(candidate, dict): + raise ValueError("orchestrator semantic assertion must be an object") + quote_text = candidate.get("evidence_quote") + if not isinstance(quote_text, str): + raise ValueError("orchestrator semantic assertion has no evidence quote") + local_start = chunk.text.find(quote_text) + if local_start < 0: + continue try: - result = json.loads(_chat_completion_content(response)) - candidates = result["assertions"] - except (KeyError, TypeError, json.JSONDecodeError) as exc: - raise ValueError("orchestrator semantic output is malformed") from exc - if not isinstance(candidates, list): - raise ValueError("orchestrator semantic output assertions must be a list") - - for candidate in candidates: - if not isinstance(candidate, dict): - raise ValueError("orchestrator semantic assertion must be an object") - quote_text = candidate.get("evidence_quote") - if not isinstance(quote_text, str): - raise ValueError("orchestrator semantic assertion has no evidence quote") - local_start = chunk.text.find(quote_text) - if local_start < 0: - continue - try: - assertion = SemanticAssertion( + assertions.append( + SemanticAssertion( relation=candidate["relation"], target_kind=candidate["target_kind"], target_label=candidate["target_label"], @@ -433,8 +432,24 @@ def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssert evidence_end=chunk.start + local_start + len(quote_text), method="contextual-orchestrator", ) - except (KeyError, TypeError, ValueError) as exc: - raise ValueError("orchestrator semantic assertion is invalid") from exc + ) + except (KeyError, TypeError, ValueError) as exc: + raise ValueError("orchestrator semantic assertion is invalid") from exc + return assertions + + def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: + merged: dict[tuple[str, str, str], SemanticAssertion] = {} + for chunk in chunks: + # ponytail: one retry handles nondeterministic structured-output truncation; + # add persistent retry telemetry only if provider instability warrants it. + for attempt in range(2): + try: + assertions = self._extract_chunk(filename, chunk) + break + except ValueError: + if attempt: + raise + for assertion in assertions: key = ( assertion.relation, assertion.target_kind, diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index 564cac6..d30eaa1 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -404,6 +404,7 @@ def test_orchestrator_extractor_uses_strict_schema_and_persists_only_evidence_re assert captured["store"] is False assert captured["model"] == "gpt-5-mini-2025-08-07" + assert captured["reasoning_effort"] == "none" assert captured["response_format"]["type"] == "json_schema" assert captured["response_format"]["json_schema"]["strict"] is True assert captured["url"] == "https://orchestrator.example/v1/chat/completions" @@ -428,6 +429,31 @@ def test_orchestrator_extractor_rejects_quote_not_present_in_input(): assert extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) == [] +def test_orchestrator_extractor_retries_one_malformed_structured_output(): + calls = 0 + good_transport = fake_orchestrator_transport({}) + + def flaky_transport(request, timeout): + nonlocal calls + calls += 1 + if calls == 1: + return {"choices": [{"message": {"content": "{"}}]} + return good_transport(request, timeout) + + extractor = ContextualOrchestratorClient( + EphemeralCredentialRegistry( + {"CONTEXTUAL_ORCHESTRATOR_TOKEN": "orchestrator-token"} + ), + base_url="https://orchestrator.example", + transport=flaky_transport, + ) + + assertions = extractor.extract("meeting.docx", chunk_text("효성중공업 C-Cube PoC")) + + assert calls == 2 + assert len(assertions) == 1 + + def test_orchestrator_extractor_fails_closed_without_credential(): extractor = ContextualOrchestratorClient( EphemeralCredentialRegistry({}), From e0a212e6fc877ac77ff0e44fbebf0e88957a53c4 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 16:44:53 +0900 Subject: [PATCH 11/15] fix: document audited Semgrep findings --- src/sdp/document_semantics.py | 3 ++- src/sdp/file_ontology.py | 3 ++- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py index 6283dd4..5879546 100644 --- a/src/sdp/document_semantics.py +++ b/src/sdp/document_semantics.py @@ -269,7 +269,8 @@ def chunk_text( def _orchestrator_http_transport(request: Request, timeout: int) -> dict[str, Any]: try: - with urlopen(request, timeout=timeout) as response: + # ContextualOrchestratorClient accepts only absolute HTTP(S) base URLs. + with urlopen(request, timeout=timeout) as response: # nosemgrep: python.lang.security.audit.dynamic-urllib-use-detected.dynamic-urllib-use-detected return json.loads(response.read().decode("utf-8")) except HTTPError as exc: raise RuntimeError(f"orchestrator request failed with HTTP {exc.code}") from None diff --git a/src/sdp/file_ontology.py b/src/sdp/file_ontology.py index 78f1a2b..eb7daa6 100644 --- a/src/sdp/file_ontology.py +++ b/src/sdp/file_ontology.py @@ -6,7 +6,8 @@ import json import re from datetime import datetime -from importlib.resources import files +# Python >=3.10 is required by pyproject.toml, so the stdlib module is valid. +from importlib.resources import files # nosemgrep: python.lang.compatibility.python37.python37-compatibility-importlib2 from typing import Any, Literal from urllib.parse import urlsplit From 324567e4437596c7fcde77ffa5f20adc2d86b3a8 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 16:53:20 +0900 Subject: [PATCH 12/15] fix: clear repository-wide Semgrep findings --- src/sdp/authz.py | 7 ++++++- src/sdp/graph_store.py | 6 ++++-- src/sdp/observability.py | 3 ++- tests/test_api.py | 6 ++++++ 4 files changed, 18 insertions(+), 4 deletions(-) diff --git a/src/sdp/authz.py b/src/sdp/authz.py index 98e83bf..d77c9c9 100644 --- a/src/sdp/authz.py +++ b/src/sdp/authz.py @@ -4,6 +4,7 @@ import os from datetime import datetime, timezone from typing import Any +from urllib.parse import urlsplit from urllib.request import urlopen import jwt @@ -112,8 +113,12 @@ def resolve_oidc_actor_context( def _load_jwks_from_url(jwks_url: str) -> dict[str, Any]: + parsed = urlsplit(jwks_url) + if parsed.scheme != "https" or not parsed.netloc: + raise ValueError("OIDC JWKS URL must be absolute HTTPS") timeout = float(os.getenv("SDP_OIDC_JWKS_TIMEOUT_SECONDS", "2")) - with urlopen(jwks_url, timeout=timeout) as response: + # The HTTPS-only check above excludes urllib's local file handlers. + with urlopen(jwks_url, timeout=timeout) as response: # nosemgrep: python.lang.security.audit.dynamic-urllib-use-detected.dynamic-urllib-use-detected return json.loads(response.read().decode("utf-8")) diff --git a/src/sdp/graph_store.py b/src/sdp/graph_store.py index c4b2075..06380d5 100644 --- a/src/sdp/graph_store.py +++ b/src/sdp/graph_store.py @@ -542,7 +542,8 @@ def _cypher( ) driver_connection = conn.connection.driver_connection with driver_connection.cursor() as cursor: - cursor.execute(statement, (json.dumps(params),)) + # pg_sql quotes both literals; user values stay in the bound parameter map. + cursor.execute(statement, (json.dumps(params),)) # nosemgrep: python.sqlalchemy.security.sqlalchemy-execute-raw-query.sqlalchemy-execute-raw-query return cursor.fetchall() def upsert_node( @@ -852,7 +853,8 @@ def semantic_search( stmt += "WHERE " + " AND ".join(conditions) + " " stmt += "ORDER BY e.embedding <=> CAST(:vec AS vector) LIMIT :limit" with self._engine.connect() as conn: - rows = conn.execute(sql(stmt), params).fetchall() + # stmt contains only closed fragments; every dynamic value is bound in params. + rows = conn.execute(sql(stmt), params).fetchall() # nosemgrep: python.sqlalchemy.security.audit.avoid-sqlalchemy-text.avoid-sqlalchemy-text return [ { "node_id": row[0], diff --git a/src/sdp/observability.py b/src/sdp/observability.py index cffe813..eca1c07 100644 --- a/src/sdp/observability.py +++ b/src/sdp/observability.py @@ -136,7 +136,8 @@ def _export_to_sink(observation: dict[str, Any]) -> None: headers={"Content-Type": "application/json"}, method="POST", ) - with urlopen(request, timeout=timeout_ms / 1000): + # This branch excludes urllib's local file handlers before the request is built. + with urlopen(request, timeout=timeout_ms / 1000): # nosemgrep: python.lang.security.audit.dynamic-urllib-use-detected.dynamic-urllib-use-detected return raise ValueError(f"unsupported SDP_LOG_SINK_URL scheme: {scheme}") diff --git a/tests/test_api.py b/tests/test_api.py index 4d89779..a859c71 100644 --- a/tests/test_api.py +++ b/tests/test_api.py @@ -10,6 +10,7 @@ from jwt.algorithms import RSAAlgorithm import sdp.catalog as app_catalog +import sdp.authz as app_authz import sdp.domain as app_domain import sdp.evidence as app_evidence import sdp.observability as app_observability @@ -444,6 +445,11 @@ def test_oidc_jwks_verification_maps_verified_token_without_token_leak(): assert token not in json.dumps(body) +def test_oidc_jwks_loader_rejects_non_https_url(): + with pytest.raises(ValueError, match="HTTPS"): + app_authz._load_jwks_from_url("file:///etc/passwd") + + def test_oidc_jwks_verification_rejects_wrong_audience(): private_key = rsa.generate_private_key(public_exponent=65537, key_size=2048) jwk = json.loads(RSAAlgorithm.to_jwk(private_key.public_key())) From 76a1c5e8a1426844a7fed4955e3068736e853072 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 20:01:33 +0900 Subject: [PATCH 13/15] fix: bound in-memory audit evidence --- src/sdp/catalog.py | 3 ++- tests/fuzz/test_fuzz_properties.py | 24 ++++++++++++++++++++++++ tests/test_api.py | 3 ++- 3 files changed, 28 insertions(+), 2 deletions(-) diff --git a/src/sdp/catalog.py b/src/sdp/catalog.py index 38e1a70..f4a101d 100644 --- a/src/sdp/catalog.py +++ b/src/sdp/catalog.py @@ -1,6 +1,7 @@ from __future__ import annotations import re +from collections import deque from dataclasses import dataclass from datetime import datetime, timezone from typing import Any, Dict, List, Optional @@ -25,7 +26,7 @@ def _seed_datasets() -> List[Dataset]: _DATA = {dataset.id: dataset for dataset in _seed_datasets()} for _dataset in _DATA.values(): _dataset.recompute_scores() -_AUDIT_LOG: list[AuditEvent] = [] +_AUDIT_LOG: deque[AuditEvent] = deque(maxlen=1000) _SCHEMA_HISTORY: dict[str, list[dict[str, Any]]] = {} diff --git a/tests/fuzz/test_fuzz_properties.py b/tests/fuzz/test_fuzz_properties.py index 9ac5a04..22954a5 100644 --- a/tests/fuzz/test_fuzz_properties.py +++ b/tests/fuzz/test_fuzz_properties.py @@ -20,6 +20,8 @@ from hypothesis import HealthCheck, given, settings from hypothesis import strategies as st +import sdp.catalog as app_catalog +import sdp.evidence as app_evidence from sdp.domain import ( DatasetCreateRequest, QueryDraftRequest, @@ -160,6 +162,28 @@ def test_execute_query_holds_invariants(req): invariants.check_execute_query(req) +def test_execute_query_keeps_fallback_audit_log_bounded(): + previous_store = app_evidence.configure_evidence_store(None) + previous_events = list(app_catalog._AUDIT_LOG) + app_catalog._AUDIT_LOG.clear() + req = QueryExecutionRequest( + language="python", + user="analyst", + purpose="analysis", + dataset_ids=["crm-customer-master"], + query="x", + dry_run=True, + ) + try: + for _ in range(1001): + invariants.check_execute_query(req) + assert len(app_catalog._AUDIT_LOG) <= 1000 + finally: + app_catalog._AUDIT_LOG.clear() + app_catalog._AUDIT_LOG.extend(previous_events) + app_evidence.configure_evidence_store(previous_store) + + # --------------------------------------------------------------------------- # # DTO parse boundary: arbitrary dicts must yield either a valid model or a # Pydantic ValidationError — never an unexpected exception type. diff --git a/tests/test_api.py b/tests/test_api.py index a859c71..51e5daa 100644 --- a/tests/test_api.py +++ b/tests/test_api.py @@ -35,7 +35,8 @@ def isolate_in_memory_app_state(): yield app_catalog._DATA.clear() app_catalog._DATA.update(data) - app_catalog._AUDIT_LOG[:] = audit_log + app_catalog._AUDIT_LOG.clear() + app_catalog._AUDIT_LOG.extend(audit_log) app_catalog._SCHEMA_HISTORY.clear() app_catalog._SCHEMA_HISTORY.update(schema_history) app_evidence._POLICY_DECISION_LOG[:] = policy_log From 247489f6e4f7d5ae285c6c3a7d6f8ec9b6244ba3 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 21 Jul 2026 20:12:26 +0900 Subject: [PATCH 14/15] fix: make protocol stubs explicit --- src/sdp/document_semantics.py | 3 ++- src/sdp/file_pilot.py | 6 ++++-- src/sdp/storage_readers.py | 6 ++++-- 3 files changed, 10 insertions(+), 5 deletions(-) diff --git a/src/sdp/document_semantics.py b/src/sdp/document_semantics.py index 5879546..657da40 100644 --- a/src/sdp/document_semantics.py +++ b/src/sdp/document_semantics.py @@ -60,7 +60,8 @@ class TextChunk: class CredentialRegistry(Protocol): - def get_credential(self, name: str) -> str | None: ... + def get_credential(self, name: str) -> str | None: + pass class EphemeralCredentialRegistry: diff --git a/src/sdp/file_pilot.py b/src/sdp/file_pilot.py index c6dc65f..4d04a8b 100644 --- a/src/sdp/file_pilot.py +++ b/src/sdp/file_pilot.py @@ -23,9 +23,11 @@ class PilotExtractor(Protocol): - def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: ... + def extract(self, filename: str, chunks: list[TextChunk]) -> list[SemanticAssertion]: + pass - def embed_one(self, text: str) -> list[float]: ... + def embed_one(self, text: str) -> list[float]: + pass def run_local_pilot( diff --git a/src/sdp/storage_readers.py b/src/sdp/storage_readers.py index d20898c..bf3ddc9 100644 --- a/src/sdp/storage_readers.py +++ b/src/sdp/storage_readers.py @@ -36,9 +36,11 @@ class ObjectRef: class ObjectReader(Protocol): def list( self, prefix: str = "", *, name_pattern: str | None = None - ) -> Iterable[ObjectRef]: ... + ) -> Iterable[ObjectRef]: + pass - def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: ... + def read(self, ref: ObjectRef, *, max_bytes: int) -> bytes: + pass class FilesystemReader: From 19028a981f84a135243221ca9708650273b0c87e Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 4 Aug 2026 13:08:15 +0900 Subject: [PATCH 15/15] feat: preview DiskSage pre-copy catalog candidates (#31) * feat: preview DiskSage catalog candidates safely * feat: model DiskSage post-copy lineage --- README.md | 20 ++ docs/disksage-copy-lineage-erd.md | 65 +++++ docs/implementation-compliance.md | 5 + migrations/0002_file_copy_lineage.sql | 164 +++++++++++++ src/sdp/api.py | 89 +++++++ src/sdp/disksage_catalog.py | 340 ++++++++++++++++++++++++++ tests/test_file_knowledge.py | 249 +++++++++++++++++++ tests/test_migrations.py | 31 +++ 8 files changed, 963 insertions(+) create mode 100644 docs/disksage-copy-lineage-erd.md create mode 100644 migrations/0002_file_copy_lineage.sql create mode 100644 src/sdp/disksage_catalog.py diff --git a/README.md b/README.md index fbccf1e..7496f28 100644 --- a/README.md +++ b/README.md @@ -18,6 +18,9 @@ `ontology_concepts`, `concept_edges`(→ `graph_edges`), `dataset_nodes`, `graph_nodes`, `embedding_vectors`, `config_entries`, `schema_migrations`. +DiskSage의 검증된 복사 이후 lineage를 위한 정규화 ERD와 pg-erd 재현 절차는 +[`docs/disksage-copy-lineage-erd.md`](docs/disksage-copy-lineage-erd.md)에 있습니다. + ## 목표 - 데이터 탐색: 키워드 카탈로그 검색 + 유사어/용어(ontology) 해석 + **그래프 순회** + **시맨틱 검색** @@ -114,6 +117,7 @@ SDP_DATABASE_DSN='postgresql+psycopg://sdp_graph_app:@loca ingest/validate 때 pySHACL로 실제 실행됩니다. - `POST /file-assets` — 관리자 정책을 통과한 자산·후보 주장 적재 +- `POST /file-assets/preview/disksage` — DiskSage 복사 전 후보의 비영속·경로 비노출 온톨로지 미리보기 - `GET /file-assets/{asset_id}` — 자산과 의미 관계 조회 - `GET /file-assets/{asset_id}/jsonld` — 기본 locator 비공개 JSON-LD - `GET /file-assets/{asset_id}/validate` — pySHACL 검증 리포트 @@ -139,6 +143,22 @@ context를 요구하고 body `actor`를 거부합니다. 일반 graph API는 gov node/edge를 수정할 수 없으며, traversal/search 결과는 tenant로 필터링되고 locator는 항상 redaction됩니다. +`POST /file-assets/preview/disksage`는 `disksage.file-catalog-candidate-batch` v1만 +받습니다. 본문은 2 MiB, 후보는 200건으로 제한되며 `src`, `dst`, `filename`, +`relative_path`, account/object id 같은 저장 위치 식별자는 계약에 존재하지 않고 +알 수 없는 필드는 거부됩니다. 생산일은 +`embedded_metadata → explicit_filename_date → filesystem_created → filesystem_modified` +순서가 고정되어 있고, 선택된 값은 같은 날짜·source의 metadata evidence에 결합되어야 +합니다. 파일명 날짜는 embedded metadata가 없을 때만 낮은 신뢰도의 보조값으로 +허용됩니다. + +이 endpoint는 deterministic archive-kind→artifact-type 제안만 만들며 LLM, graph +store, file-asset ingest를 호출하지 않습니다. 응답은 title/author/context/evidence +값을 되돌려주지 않고 `content_sha256`와 verified distribution이 없으므로 +`persistable_as_file_asset=false`를 명시합니다. 또한 copy/eviction 허가가 아니며, +기존 create-file 정책을 통과한 관리자만 사용할 수 있습니다. 중앙 policy decision +증빙은 기록되지만 catalog/file asset 자체는 저장되지 않습니다. + GitHub Secret에 값을 저장하는 것만으로는 런타임 주입이 되지 않습니다. 배포 호스트는 secret manager에서 token을 읽는 `CredentialRegistry` 구현을 만든 뒤 `sdp.api.create_app(registry)`로 ASGI 앱을 구성해야 합니다. KV에 diff --git a/docs/disksage-copy-lineage-erd.md b/docs/disksage-copy-lineage-erd.md new file mode 100644 index 0000000..e3cda49 --- /dev/null +++ b/docs/disksage-copy-lineage-erd.md @@ -0,0 +1,65 @@ +# DiskSage post-copy lineage ERD + +This relational read-model records only integrity-verified copy lineage after +the corresponding file asset and distribution graph nodes exist. It does not +authorize a copy, provider write, sync, or local eviction. Source paths are not +stored; the ingress contract supplies only `source_locator_sha256`. + +```mermaid +erDiagram + graph_nodes ||--o| file_asset_records : "content-addressed projection" + graph_nodes ||--o| file_distribution_records : "distribution projection" + file_asset_records ||--o{ file_distribution_records : "has location" + file_asset_records ||--o{ cloud_copy_receipts : "has verified copy" + file_distribution_records ||--o{ cloud_copy_receipts : "is destination of" + cloud_copy_receipts ||--o{ file_metadata_evidence_records : "selects production-time evidence" + cloud_copy_receipts ||--o{ cloud_sync_evidence_records : "has provider observations" +``` + +The schema deliberately leaves the generic `graph_edges` mirror unchanged. +Its seed path can create relationships before both endpoint nodes exist, while +post-copy lineage requires pre-existing asset and distribution nodes. All seven +new foreign keys use `ON DELETE RESTRICT` so catalog or graph cleanup cannot +silently erase receipt evidence. + +## Reproducible pg-erd snapshot + +After applying `migrations/0001_init_graph_vector.sql` and +`migrations/0002_file_copy_lineage.sql` to a temporary PostgreSQL 17.10 database, +capture the catalog through the Unix-socket-only pg-erd CLI: + +```bash +pg-erd-snapshot \ + --host /tmp \ + --database semantic_data_portal_erd \ + --schema public \ + --pretty > semantic-data-portal.snapshot.json +``` + +The validated snapshot contained 12 relations, 100 columns, 44 constraints, +24 indexes, 16 primary-key columns, and 7 foreign-key edges. The pre-migration +snapshot contained 7 relations and no foreign-key edges. + +The seven observed foreign-key edges were: + +| Child | Column | Parent | Column | +| --- | --- | --- | --- | +| `file_asset_records` | `graph_node_id` | `graph_nodes` | `node_id` | +| `file_distribution_records` | `graph_node_id` | `graph_nodes` | `node_id` | +| `file_distribution_records` | `asset_node_id` | `file_asset_records` | `graph_node_id` | +| `cloud_copy_receipts` | `asset_node_id` | `file_asset_records` | `graph_node_id` | +| `cloud_copy_receipts` | `destination_distribution_node_id` | `file_distribution_records` | `graph_node_id` | +| `file_metadata_evidence_records` | `receipt_id` | `cloud_copy_receipts` | `receipt_id` | +| `cloud_sync_evidence_records` | `receipt_id` | `cloud_copy_receipts` | `receipt_id` | + +## Persistence boundary + +- `local_copy_verified` must be true for every persisted receipt. +- `provider_sync_confirmed` remains distinct from local copy verification and + may only be updated from a bound provider evidence record. +- Human-review fields are all required for review-gated copies and must record + an approved disposition; they must all be absent for non-review candidates. +- Production time remains bound to ordered metadata evidence. The portal does + not reinterpret filename dates as embedded production metadata. +- Provider authority, quota, and tenant policy are upstream copy gates. A row in + this read-model is evidence, not permission to copy or evict. diff --git a/docs/implementation-compliance.md b/docs/implementation-compliance.md index edd2556..a2c5c47 100644 --- a/docs/implementation-compliance.md +++ b/docs/implementation-compliance.md @@ -142,6 +142,7 @@ - LLM 경계: `src/sdp/document_semantics.py::ContextualOrchestratorClient`가 `/v1/chat/completions`와 `/v1/embeddings`만 호출한다. 포털에는 OpenAI/provider key가 없다. - provider 중립성: Synology는 filesystem 배치일 뿐 필수 구성요소가 아니며, 파일 정체성은 저장소 URL과 독립적이다. - 정책/API: `POST /file-assets`, `GET /file-assets/{asset_id}`, `/jsonld`, `/validate`는 검증된 OIDC Bearer actor context와 저장된 `FileAsset.tenant_id`를 중앙 policy decision에 전달한다. body/query `actor`는 권한 근거가 아니며 locator 공개는 동일 tenant의 admin 또는 platform-admin이 필요하다. +- DiskSage pre-copy adapter: `POST /file-assets/preview/disksage`는 2 MiB/200건으로 제한된 strict v1 batch만 받아 원본 path/file name/account/object id 없이 deterministic `hasArtifactType` 제안을 반환한다. 선택 생산일은 `embedded metadata → explicit filename date → filesystem creation → modification`과 matching evidence를 강제한다. content metadata는 응답에 echo하지 않고 graph/file asset을 저장하거나 LLM을 호출하지 않으며 copy/eviction 허가도 만들지 않는다. create-file policy decision 증빙만 기록한다. - 그래프 격리: 같은 SHA 및 파일 관계 대상의 cross-tenant merge를 거부한다. generic graph mutation은 file/distribution node를 다룰 수 없고 graph traversal/semantic search는 OIDC tenant policy로 file node를 필터링하며 locator를 redaction한다. - 운영 패키징: SHACL/OWL TTL은 `sdp/resources/*.ttl` package data로 wheel/container에 포함되고, `orchestrator_base_url` 구성 시 app lifespan은 runtime `CredentialRegistry` token이 없으면 fail-closed한다. - 벡터 일관성: KV `embedding_dimension`을 orchestrator `dimensions`에 전달하고 같은 orchestrator embedder를 memory/Postgres graph store의 ingest와 검색에 주입한다. @@ -152,6 +153,10 @@ - `tests/test_file_knowledge.py::test_orchestrator_client_uses_sync_embeddings_endpoint` - `tests/test_file_knowledge.py::test_local_pilot_deduplicates_content_and_writes_no_raw_text` - `tests/test_file_knowledge.py::test_file_asset_api_requires_policy_and_redacts_jsonld_locator` + - `tests/test_file_knowledge.py::test_disksage_pre_copy_contract_enforces_metadata_precedence_and_evidence_binding` + - `tests/test_file_knowledge.py::test_disksage_pre_copy_contract_rejects_paths_and_unbounded_shape` + - `tests/test_file_knowledge.py::test_disksage_pre_copy_api_requires_create_policy_and_never_uses_graph_store` + - `tests/test_file_knowledge.py::test_disksage_pre_copy_api_redacts_validation_input_and_limits_body` - `tests/test_graph_engine.py::test_config_loads_from_kv_mapping` ## 8) 다음 단계 (현재 브랜치에서 미반영 권고) diff --git a/migrations/0002_file_copy_lineage.sql b/migrations/0002_file_copy_lineage.sql new file mode 100644 index 0000000..8afc76a --- /dev/null +++ b/migrations/0002_file_copy_lineage.sql @@ -0,0 +1,164 @@ +-- DiskSage post-copy catalog lineage. +-- +-- These tables are a normalized read-model over file_asset/distribution graph +-- nodes. They intentionally do not add foreign keys to graph_edges: graph seed +-- data may create an edge before both relational mirror nodes exist. Post-copy +-- lineage is stricter and can only reference already-persisted graph nodes. + +CREATE TABLE IF NOT EXISTS file_asset_records ( + graph_node_id TEXT PRIMARY KEY + REFERENCES graph_nodes (node_id) ON DELETE RESTRICT, + tenant_id TEXT NOT NULL, + content_sha256 CHAR(64) NOT NULL UNIQUE, + byte_size BIGINT NOT NULL CHECK (byte_size >= 0), + media_type TEXT NOT NULL, + asset_title TEXT NOT NULL, + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + updated_at TIMESTAMPTZ NOT NULL DEFAULT now(), + CONSTRAINT file_asset_records_digest_check CHECK ( + content_sha256 ~ '^[0-9a-f]{64}$' + AND graph_node_id = 'urn:sha256:' || content_sha256 + ) +); +CREATE INDEX IF NOT EXISTS file_asset_records_tenant_idx + ON file_asset_records (tenant_id); + +CREATE TABLE IF NOT EXISTS file_distribution_records ( + graph_node_id TEXT PRIMARY KEY + REFERENCES graph_nodes (node_id) ON DELETE RESTRICT, + asset_node_id TEXT NOT NULL + REFERENCES file_asset_records (graph_node_id) ON DELETE RESTRICT, + provider TEXT NOT NULL CHECK ( + provider IN ( + 'filesystem', 's3', 's3_compatible', 'azure_blob', + 'icloud', 'onedrive', 'google-drive' + ) + ), + account_scope TEXT CHECK ( + account_scope IS NULL + OR account_scope IN ('personal', 'organization', 'shared', 'unknown') + ), + endpoint_id TEXT NOT NULL, + available BOOLEAN NOT NULL DEFAULT TRUE, + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + updated_at TIMESTAMPTZ NOT NULL DEFAULT now() +); +CREATE INDEX IF NOT EXISTS file_distribution_records_asset_idx + ON file_distribution_records (asset_node_id); + +CREATE TABLE IF NOT EXISTS cloud_copy_receipts ( + receipt_id CHAR(64) PRIMARY KEY, + asset_node_id TEXT NOT NULL + REFERENCES file_asset_records (graph_node_id) + ON DELETE RESTRICT, + destination_distribution_node_id TEXT NOT NULL + REFERENCES file_distribution_records (graph_node_id) + ON DELETE RESTRICT, + candidate_fingerprint CHAR(64) NOT NULL, + review_fingerprint CHAR(64) NOT NULL, + lineage_fingerprint CHAR(64) NOT NULL, + source_locator_sha256 CHAR(64) NOT NULL, + content_blake3 CHAR(64) NOT NULL, + copied_bytes BIGINT NOT NULL CHECK (copied_bytes >= 0), + source_modified_ms BIGINT NOT NULL CHECK (source_modified_ms >= 0), + copied_at_ms BIGINT NOT NULL CHECK (copied_at_ms > 0), + production_time_ms BIGINT NOT NULL CHECK (production_time_ms > 0), + production_time_source TEXT NOT NULL, + production_time_confidence TEXT NOT NULL CHECK ( + production_time_confidence IN ('high', 'medium', 'low', 'unknown') + ), + filesystem_created_ms BIGINT NOT NULL CHECK (filesystem_created_ms >= 0), + filesystem_modified_ms BIGINT NOT NULL CHECK (filesystem_modified_ms > 0), + copy_verification_method TEXT NOT NULL, + local_copy_verified BOOLEAN NOT NULL CHECK (local_copy_verified), + provider_write_executed BOOLEAN NOT NULL DEFAULT FALSE, + provider_sync_confirmed BOOLEAN NOT NULL DEFAULT FALSE, + requires_review BOOLEAN NOT NULL, + review_reason_codes JSONB NOT NULL DEFAULT '[]'::jsonb CHECK ( + jsonb_typeof(review_reason_codes) = 'array' + ), + review_decision_id TEXT, + review_disposition TEXT CHECK ( + review_disposition IS NULL + OR review_disposition IN ('approved', 'held') + ), + reviewed_at_ms BIGINT CHECK (reviewed_at_ms > 0), + reviewed_by TEXT, + review_rationale TEXT, + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + CONSTRAINT cloud_copy_receipts_digest_check CHECK ( + receipt_id ~ '^[0-9a-f]{64}$' + AND candidate_fingerprint ~ '^[0-9a-f]{64}$' + AND review_fingerprint ~ '^[0-9a-f]{64}$' + AND lineage_fingerprint ~ '^[0-9a-f]{64}$' + AND source_locator_sha256 ~ '^[0-9a-f]{64}$' + AND content_blake3 ~ '^[0-9a-f]{64}$' + ), + CONSTRAINT cloud_copy_receipts_review_check CHECK ( + ( + requires_review + AND jsonb_array_length(review_reason_codes) > 0 + AND review_decision_id IS NOT NULL + AND review_disposition = 'approved' + AND reviewed_at_ms IS NOT NULL + AND reviewed_by IS NOT NULL + AND review_rationale IS NOT NULL + ) + OR ( + NOT requires_review + AND jsonb_array_length(review_reason_codes) = 0 + AND review_decision_id IS NULL + AND review_disposition IS NULL + AND reviewed_at_ms IS NULL + AND reviewed_by IS NULL + AND review_rationale IS NULL + ) + ) +); +CREATE INDEX IF NOT EXISTS cloud_copy_receipts_asset_idx + ON cloud_copy_receipts (asset_node_id); +CREATE INDEX IF NOT EXISTS cloud_copy_receipts_distribution_idx + ON cloud_copy_receipts (destination_distribution_node_id); + +CREATE TABLE IF NOT EXISTS file_metadata_evidence_records ( + receipt_id CHAR(64) NOT NULL + REFERENCES cloud_copy_receipts (receipt_id) ON DELETE RESTRICT, + evidence_order INTEGER NOT NULL CHECK (evidence_order >= 0), + evidence_field TEXT NOT NULL, + evidence_value TEXT NOT NULL, + evidence_source TEXT NOT NULL, + confidence TEXT NOT NULL CHECK (confidence IN ('high', 'medium', 'low', 'unknown')), + selected BOOLEAN NOT NULL DEFAULT FALSE, + PRIMARY KEY (receipt_id, evidence_order) +); +CREATE UNIQUE INDEX IF NOT EXISTS file_metadata_evidence_selected_idx + ON file_metadata_evidence_records (receipt_id) + WHERE selected; + +CREATE TABLE IF NOT EXISTS cloud_sync_evidence_records ( + evidence_record_id TEXT PRIMARY KEY, + receipt_id CHAR(64) NOT NULL + REFERENCES cloud_copy_receipts (receipt_id) ON DELETE RESTRICT, + evidence_kind TEXT NOT NULL, + provider_evidence_id TEXT NOT NULL, + confirmed_at_ms BIGINT NOT NULL CHECK (confirmed_at_ms > 0), + observed_bytes BIGINT NOT NULL CHECK (observed_bytes >= 0), + destination_blake3 CHAR(64) NOT NULL CHECK ( + destination_blake3 ~ '^[0-9a-f]{64}$' + ), + sync_complete BOOLEAN NOT NULL, + remote_object_id TEXT, + remote_revision TEXT, + remote_location_bound BOOLEAN, + sync_reason_codes JSONB NOT NULL DEFAULT '[]'::jsonb CHECK ( + jsonb_typeof(sync_reason_codes) = 'array' + ), + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + UNIQUE (receipt_id, provider_evidence_id) +); +CREATE INDEX IF NOT EXISTS cloud_sync_evidence_records_receipt_idx + ON cloud_sync_evidence_records (receipt_id, confirmed_at_ms DESC); + +INSERT INTO schema_migrations (migration_id) +VALUES ('0002_file_copy_lineage') +ON CONFLICT (migration_id) DO NOTHING; diff --git a/src/sdp/api.py b/src/sdp/api.py index 3a2604d..77ac7eb 100644 --- a/src/sdp/api.py +++ b/src/sdp/api.py @@ -9,6 +9,7 @@ from fastapi import Depends, FastAPI, HTTPException, Query, Request, Response from fastapi.responses import HTMLResponse from fastapi.middleware.cors import CORSMiddleware +from pydantic import ValidationError from sdp_core import ( ActorContext, buyer_demo_activation_plan, @@ -30,6 +31,12 @@ SemanticSearchRequest, ) from .document_semantics import CredentialRegistry, set_credential_registry, validate_runtime_credentials +from .disksage_catalog import ( + MAX_DISKSAGE_CATALOG_BODY_BYTES, + DiskSageCatalogCandidateBatch, + DiskSageCatalogPreviewResponse, + build_disksage_catalog_preview, +) from .graph_store import get_store, set_store from .file_ontology import ( FileAsset, @@ -893,6 +900,88 @@ def ingest_file_asset( } +@app.post( + "/file-assets/preview/disksage", + response_model=DiskSageCatalogPreviewResponse, + openapi_extra={ + "requestBody": { + "required": True, + "content": { + "application/json": { + "schema": DiskSageCatalogCandidateBatch.model_json_schema() + } + }, + } + }, +) +async def preview_disksage_file_catalog( + request: Request, + actor: ActorContext = Depends(_authenticated_actor), +) -> DiskSageCatalogPreviewResponse: + """Validate a path-redacted pre-copy batch without creating file assets.""" + + content_type = ( + request.headers.get("content-type", "").partition(";")[0].strip().lower() + ) + if content_type != "application/json": + raise HTTPException(status_code=415, detail="application/json is required") + content_encoding = ( + request.headers.get("content-encoding", "identity").strip().lower() + ) + if content_encoding not in {"", "identity"}: + raise HTTPException(status_code=415, detail="content encoding is not supported") + content_length = request.headers.get("content-length") + if content_length is not None: + try: + declared_body_bytes = int(content_length) + except ValueError: + raise HTTPException(status_code=400, detail="invalid Content-Length") + if declared_body_bytes < 0: + raise HTTPException(status_code=400, detail="invalid Content-Length") + if declared_body_bytes > MAX_DISKSAGE_CATALOG_BODY_BYTES: + raise HTTPException( + status_code=413, + detail="DiskSage catalog preview body exceeds the 2 MiB limit", + ) + + chunks: list[bytes] = [] + body_bytes = 0 + async for chunk in request.stream(): + body_bytes += len(chunk) + if body_bytes > MAX_DISKSAGE_CATALOG_BODY_BYTES: + raise HTTPException( + status_code=413, + detail="DiskSage catalog preview body exceeds the 2 MiB limit", + ) + chunks.append(chunk) + raw_body = b"".join(chunks) + try: + batch = DiskSageCatalogCandidateBatch.model_validate_json(raw_body) + except ValidationError as exc: + safe_errors = [ + { + "loc": [str(item) for item in error["loc"]], + "type": error["type"], + "msg": "request does not satisfy the bounded DiskSage catalog contract", + } + for error in exc.errors( + include_url=False, + include_context=False, + include_input=False, + ) + ] + raise HTTPException(status_code=422, detail=safe_errors) + + preview = build_disksage_catalog_preview(batch) + _enforce_file_policy( + actor, + preview.preview_id, + "create_file_asset", + actor.tenant_id, + ) + return preview + + def _read_file_asset(asset_id: str, actor: ActorContext): asset = get_file_asset(get_store(), asset_id) if asset is None: diff --git a/src/sdp/disksage_catalog.py b/src/sdp/disksage_catalog.py new file mode 100644 index 0000000..d701ae9 --- /dev/null +++ b/src/sdp/disksage_catalog.py @@ -0,0 +1,340 @@ +"""Bounded, non-persisting DiskSage pre-copy catalog preview contracts.""" + +from __future__ import annotations + +import hashlib +import json +from collections import Counter +from datetime import datetime, timedelta, timezone +from typing import Literal + +from pydantic import BaseModel, ConfigDict, Field, model_validator + + +DISKSAGE_CATALOG_SCHEMA = "disksage.file-catalog-candidate-batch" +DISKSAGE_CATALOG_VERSION = 1 +MAX_DISKSAGE_CATALOG_BODY_BYTES = 2 * 1024 * 1024 +PRODUCTION_TIME_PRECEDENCE = ( + "embedded_metadata", + "explicit_filename_date", + "filesystem_created", + "filesystem_modified", +) + +CloudProvider = Literal["icloud", "onedrive", "google-drive"] +CloudAccountScope = Literal["personal", "organization", "shared", "unknown"] +ArchiveKind = Literal[ + "document", + "media", + "archive", + "dataset", + "backup", + "creative", + "incomplete-download", +] +Confidence = Literal["high", "medium", "low", "unknown"] +ProductionTimeSourceClass = Literal[ + "embedded_metadata", + "explicit_filename_date", + "filesystem_created", + "filesystem_modified", +] + +_HEX_64_PATTERN = r"^[0-9a-f]{64}$" +_MAX_DATETIME_EPOCH_MS = 253_402_300_799_999 +_ARTIFACT_TYPE_LABELS: dict[str, str] = { + "document": "Document", + "media": "Media asset", + "archive": "Archive package", + "dataset": "Dataset", + "backup": "Backup", + "creative": "Creative work", + "incomplete-download": "Incomplete download", +} + + +class _StrictModel(BaseModel): + model_config = ConfigDict( + extra="forbid", + str_strip_whitespace=True, + ) + + +class DiskSageMetadataEvidence(_StrictModel): + field: str = Field(min_length=1, max_length=128) + value: str = Field(min_length=1, max_length=2048) + source: str = Field(min_length=1, max_length=256) + confidence: Confidence + + +class DiskSageDatasetColumnProfile(_StrictModel): + name: str = Field(min_length=1, max_length=256) + inferred_type: str = Field(min_length=1, max_length=64) + observed_values: int = Field(ge=0) + missing_values: int = Field(ge=0) + sensitive_name: bool + + +class DiskSageDatasetProfile(_StrictModel): + format: str = Field(min_length=1, max_length=64) + sampled_rows: int = Field(ge=0) + sampled_worksheets: int = Field(ge=0) + worksheet_names: list[str] = Field(default_factory=list, max_length=128) + profile_complete: bool + sample_truncated: bool + columns: list[DiskSageDatasetColumnProfile] = Field(default_factory=list, max_length=512) + quality_warnings: list[str] = Field(default_factory=list, max_length=128) + + @model_validator(mode="after") + def bound_nested_strings(self) -> "DiskSageDatasetProfile": + if any(not value or len(value) > 256 for value in self.worksheet_names): + raise ValueError("worksheet_names entries must contain 1..256 characters") + if any(not value or len(value) > 256 for value in self.quality_warnings): + raise ValueError("quality_warnings entries must contain 1..256 characters") + return self + + +def production_time_source_class(source: str) -> ProductionTimeSourceClass: + if source.startswith("embedded:"): + return "embedded_metadata" + if source == "filename:path-token": + return "explicit_filename_date" + if source == "filesystem:created": + return "filesystem_created" + if source == "filesystem:modified-fallback": + return "filesystem_modified" + raise ValueError("unsupported production_time_source") + + +class DiskSageCatalogCandidate(_StrictModel): + candidate_fingerprint: str = Field(pattern=_HEX_64_PATTERN) + review_fingerprint: str = Field(pattern=_HEX_64_PATTERN) + destination_provider: CloudProvider + destination_account_scope: CloudAccountScope + archive_kind: ArchiveKind + bytes: int = Field(ge=0) + created_ms: int = Field(ge=0, le=_MAX_DATETIME_EPOCH_MS) + modified_ms: int = Field(gt=0, le=_MAX_DATETIME_EPOCH_MS) + production_time_ms: int = Field(gt=0, le=_MAX_DATETIME_EPOCH_MS) + production_time_source: str = Field(min_length=1, max_length=256) + production_time_confidence: Confidence + requires_review: bool + review_reasons: list[str] = Field(default_factory=list, max_length=128) + content_title: str | None = Field(default=None, min_length=1, max_length=1024) + content_authors: list[str] = Field(default_factory=list, max_length=64) + content_context: list[str] = Field(default_factory=list, max_length=128) + duration_ms: int | None = Field(default=None, ge=0) + dataset_profile: DiskSageDatasetProfile | None = None + metadata_evidence: list[DiskSageMetadataEvidence] = Field( + default_factory=list, + max_length=256, + ) + blocked_reason: str | None = Field(default=None, min_length=1, max_length=256) + + @model_validator(mode="after") + def validate_metadata_lineage(self) -> "DiskSageCatalogCandidate": + bounded_lists = ( + ("review_reasons", self.review_reasons, 256), + ("content_authors", self.content_authors, 256), + ("content_context", self.content_context, 1024), + ) + for field_name, values, max_length in bounded_lists: + if any(not value or len(value) > max_length for value in values): + raise ValueError( + f"{field_name} entries must contain 1..{max_length} characters" + ) + + source_class = production_time_source_class(self.production_time_source) + expected_field, evidence_source = { + "embedded_metadata": ("production-date", self.production_time_source), + "explicit_filename_date": ("filename-date-hint", "filename:path-token"), + "filesystem_created": ("filesystem-created-date", "filesystem:created"), + "filesystem_modified": ( + "filesystem-modified-date", + "filesystem:modified", + ), + }[source_class] + selected_date = ( + datetime(1970, 1, 1, tzinfo=timezone.utc) + + timedelta(milliseconds=self.production_time_ms) + ).date().isoformat() + selected_evidence = [ + evidence + for evidence in self.metadata_evidence + if evidence.field == expected_field and evidence.source == evidence_source + ] + if not any(evidence.value == selected_date for evidence in selected_evidence): + raise ValueError( + "selected production time must bind to matching metadata evidence" + ) + + evidence_classes: set[ProductionTimeSourceClass] = set() + for evidence in self.metadata_evidence: + if evidence.field == "production-date" and evidence.source.startswith( + "embedded:" + ): + evidence_classes.add("embedded_metadata") + elif ( + evidence.field == "filename-date-hint" + and evidence.source == "filename:path-token" + ): + evidence_classes.add("explicit_filename_date") + elif ( + evidence.field == "filesystem-created-date" + and evidence.source == "filesystem:created" + ): + evidence_classes.add("filesystem_created") + elif ( + evidence.field == "filesystem-modified-date" + and evidence.source == "filesystem:modified" + ): + evidence_classes.add("filesystem_modified") + selected_rank = PRODUCTION_TIME_PRECEDENCE.index(source_class) + if any( + PRODUCTION_TIME_PRECEDENCE.index(evidence_class) < selected_rank + for evidence_class in evidence_classes + ): + raise ValueError( + "selected production time violates metadata precedence" + ) + if source_class != "embedded_metadata" and self.production_time_confidence != "low": + raise ValueError("non-embedded production time must remain low confidence") + if self.requires_review != bool(self.review_reasons): + raise ValueError("requires_review must match the presence of review_reasons") + return self + + +class DiskSageCatalogCandidateBatch(_StrictModel): + schema_id: Literal["disksage.file-catalog-candidate-batch"] = Field(alias="schema") + version: Literal[1] + production_time_precedence: tuple[ + Literal["embedded_metadata"], + Literal["explicit_filename_date"], + Literal["filesystem_created"], + Literal["filesystem_modified"], + ] + generated_at_ms: int = Field(gt=0, le=_MAX_DATETIME_EPOCH_MS) + candidates: list[DiskSageCatalogCandidate] = Field(min_length=1, max_length=200) + + @model_validator(mode="after") + def validate_batch(self) -> "DiskSageCatalogCandidateBatch": + if self.production_time_precedence != PRODUCTION_TIME_PRECEDENCE: + raise ValueError("production_time_precedence must use the fixed policy order") + fingerprints = [item.candidate_fingerprint for item in self.candidates] + if len(fingerprints) != len(set(fingerprints)): + raise ValueError("candidate_fingerprint values must be unique within a batch") + return self + + +class DiskSageOntologyMatch(_StrictModel): + relation: Literal["hasArtifactType"] = "hasArtifactType" + target_kind: Literal["artifact_type"] = "artifact_type" + target_label: str + confidence: Literal[1.0] = 1.0 + method: Literal["disksage-archive-kind-v1"] = "disksage-archive-kind-v1" + review_status: Literal["proposed"] = "proposed" + + +class DiskSageCatalogProjection(_StrictModel): + candidate_fingerprint: str = Field(pattern=_HEX_64_PATTERN) + source_class: ProductionTimeSourceClass + ontology_matches: list[DiskSageOntologyMatch] + dataset_profile_present: bool + persistable_as_file_asset: Literal[False] = False + missing_file_asset_requirements: tuple[ + Literal["content_sha256"], + Literal["verified_distribution"], + ] = ("content_sha256", "verified_distribution") + + +class DiskSageCatalogPreviewResponse(_StrictModel): + schema_id: Literal["disksage.file-catalog-preview"] = Field(alias="schema") + version: Literal[1] + preview_id: str = Field(pattern=r"^urn:sha256:[0-9a-f]{64}$") + structural_validation: Literal["accepted"] + candidate_count: int = Field(ge=1, le=200) + total_bytes: int = Field(ge=0) + requires_review_count: int = Field(ge=0) + blocked_count: int = Field(ge=0) + production_source_counts: dict[ProductionTimeSourceClass, int] + projections: list[DiskSageCatalogProjection] + persisted: Literal[False] = False + llm_used: Literal[False] = False + copy_authorized: Literal[False] = False + eviction_authorized: Literal[False] = False + persistable_as_file_asset: Literal[False] = False + content_sha256_required: Literal[True] = True + notices: tuple[str, ...] + + +def preview_id(batch: DiskSageCatalogCandidateBatch) -> str: + """Bind the response to non-sensitive candidate/review identities only.""" + + identity = { + "schema": batch.schema_id, + "version": batch.version, + "precedence": batch.production_time_precedence, + "generated_at_ms": batch.generated_at_ms, + "candidates": [ + { + "candidate_fingerprint": candidate.candidate_fingerprint, + "review_fingerprint": candidate.review_fingerprint, + } + for candidate in batch.candidates + ], + } + digest = hashlib.sha256( + json.dumps( + identity, + ensure_ascii=True, + separators=(",", ":"), + sort_keys=True, + ).encode("utf-8") + ).hexdigest() + return f"urn:sha256:{digest}" + + +def build_disksage_catalog_preview( + batch: DiskSageCatalogCandidateBatch, +) -> DiskSageCatalogPreviewResponse: + """Build a deterministic ontology preview without persistence or LLM calls.""" + + projections: list[DiskSageCatalogProjection] = [] + source_counts: Counter[ProductionTimeSourceClass] = Counter() + for candidate in batch.candidates: + source_class = production_time_source_class(candidate.production_time_source) + source_counts[source_class] += 1 + projections.append( + DiskSageCatalogProjection( + candidate_fingerprint=candidate.candidate_fingerprint, + source_class=source_class, + ontology_matches=[ + DiskSageOntologyMatch( + target_label=_ARTIFACT_TYPE_LABELS[candidate.archive_kind] + ) + ], + dataset_profile_present=candidate.dataset_profile is not None, + ) + ) + + return DiskSageCatalogPreviewResponse( + schema="disksage.file-catalog-preview", + version=1, + preview_id=preview_id(batch), + structural_validation="accepted", + candidate_count=len(batch.candidates), + total_bytes=sum(candidate.bytes for candidate in batch.candidates), + requires_review_count=sum(candidate.requires_review for candidate in batch.candidates), + blocked_count=sum(candidate.blocked_reason is not None for candidate in batch.candidates), + production_source_counts={ + source: source_counts[source] for source in PRODUCTION_TIME_PRECEDENCE + }, + projections=projections, + notices=( + "preview-only-no-persistence", + "preview-does-not-authorize-copy-or-eviction", + "content-sha256-and-verified-distribution-required-before-file-asset-ingest", + "ontology-matches-remain-proposed-until-steward-review", + "response-omits-content-metadata-and-storage-coordinates", + ), + ) diff --git a/tests/test_file_knowledge.py b/tests/test_file_knowledge.py index d30eaa1..6edf3fc 100644 --- a/tests/test_file_knowledge.py +++ b/tests/test_file_knowledge.py @@ -2,6 +2,7 @@ import json from copy import deepcopy +from datetime import datetime, timezone from importlib.resources import files from io import BytesIO from pathlib import Path @@ -22,6 +23,13 @@ chunk_text, extract_document_text, ) +from sdp.disksage_catalog import ( + DISKSAGE_CATALOG_SCHEMA, + MAX_DISKSAGE_CATALOG_BODY_BYTES, + PRODUCTION_TIME_PRECEDENCE, + DiskSageCatalogCandidateBatch, + build_disksage_catalog_preview, +) from sdp.file_ontology import ( FileAsset, @@ -686,6 +694,247 @@ def verify(token): monkeypatch.setattr(api_module.authz, "verify_oidc_jwks_token", verify) +def disksage_candidate(**overrides: object) -> dict[str, object]: + production_time = int( + datetime(2026, 1, 2, 10, 30, tzinfo=timezone.utc).timestamp() * 1000 + ) + values: dict[str, object] = { + "candidate_fingerprint": "c" * 64, + "review_fingerprint": "d" * 64, + "destination_provider": "icloud", + "destination_account_scope": "personal", + "archive_kind": "document", + "bytes": 4096, + "created_ms": production_time + 1000, + "modified_ms": production_time + 2000, + "production_time_ms": production_time, + "production_time_source": "embedded:ooxml:created", + "production_time_confidence": "high", + "requires_review": False, + "review_reasons": [], + "content_title": "private title", + "content_authors": ["private author"], + "content_context": ["private context"], + "duration_ms": None, + "dataset_profile": None, + "metadata_evidence": [ + { + "field": "production-date", + "value": "2026-01-02", + "source": "embedded:ooxml:created", + "confidence": "high", + }, + { + "field": "filename-date-hint", + "value": "2025-12-31", + "source": "filename:path-token", + "confidence": "low", + }, + { + "field": "filesystem-created-date", + "value": "2026-01-02", + "source": "filesystem:created", + "confidence": "low", + }, + { + "field": "filesystem-modified-date", + "value": "2026-01-02", + "source": "filesystem:modified", + "confidence": "medium", + }, + ], + "blocked_reason": None, + } + values.update(overrides) + return values + + +def disksage_batch(**overrides: object) -> dict[str, object]: + values: dict[str, object] = { + "schema": DISKSAGE_CATALOG_SCHEMA, + "version": 1, + "production_time_precedence": list(PRODUCTION_TIME_PRECEDENCE), + "generated_at_ms": 1_784_900_000_000, + "candidates": [disksage_candidate()], + } + values.update(overrides) + return values + + +def test_disksage_pre_copy_preview_is_pure_and_does_not_echo_private_content(): + batch = DiskSageCatalogCandidateBatch.model_validate(disksage_batch()) + + preview = build_disksage_catalog_preview(batch) + serialized = preview.model_dump_json() + + assert preview.structural_validation == "accepted" + assert preview.candidate_count == 1 + assert preview.production_source_counts["embedded_metadata"] == 1 + assert preview.projections[0].ontology_matches[0].target_label == "Document" + assert preview.persisted is False + assert preview.llm_used is False + assert preview.copy_authorized is False + assert preview.eviction_authorized is False + assert preview.persistable_as_file_asset is False + assert preview.content_sha256_required is True + assert "private title" not in serialized + assert "private author" not in serialized + assert "private context" not in serialized + + +def test_disksage_pre_copy_contract_enforces_metadata_precedence_and_evidence_binding(): + filename_selected = disksage_candidate( + production_time_source="filename:path-token", + production_time_confidence="low", + production_time_ms=int( + datetime(2025, 12, 31, tzinfo=timezone.utc).timestamp() * 1000 + ), + requires_review=True, + review_reasons=["production-date-not-from-embedded-metadata"], + ) + with pytest.raises(ValueError, match="violates metadata precedence"): + DiskSageCatalogCandidateBatch.model_validate( + disksage_batch(candidates=[filename_selected]) + ) + + missing_selected_evidence = disksage_candidate(metadata_evidence=[]) + with pytest.raises(ValueError, match="bind to matching metadata evidence"): + DiskSageCatalogCandidateBatch.model_validate( + disksage_batch(candidates=[missing_selected_evidence]) + ) + + +@pytest.mark.parametrize( + ("production_time_source", "production_time", "evidence_fields", "source_class"), + [ + ( + "filename:path-token", + datetime(2025, 12, 31, tzinfo=timezone.utc), + { + "filename-date-hint", + "filesystem-created-date", + "filesystem-modified-date", + }, + "explicit_filename_date", + ), + ( + "filesystem:created", + datetime(2026, 1, 2, tzinfo=timezone.utc), + {"filesystem-created-date", "filesystem-modified-date"}, + "filesystem_created", + ), + ( + "filesystem:modified-fallback", + datetime(2026, 1, 2, tzinfo=timezone.utc), + {"filesystem-modified-date"}, + "filesystem_modified", + ), + ], +) +def test_disksage_pre_copy_contract_accepts_only_the_next_available_fallback( + production_time_source, + production_time, + evidence_fields, + source_class, +): + baseline = disksage_candidate() + evidence = [ + item + for item in baseline["metadata_evidence"] + if item["field"] in evidence_fields + ] + candidate = disksage_candidate( + production_time_source=production_time_source, + production_time_confidence="low", + production_time_ms=int(production_time.timestamp() * 1000), + metadata_evidence=evidence, + requires_review=True, + review_reasons=["production-date-not-from-embedded-metadata"], + ) + + batch = DiskSageCatalogCandidateBatch.model_validate( + disksage_batch(candidates=[candidate]) + ) + preview = build_disksage_catalog_preview(batch) + + assert preview.projections[0].source_class == source_class + + +def test_disksage_pre_copy_contract_rejects_paths_and_unbounded_shape(): + payload = disksage_batch() + candidate = dict(payload["candidates"][0]) + candidate["src"] = "/Users/example/Downloads/private-report.pdf" + payload["candidates"] = [candidate] + + with pytest.raises(ValueError, match="Extra inputs are not permitted"): + DiskSageCatalogCandidateBatch.model_validate(payload) + + schema = DiskSageCatalogCandidateBatch.model_json_schema() + candidate_schema = schema["$defs"]["DiskSageCatalogCandidate"]["properties"] + assert "src" not in candidate_schema + assert "dst" not in candidate_schema + assert "relative_path" not in candidate_schema + assert "filename" not in candidate_schema + assert "account_id" not in candidate_schema + assert "object_id" not in candidate_schema + + +def test_disksage_pre_copy_api_requires_create_policy_and_never_uses_graph_store( + monkeypatch, +): + install_fake_oidc(monkeypatch) + + def fail_if_graph_store_is_used(): + raise AssertionError("pre-copy catalog preview must not touch graph persistence") + + monkeypatch.setattr(api_module, "get_store", fail_if_graph_store_is_used) + denied = client.post( + "/file-assets/preview/disksage", + json=disksage_batch(), + headers=auth_headers("analyst"), + ) + accepted = client.post( + "/file-assets/preview/disksage", + json=disksage_batch(), + headers=auth_headers("admin"), + ) + + assert denied.status_code == 403 + assert accepted.status_code == 200 + assert accepted.json()["schema"] == "disksage.file-catalog-preview" + assert accepted.json()["persisted"] is False + assert accepted.json()["copy_authorized"] is False + assert "private title" not in accepted.text + assert "/Users/" not in accepted.text + + +def test_disksage_pre_copy_api_redacts_validation_input_and_limits_body(monkeypatch): + install_fake_oidc(monkeypatch) + payload = disksage_batch() + candidate = dict(payload["candidates"][0]) + candidate["src"] = "/Users/example/Downloads/do-not-echo.pdf" + payload["candidates"] = [candidate] + + invalid = client.post( + "/file-assets/preview/disksage", + json=payload, + headers=auth_headers("admin"), + ) + oversized = client.post( + "/file-assets/preview/disksage", + content=b"x" * (MAX_DISKSAGE_CATALOG_BODY_BYTES + 1), + headers={ + **auth_headers("admin"), + "Content-Type": "application/json", + }, + ) + + assert invalid.status_code == 422 + assert "do-not-echo" not in invalid.text + assert "/Users/" not in invalid.text + assert oversized.status_code == 413 + + def test_file_asset_api_requires_policy_and_redacts_jsonld_locator(monkeypatch): store = InMemoryGraphStore() monkeypatch.setattr(api_module, "get_store", lambda: store) diff --git a/tests/test_migrations.py b/tests/test_migrations.py index 03523da..343d548 100644 --- a/tests/test_migrations.py +++ b/tests/test_migrations.py @@ -59,3 +59,34 @@ def test_render_sql_substitutes_configured_embedding_dimension(): def test_embedding_dimension_reads_config_default(): runner = _load_runner() assert runner._embedding_dimension() == 128 + + +def test_file_copy_lineage_migration_defines_normalized_fk_graph(): + sql = (MIGRATIONS_DIR / "0002_file_copy_lineage.sql").read_text( + encoding="utf-8" + ) + for table in [ + "file_asset_records", + "file_distribution_records", + "cloud_copy_receipts", + "file_metadata_evidence_records", + "cloud_sync_evidence_records", + ]: + assert f"CREATE TABLE IF NOT EXISTS {table}" in sql + assert sql.count("REFERENCES") == 7 + assert sql.count("ON DELETE RESTRICT") == 7 + assert "source_locator_sha256" in sql + assert "source_relative_path" not in sql + assert "provider_sync_confirmed" in sql + assert "local_copy_verified" in sql + assert "content_blake3" in sql + assert "production_time_source" in sql + assert "file_metadata_evidence_selected_idx" in sql + + +def test_all_migration_files_are_statement_splitter_compatible(): + runner = _load_runner() + for path in sorted(MIGRATIONS_DIR.glob("*.sql")): + statements = runner._statements(path.read_text(encoding="utf-8")) + assert statements, f"migration {path.name} must contain executable statements" + assert statements[-1].startswith("INSERT INTO schema_migrations")