From bf1e644e9aa7f511d4e9aafa3ae389b96cd0b1f9 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Wed, 12 Aug 2026 09:13:39 +0900 Subject: [PATCH 01/18] docs(adr): define LLM-native email writing guidance boundary --- ...kspan-backed-llm-email-writing-guidance.md | 197 ++++++++++++++++++ 1 file changed, 197 insertions(+) create mode 100644 docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md diff --git a/docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md b/docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md new file mode 100644 index 000000000..932d89fc3 --- /dev/null +++ b/docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md @@ -0,0 +1,197 @@ +# ADR 0001: Inkspan-backed, LLM-native email writing guidance + +Status: Proposed + +## Context + +Naruon currently supports one-shot reply drafting, but it does not provide a Grammarly-like authoring loop in which the user writes directly, sees passage-level guidance, understands each proposed change, and individually applies or ignores suggestions. The required guidance spans spelling, grammar, spacing, punctuation, clarity, concision, discourse structure, workplace pragmatics, audience appropriateness, technical precision, actionability, and preservation of the user's intended request. + +These are contextual language judgments. A phrase can be appropriate in a quotation and inappropriate as a direct rebuke; the same pragmatic problem can be expressed with unrelated words; identical prose can reasonably be interpreted differently when the source thread, recipient roles, or requested outcome changes. A keyword list, regular expression, sender-domain table, fixed “aggressive phrase” dictionary, or positional repair rule cannot establish those meanings. + +ContextualWisdomLab/inkspan already owns the modular WYSIWYG editor, revision evidence, revision-scoped W3C text-position evidence, guarded document mutation, safe serialization, and host-owned model-assistance boundary. ContextualWisdomLab/fast-mlsirm now provides a provider-neutral `ContextualOrchestratorJudge`, strict criterion-level JSON parsing, explicit polytomous categories, runtime-derived acceptance, an IRT response-row bridge, and a deliberate prohibition on keyword or positional repair. Naruon should compose these capabilities rather than recreating an editor or inventing a lexical classifier. + +## Alternatives considered + +- **Keep the current textarea and add a single “rewrite professionally” button.** Rejected because whole-document replacement hides individual reasons, weakens author control, and cannot safely bind delayed output to the reviewed revision. +- **Add a keyword/regex tone and grammar checker in Naruon.** Rejected because lexical triggers are not evidence of intent, pragmatics, grammaticality, or technical correctness and would fail across paraphrases, quotations, languages, and recipient context. +- **Let Inkspan call an LLM and own email semantics.** Rejected because Inkspan must remain provider-neutral and reusable; email thread context, tenant policy, provider credentials, retention, and semantic judgment belong to Naruon and its orchestration services. +- **Use one LLM response as both candidate generator and unquestioned authority.** Rejected because LLM judges are fallible measuring instruments with position, verbosity, self-preference, artifact, multilingual, and drift risks. +- **Naruon owns email context and review orchestration; Inkspan owns the revision-safe editor surface; fast-mlsirm owns judge calibration and criterion-response contracts.** Selected because it separates product semantics, deterministic document integrity, model routing, and measurement validation. + +## Decision + +Naruon will implement email writing guidance as an LLM-native, context-grounded review workflow rendered through Inkspan's generic writing-diagnostic contract. + +### Semantic judgment + +All claims that prose is misspelled, ungrammatical, unclear, verbose, poorly structured, pragmatically inappropriate, technically imprecise, insufficiently actionable, or likely to alter the author's intended request must originate from an LLM review workflow. Naruon production code must not manufacture those judgments from keywords, regexes, phrase dictionaries, sender domains, recipient counts, language names, or word positions. + +Deterministic code remains mandatory for non-semantic integrity: + +- request authentication and tenant/workspace scope; +- source-email and thread re-read under server authority; +- JSON/schema, type, enum, length, count, duplicate, and resource-bound validation; +- prompt-data delimiting and prompt-injection resistance; +- document revision, projection, Unicode code-point, grapheme, and selector validation; +- stale-result, overlap, and conflict rejection; +- provider allowlisting, credential lookup, timeouts, retries, circuit breaking, and audit metadata; +- safe rendering and ordinary editor transaction/undo behavior. + +These deterministic checks may reject malformed or stale model output. They may not replace a failed model decision with a lexical heuristic. + +### Model workflow + +The online workflow has distinct roles: + +1. **Context builder** re-reads the selected source email, thread, subject, sender, recipients, reply purpose, and user-authored draft under the current authorization scope. +2. **Candidate reviewer** produces bounded passage-level diagnostic proposals and whole-document guidance from that context. +3. **Independent criterion judge** evaluates each candidate against an explicit rubric through contextual-orchestrator. The preferred adapter is fast-mlsirm's `ContextualOrchestratorJudge` after an immutable, hash-locked compatible release is available. +4. **Deterministic result validator** admits only exact-schema, revision-bound, selector-valid diagnostics. +5. **Optional adjudication** uses an independent model or deeper contextual-orchestrator conduct workflow when calibrated uncertainty or cross-model disagreement is high. +6. **Inkspan** displays admitted proposals and applies only explicit user-approved replacements under the current revision. + +The model workflow must be capable of abstaining. Provider failure, malformed output, insufficient context, judge disagreement, or calibration-policy rejection returns no admitted diagnostic for that claim. Naruon continues to provide editing and sending; it does not fall back to keyword matching. + +### fast-mlsirm role + +fast-mlsirm is the judge measurement and calibration layer, not the editor, mail client, or source of email context. + +It is used to: + +- define criterion-level dichotomous or polytomous response contracts; +- validate judge response matrices; +- fit and compare item difficulty, discrimination, rater/model effects, and latent writing-quality dimensions where supported; +- detect differential item functioning across language, recipient-role, organizational-context, thread-depth, and other prespecified groups; +- measure calibration, reliability, prompt/model test-retest behavior, and drift; +- publish a versioned judge-admission policy consumed by Naruon. + +Live email review must not block on model fitting. Runtime review consumes an already-published policy version; fitting, simulation, DIF, ablation, and drift analysis run offline or nearline. A future fast-mlsirm service may expose policy artifacts, but it does not receive authority to send or mutate email. + +### contextual-orchestrator role + +Every semantic model call routes through contextual-orchestrator or a contract-compatible provider-neutral port. The orchestrator decides provider/model routing, single-model versus multi-agent computation, workflow depth, role-specific reasoning effort, tracing, failure policy, and cost evidence. Naruon does not use keyword routing to select a model or decide a semantic category. + +Operational routing may depend on explicit user mode, document size, configured compute policy, provider availability, calibrated model disagreement, or structured uncertainty. These are routing controls, not lexical classifiers. + +### User authority + +Writing diagnostics are advisory. The user can inspect, apply, ignore, dismiss, or request another explanation. Remaining suggestions do not block editing or sending by default. Naruon may expose an enterprise policy surface later, but such a policy must be explicit, separately authorized, and cannot be inferred from a diagnostic confidence or priority label. + +A suggestion never silently weakens a deadline, changes a requested deliverable, alters responsibility, removes a material fact, or converts a firm request into a noncommittal statement. Intent, fact, request-strength, deadline, actor, and actionability preservation are explicit judge criteria. + +## Rubric contract + +The first review policy will use independently scored observable criteria such as: + +- `issue_support`: the cited span actually supports the claimed issue; +- `span_fidelity`: the selector targets the smallest sufficient passage; +- `replacement_correctness`: the proposed wording resolves the identified issue; +- `intent_preservation`: the proposal preserves the author's intended outcome; +- `fact_preservation`: names, quantities, dates, technical claims, and commitments are not invented or removed; +- `request_strength_preservation`: firmness and accountability are not softened without explicit author direction; +- `audience_pragmatics`: wording fits the recipient, copied audience, hierarchy, and thread context; +- `technical_precision`: terminology, metrics, and causal claims remain technically defensible; +- `actionability`: responsible actor, requested artifact, timing, and response channel are clear where the source intent requires them; +- `explanation_quality`: the explanation is specific, evidence-based, and useful to the author. + +Category anchors are ordered and versioned. Overall acceptance is derived from criterion evidence and a published policy, not from a hidden word list or a model's free-form “accepted” token alone. + +## Data flow + +```mermaid +sequenceDiagram + participant User + participant Inkspan + participant Naruon + participant Orchestrator as contextual-orchestrator + participant Judge as fast-mlsirm judge adapter + + User->>Inkspan: Write or edit reply + Inkspan->>Naruon: Review request + revision + projection + Naruon->>Naruon: Re-read source email/thread under scope + Naruon->>Orchestrator: Candidate review with untrusted context + Orchestrator-->>Naruon: Structured diagnostic candidates + Naruon->>Judge: Criterion-level independent evaluation + Judge->>Orchestrator: Provider-neutral judge call + Orchestrator-->>Judge: Strict JSON criterion result + Judge-->>Naruon: Validated scores/categories + trace metadata + Naruon->>Naruon: Policy admission + selector/revision validation + Naruon-->>Inkspan: Admitted diagnostics only + User->>Inkspan: Apply / ignore / dismiss + Inkspan-->>Naruon: Privacy-minimized feedback event +``` + +## Consequences + +Naruon gains contextual, passage-level writing assistance that can generalize across wording and language instead of being tied to a brittle phrase list. The architecture introduces a deeper model-evaluation pipeline and therefore requires versioned rubrics, calibration evidence, provider operation, and degraded-mode behavior. Inkspan remains reusable and deterministic; fast-mlsirm remains a separable psychometric/evaluation component. + +No semantic result is guaranteed merely because an LLM returned valid JSON. Admission means the result passed the current model, rubric, calibration, and structural policy; it remains an advisory proposal. + +## Failure and recovery + +- **Candidate model unavailable:** return a review-unavailable status; preserve editing and sending. +- **Malformed or oversized candidate output:** reject; do not repair by keyword or nearest-text search. +- **Judge unavailable or malformed:** abstain or retry through contextual-orchestrator policy; never accept the candidate solely because generation succeeded. +- **Judge disagreement or low calibrated confidence:** withhold, mark for deeper adjudication, or request human review according to the published policy. +- **Document changed while reviewing:** return stale diagnostics and require a fresh review; never mutate newer content. +- **Selector no longer identifies the reviewed text:** reject that diagnostic. +- **Replacement violates Inkspan safety/schema policy:** reject without mutation. +- **fast-mlsirm calibration service unavailable:** continue using the last signed, non-expired approved policy if tenant policy permits; otherwise disable semantic review rather than inventing a fallback. + +## Security and privacy impact + +Email bodies, drafts, participants, and organizational context may contain PII and confidential business content. Naruon will not mask away names, roles, dates, or relationships when doing so would destroy the requested pragmatic and actionability analysis. Instead it will apply compensating controls: + +- tenant-approved providers and contractual no-training/no-secondary-use terms; +- encrypted transport and encrypted credential storage; +- server-authoritative tenant/workspace authorization; +- prompt/data separation and untrusted-content treatment; +- no raw email body, draft, replacement, prompt, or model output in ordinary logs or product telemetry; +- short-lived or zero-retention review sessions by default; +- explicit provider, region, model, rubric, policy, and prompt provenance; +- bounded payloads and abuse protection; +- audit events containing opaque identifiers and policy outcomes rather than source text; +- synthetic or carefully de-identified benchmark corpora for persistent evaluation artifacts. + +Raw model output is not returned to the browser by default. fast-mlsirm's judge records used in persistent calibration datasets must follow separate approved data-classification and retention policy. + +## Accessibility + +Naruon will use Inkspan's diagnostic surface rather than a hover-only overlay. Users must be able to discover the number of suggestions, navigate among them, hear category and explanation, apply or ignore them, and return focus predictably to the edited range. Color is never the only signal. Review arrival does not steal focus. + +## Compatibility and migration + +The existing draft endpoint may remain for one-shot generation during migration, but the reply composer will move from the plain textarea to the first released Inkspan version that contains the accepted writing-diagnostic contract. Naruon must consume an immutable released package with its lockfile and package-verification evidence; it will not depend on an unreleased mutable branch in production. + +The first implementation is additive and does not require a database migration if review sessions remain ephemeral. Any later persistent tables must use two-or-more-word `snake_case` names such as `email_review_session`, `writing_diagnostic_record`, `diagnostic_feedback_event`, and `judge_policy_artifact`. + +Rollback disables the review endpoint and diagnostic props while retaining the Inkspan editor and existing mail-send path. No canonical email content migration is required. + +## Verification and acceptance evidence + +Implementation cannot be accepted without: + +- semantic contrast tests in which the same keywords occur with different meanings; +- paraphrase tests in which the same issue appears without shared trigger words; +- quotation, code, proper-name, and technical-term false-positive tests; +- identical drafts evaluated under materially different authorized thread/recipient contexts; +- Korean, English, mixed-language, CJK, emoji, combining-mark, and hostile-Unicode cases; +- gold human span/category/replacement data with inter-rater evidence; +- span precision/recall, category macro-F1, accepted-suggestion precision, intent/fact/request-strength preservation, Brier score, calibration error, and unsupported-claim rate; +- fast-mlsirm criterion response-matrix validation, category-count ablation, item/rater behavior, DIF, and drift studies where sample size supports them; +- independent-model or role ablations for candidate reviewer, judge, and adjudicator; +- prompt-injection, malformed JSON, duplicate keys, oversized payload, stale revision, overlap, and provider failure tests; +- no-keyword-fallback source and contract tests; +- Python 3.14, frontend, security, coverage, container, package-lock, and exact-head repository gates; +- production statement/branch coverage and public docstring/JSDoc requirements defined by the repository; +- current authoritative APA 7th doctoring and cross-repository traceability. + +Tests may use `NVIDIA_NIM_API_KEY` for scheduled live-model evaluation. They must not use `COPILOT_GITHUB_TOKEN` for model calls. + +## Research and standards traceability + +The accompanying doctoring record covers LLM-as-a-Judge reliability, bias, multilingual evaluation, psychometric calibration, human-centered writing support, W3C selector semantics, NIST AI 600-1, ISO/IEC 23894, and ISO/IEC 42001. This ADR does not claim that any model or language profile is validated before those studies pass. + +## Rollback or supersession + +Rollback removes or disables semantic review while preserving the editor, mail data, and send path. Supersession requires a new ADR if Naruon is proposed to use keyword-based semantic classification, allow unjudged model output to mutate drafts, move email semantics into Inkspan, put fast-mlsirm fitting in the synchronous critical path, mask context needed for valid interpretation, or make suggestions an implicit send gate. \ No newline at end of file From 5782ee3a28a0a876f0d369f47daffc7be789e665 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Wed, 12 Aug 2026 09:15:14 +0900 Subject: [PATCH 02/18] docs(spec): design Inkspan LLM email writing guidance --- ...kspan-llm-email-writing-guidance-design.md | 578 ++++++++++++++++++ 1 file changed, 578 insertions(+) create mode 100644 docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md diff --git a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md new file mode 100644 index 000000000..b5cd6cd06 --- /dev/null +++ b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md @@ -0,0 +1,578 @@ +# Inkspan-Based LLM Email Writing Guidance Design + +**Date:** 2026-08-12 +**Status:** Proposed design; no shipped feature is claimed +**Repositories:** `ContextualWisdomLab/naruon`, `ContextualWisdomLab/inkspan`, `ContextualWisdomLab/fast-mlsirm`, and `ContextualWisdomLab/contextual-orchestrator` + +## Objective + +Replace Naruon's plain reply textarea and one-shot rewrite experience with a Grammarly-like, revision-safe authoring workflow: + +- the user writes directly in Inkspan; +- Naruon reviews the current reply against the source email, complete thread, recipients, reply purpose, and tenant policy; +- an LLM produces passage-level candidate diagnostics; +- an independent LLM-as-a-Judge evaluates those candidates under a versioned rubric; +- deterministic code admits only strict, current-revision, selector-valid results; +- Inkspan displays, navigates, applies, ignores, and dismisses suggestions; +- fast-mlsirm calibrates judge behavior and publishes versioned admission evidence; +- no keyword, regex, phrase dictionary, domain list, or positional repair acts as a semantic fallback. + +The feature is writing guidance, not a send-risk score or mandatory gate. + +## User experience + +### Continuous authoring + +The reply composer is an Inkspan editor in email-compatible HTML mode. The user can type, paste safe rich content, use formatting, undo, and send even when semantic review is unavailable. + +### Incremental review + +After a bounded debounce, Naruon reviews the changed paragraph plus enough authorized document/thread context to interpret it. Incremental review is optimized for responsive feedback, but semantic correctness is still model-based. It does not run a local trigger-word detector before deciding what category to assign. + +### Deep review + +The author can explicitly request `전체 메일 검토`. Deep review evaluates the complete draft for cross-sentence structure, repetition, actor ambiguity, request completeness, audience pragmatics, technical precision, and intent preservation. It may use contextual-orchestrator conduct mode with separate reviewer, critic, judge, and adjudicator roles. + +### Suggestions + +Each admitted diagnostic shows: + +- affected passage; +- category and concise title; +- evidence-based explanation; +- optional replacement; +- confidence/admission provenance appropriate for the UI; +- Apply, Ignore, Dismiss, and Explain actions. + +Apply changes only the bound passage. Ignore and Dismiss do not alter the document. New asynchronous results do not steal focus. Stale results are visibly invalidated and cannot apply. + +### Whole-document guidance + +Some issues cannot be represented as one replacement range. The response may therefore also contain non-mutating document guidance, such as: + +- inferred purpose summary; +- likely reader interpretation; +- unclear actor or missing deliverable; +- missing response deadline or channel; +- structural reordering suggestion; +- unresolved factual/technical verification question. + +Document guidance is advisory and cannot mutate the editor without a separate, explicit author action. + +## Architecture + +```mermaid +flowchart TB + U[Author] --> I[Inkspan reply editor] + I -->|revision + projection + draft| R[Naruon email-writing API] + R --> A[Authorized email/thread context builder] + A --> O[contextual-orchestrator] + O --> C[Candidate reviewer] + C --> J[Independent fast-mlsirm judge adapter] + J --> O + J --> P[Versioned judge admission policy] + P --> V[Deterministic review validator] + V -->|admitted diagnostics| I + I -->|apply/ignore/dismiss feedback| F[Naruon feedback endpoint] + F --> E[Evaluation evidence pipeline] + E --> M[fast-mlsirm calibration / DIF / drift] + M --> P +``` + +### Component boundaries + +#### `email_writing_api` + +Owns authenticated HTTP contracts, tenant/workspace scope, rate limits, payload bounds, idempotency, and error mapping. + +#### `email_writing_context_service` + +Re-reads the source email and thread by opaque server-authorized identifiers. It creates a bounded context bundle with subject, selected source message, relevant prior messages, sender/recipient metadata, explicit reply objective, and current draft. Client-provided recipient roles or thread text are never authoritative. + +#### `email_writing_review_service` + +Builds untrusted-data-delimited prompts, calls contextual-orchestrator, validates the candidate response, invokes the independent judge, applies the published admission policy, and returns a structured review result. + +#### `writing_review_judge_port` + +Defines the Naruon-side interface to an LLM judge. The preferred adapter consumes a released, hash-locked fast-mlsirm package and delegates every model call through contextual-orchestrator. The port allows testing and future service deployment without leaking fast-mlsirm internals into the API router. + +#### `writing_diagnostic_validator` + +Performs only deterministic structural and revision validation. It does not infer semantics. + +#### `InkspanReplyEditor` + +A small Naruon `'use client'` boundary around the released Inkspan component. It owns editor refs, revision capture, diagnostic props, current-review state, feedback callbacks, email serialization, focus, and send integration. + +#### `judge_policy_registry` + +Provides signed/versioned policy artifacts: rubric version, category anchors, accepted model/provider set, calibration scope, language profiles, thresholds or posterior admission rules, expiry, and rollback version. The first implementation may load a checked-in bounded policy artifact; long-term persistence belongs in a two-word snake_case registry object. + +## API contracts + +### `POST /api/email-writing/reviews` + +Creates a bounded semantic review for the current exact editor revision. + +```json +{ + "source_email_id": 184, + "document_revision": "\"sha256-BASE64URL\"", + "projection_name": "inkspan-prosemirror-text", + "projection_version": 1, + "draft_plain_text": "안녕하세요. ...", + "language_tag": "ko", + "review_mode": "incremental", + "changed_selector": { + "type": "TextPositionSelector", + "start": 0, + "end": 42 + }, + "reply_objective": "자료 범위와 회신 일정을 명확히 확인" +} +``` + +Rules: + +- `source_email_id` is scoped and re-read server-side. +- `document_revision`, projection, and draft must describe one exact Inkspan snapshot. +- `changed_selector` is required for incremental mode and omitted for deep mode. +- `reply_objective` is user guidance, not authority to override system policy. +- unexpected fields are rejected. +- raw email/thread data is not accepted from the browser as authoritative context. + +Response: + +```json +{ + "review_session_id": "email_review_01J...", + "document_revision": "\"sha256-BASE64URL\"", + "projection_name": "inkspan-prosemirror-text", + "projection_version": 1, + "review_status": "completed", + "diagnostics": [ + { + "diagnostic_id": "writing_diagnostic_01J...", + "document_revision": "\"sha256-BASE64URL\"", + "projection_name": "inkspan-prosemirror-text", + "projection_version": 1, + "selector": { + "type": "TextPositionSelector", + "start": 12, + "end": 35 + }, + "category_code": "audience_pragmatics", + "priority": "important", + "title": "질문의 목적을 먼저 제시하는 편이 명확합니다", + "explanation": "현재 표현은 확인 요청보다 상대 답변의 타당성을 평가하는 반문으로 읽힐 수 있습니다.", + "suggested_replacement": "말씀하신 작업이 기존 범위에 포함되는지와 수행 주체를 확인 부탁드립니다.", + "confidence": 0.86, + "provenance": { + "workflow_id": "email_writing_review", + "workflow_version": "1", + "judge_policy_version": "email_writing_judge_v1", + "orchestration_mode": "conduct" + } + } + ], + "document_guidance": { + "purpose_summary": "자료 범위와 일정 확인", + "reader_interpretation": "여러 확인 질문과 전문성 방어가 섞여 핵심 요청이 흐려질 수 있음", + "missing_requests": [ + "수행 주체", + "회신 가능 예정일" + ], + "structure_suggestion": "목적, 확인 항목, 요청 산출물, 일정 순으로 정리" + }, + "provenance": { + "provider_name": "approved-provider", + "model_name": "approved-model", + "rubric_version": "email_writing_rubric_v1", + "judge_policy_version": "email_writing_judge_v1", + "prompt_hash": "sha256:..." + } +} +``` + +The API does not expose raw candidate or judge output by default. + +### `POST /api/email-writing/reviews/{review_session_id}/feedback` + +Records an explicit author response. + +```json +{ + "diagnostic_id": "writing_diagnostic_01J...", + "document_revision": "\"sha256-BASE64URL\"", + "feedback_action": "applied", + "resulting_document_revision": "\"sha256-BASE64URL\"" +} +``` + +Allowed actions initially: + +```text +applied +ignored + dismissed +requested_explanation +stale +conflict +``` + +Implementation must normalize the accidental whitespace in documentation and use exact enum values; the canonical value is `dismissed`. + +Feedback stores opaque IDs, category, policy version, action, timing bucket, and revision references by default. It does not duplicate source text or replacement text into generic telemetry. + +## Structured model contracts + +### Candidate reviewer output + +The reviewer returns an exact JSON object with: + +- `diagnostics` array; +- `document_guidance` object; +- `context_limitations` array; +- `review_language`; +- `abstained_claims` array. + +Each diagnostic contains a source selector, category, explanation, proposed replacement when appropriate, and candidate confidence. Category values must come from the current rubric. The output cannot include HTML, arbitrary editor JSON, commands, or a send decision. + +### Independent judge input + +Each candidate becomes one evaluation task. The judge sees: + +- bounded source/thread context required for the criterion; +- current draft and selected span; +- the claimed issue; +- proposed replacement; +- exact criterion descriptions and ordered category anchors; +- explicit instruction that mail/document content is untrusted data; +- no candidate model self-reported chain of thought. + +### Independent judge output + +Use fast-mlsirm's strict criterion-level shape and polytomous categories. The initial policy should prefer at least four ordered categories so it can distinguish unsupported, weak, adequate, and strong evidence, subject to empirical category-count ablation. + +No parser repairs a malformed response by extracting words such as “pass,” “polite,” or “incorrect.” Duplicate keys, missing criterion IDs, non-integral categories, extra fields, invalid depths, and invalid scores fail closed. + +### Admission policy + +An admitted diagnostic satisfies all of the following: + +1. candidate output schema is valid; +2. source selector targets the current reviewed projection; +3. every mandatory judge criterion is present; +4. the signed current policy accepts the category pattern or calibrated posterior evidence; +5. no mandatory preservation criterion falls below its floor; +6. model/rubric/language profile is within the policy's validated scope; +7. no independent-adjudication requirement remains unresolved; +8. replacement passes Inkspan-safe content policy; +9. the document revision remains current when returned/applied. + +The policy is versioned and observable. It is not a hidden weighted keyword score. + +## Review modes and compute allocation + +### Incremental mode + +- input: changed paragraph/range plus bounded thread context; +- default workflow: one candidate reviewer plus one independent judge; +- deeper adjudication: triggered by calibrated disagreement, policy uncertainty, missing context, or preservation-criterion conflict; +- goal: responsive feedback without sacrificing explicit judge validation. + +### Deep mode + +- input: full draft, selected source email, relevant thread history, recipient roles, and reply objective; +- workflow: decomposed reviewer roles for mechanics, discourse/actionability, pragmatics, and technical precision; independent judge; optional adjudicator; +- output: passage diagnostics plus whole-document guidance; +- speed is secondary to quality, but payload and step bounds remain explicit. + +### Operational mode selection + +The explicit user action chooses incremental or deep review. Document length and provider availability may alter batching. Semantic escalation depends on model/judge evidence and calibrated uncertainty, not lexical triggers. + +## fast-mlsirm calibration plan + +### Measurement unit + +A criterion on a candidate diagnostic is an evaluation item. A model/provider/prompt configuration acts as a rater. Human expert decisions provide reference evidence, not assumed infallible truth. + +### Response matrix + +For each diagnostic candidate, collect ordered category responses for multiple criteria and, where feasible, repeated raters/models/prompts. Use fast-mlsirm's response-matrix validation before fitting. + +### Analyses + +- item/category frequency and sparse-category checks; +- criterion difficulty and discrimination; +- model/rater severity and interaction where supported; +- latent writing-quality and preservation dimensions; +- reliability and prompt/model test-retest; +- Brier score and calibration curves for acceptance probabilities; +- differential item functioning by language, organization role, recipient configuration, thread depth, document length, and review mode; +- temporal drift across model, prompt, rubric, and policy releases; +- category-count ablation; +- reasoning-effort and single-model versus multi-agent ablation; +- human consequence analysis for overcorrection and missed issues. + +### Policy publication + +A calibration run emits a signed or integrity-bound `judge_policy_artifact` with: + +- policy version and creation/expiry times; +- compatible fast-mlsirm, Naruon, Inkspan, and orchestrator contract versions; +- approved model/provider/rubric/language profiles; +- category anchors; +- admission and preservation rules; +- calibration and DIF summary; +- sample and dataset provenance hashes; +- known limitations and rollback policy. + +Naruon runtime consumes the artifact but does not refit it. + +## Benchmark design + +### Human-authored cases + +Build a consented, de-identified or synthetic reconstruction corpus covering: + +- internal status requests; +- vendor/client scope disputes; +- deadline and responsibility clarification; +- incident reporting; +- meeting coordination; +- technical review; +- executive updates; +- apology/correction; +- Korean, English, mixed-language, and code-switched mail. + +Do not persist real confidential mail bodies merely because they are available in Naruon. + +### Contrast sets proving semantic—not keyword—behavior + +1. **Same words, different meaning:** a phrase appears inside a quotation, a neutral incident transcript, and a direct interpersonal rebuke. +2. **Same issue, different words:** multiple paraphrases express public blame without sharing trigger terms. +3. **Proper name/code protection:** a suspected spelling form appears as a product name, identifier, URL, file path, quotation, or code sample. +4. **Context shift:** an identical draft is evaluated with a one-to-one peer recipient, a large executive CC list, and an external customer thread. +5. **Intent preservation:** a forceful but legitimate deadline request must not be weakened merely to sound polite. +6. **Technical precision:** a confident but unsuitable metric request receives technical guidance even without any “rude” wording. +7. **Non-issue negative controls:** terse but acceptable operational messages should not be expanded gratuitously. + +### Metrics + +- issue-level and category-level precision, recall, macro-F1; +- selector span intersection-over-union and exact/smallest-sufficient-span rate; +- replacement grammaticality and correctness; +- intent, fact, actor, deadline, request-strength, and technical-claim preservation; +- unsupported-claim and hallucinated-fact rate; +- accepted-suggestion precision and ignore rate; +- human inter-rater agreement and adjudicated disagreement; +- Brier score, expected calibration error, reliability diagrams; +- judge/category response validity and sparse-category behavior; +- DIF and drift evidence; +- latency, token, step, and cost distributions by mode; +- stale-result and conflict rates; +- accessibility task success. + +No single scalar “email risk score” is a release criterion. + +## Data model + +Phase 1 can keep review state ephemeral. If persistence is introduced, objects use two-or-more-word `snake_case` names: + +```text +email_review_session +writing_diagnostic_record +diagnostic_feedback_event +judge_policy_artifact +judge_calibration_run +review_model_profile +review_rubric_version +review_evaluation_case +``` + +`email_review_session` stores opaque IDs, owner scope, source email reference, document revision, model/rubric/policy versions, status, and timestamps. Raw source email/draft text is not duplicated by default. + +`writing_diagnostic_record` is optional and encrypted if policy requires persistence. It never becomes canonical email content. + +## Privacy and PII controls + +Naruon does not blindly mask names, roles, organizations, dates, quantities, or relationships required to judge pragmatics and actionability. It instead uses: + +- tenant policy and explicit provider approval; +- encryption in transit and at rest where persisted; +- encrypted credential registry; +- region/provider/model restrictions; +- contractual no-training/no-secondary-use controls; +- short-lived review context and no ordinary raw-content logging; +- opaque IDs and privacy-minimized feedback telemetry; +- synthetic/de-identified long-lived evaluation corpora; +- access-controlled export and deletion procedures; +- audit events that record policy/model outcomes without copying mail text. + +A tenant can disable remote review or require an approved local provider. If no allowed model is available, the feature is unavailable; no lexical fallback runs. + +## Prompt-injection boundary + +Every source email, quoted thread, draft, recipient display name, signature, attachment-derived text, and user reply objective is untrusted content. The prompt uses explicit delimited JSON/data blocks and tells every role that content cannot change the rubric, tool access, output schema, provider policy, or system instruction. + +Candidate and judge roles receive only the minimum required context. They have no arbitrary tool access. The response parser accepts only the exact schema and rejects duplicate keys, extra fields, unsupported identifiers, excessive nesting, and unbounded strings. + +## Error contract + +Recommended HTTP/application statuses: + +```text +review_completed +review_abstained +review_unavailable +review_stale +review_rejected +context_insufficient +policy_unavailable +provider_unavailable +judge_disagreement +``` + +External responses do not expose raw provider errors, prompts, credentials, server URLs, or email text. Server logs retain typed error codes and trace IDs, not raw source content. + +## Observability + +Allowed ordinary metrics: + +- review count/status; +- mode; +- document/changed-range length bucket; +- diagnostic count/category counts; +- policy/rubric/model profile identifiers; +- latency, token, step, retry, and cost buckets; +- admission/abstention/disagreement counts; +- feedback action counts; +- stale/conflict counts; +- provider health. + +Disallowed by default: + +- source email/draft text; +- selected span text; +- replacement text; +- full explanation; +- prompts and raw responses; +- participant addresses/names; +- document envelopes; +- secrets. + +## Accessibility + +The integration must preserve Inkspan's diagnostic accessibility contract: + +- keyboard navigation among suggestions; +- accessible category, count, and explanation; +- non-color-only range indication; +- predictable focus when opening, applying, ignoring, or closing a card; +- polite status for review completion and stale invalidation; +- no focus theft from asynchronous arrival; +- screen-reader equivalent to pointer/hover behavior; +- mobile/touch actions with sufficiently large targets. + +## Security and threat cases + +Tests include: + +- source email instructing the model to ignore the system or approve every phrase; +- quoted JSON pretending to be judge output; +- duplicate-key and deep-nesting model responses; +- oversized thread, draft, explanation, replacement, and diagnostic arrays; +- hostile Unicode, bidi controls, combining marks, grapheme splits, HTML, links, and inline images; +- cross-tenant source email IDs; +- stale revision and selector reuse; +- feedback ID forgery; +- provider base URL SSRF/DNS-rebinding attempts under existing Naruon controls; +- raw-content leakage through logs, metrics, errors, or traces; +- judge self-preference and candidate/judge same-model ablation; +- malicious replacement attempting unsafe markup or credential exfiltration. + +## Testing layers + +### Backend contract tests + +- request/response schemas and extra-field rejection; +- auth and source re-read; +- prompt/data delimiting; +- exact candidate and judge parser behavior; +- no keyword/positional repair; +- policy admission and abstention; +- provider errors and stale results; +- privacy-safe errors/telemetry. + +### fast-mlsirm integration tests + +- released adapter import and version compatibility; +- criterion/category contract; +- response-matrix conversion and validation; +- policy artifact parsing; +- offline calibration fixtures and deterministic seeds; +- scheduled live NIM evaluation with `NVIDIA_NIM_API_KEY`. + +### Frontend/Inkspan integration tests + +- real Inkspan component instead of textarea; +- revision capture and review request; +- diagnostics rendering and action callbacks; +- stale invalidation after typing; +- Apply/Ignore/Dismiss/Explain; +- undo, focus, SSR, accessibility, email serialization, and send behavior; +- review outage does not disable editing/sending. + +### End-to-end tests + +- import/read an actual test email fixture; +- author a response; +- receive model-backed diagnostics through a controlled provider fixture; +- apply one suggestion and ignore another; +- verify the serialized outgoing message and thread headers; +- assert no raw mail content enters captured logs/metrics; +- run a live-model scheduled evaluation separately from deterministic CI. + +## Release and migration plan + +1. Inkspan design/ADR PR accepted. +2. Inkspan implementation plan and TDD implementation. +3. Inkspan release containing the public diagnostic contract. +4. fast-mlsirm immutable release containing the strict judge/IRT contract used by Naruon. +5. Naruon backend adapter, benchmark, and policy artifact. +6. Naruon reply editor migration to the released Inkspan package. +7. current-head CI/security/coverage/review evidence. +8. CHANGELOG/version/release evaluation. + +No unreleased branch or mutable source archive is a production dependency. Naruon's hash-locked dependency policy remains authoritative. + +## Rollback + +- disable semantic review feature flag or tenant policy; +- preserve Inkspan authoring and sending; +- keep the last safe composer value and normal draft state; +- stop emitting diagnostic requests; +- retain only bounded audit evidence required by policy; +- revert to one-shot draft generation if explicitly enabled, without claiming Grammarly-like review; +- no canonical email/database migration is required to remove ephemeral review state. + +## Documentation and traceability updates required during implementation + +- Naruon ADR index and architecture; +- API contract; +- threat model; +- test strategy; +- operability and data-retention guidance; +- product-event dictionary; +- Inkspan integration/version contract; +- fast-mlsirm judge policy/calibration traceability; +- contextual-orchestrator workflow/role contract; +- APA 7th doctoring; +- CHANGELOG and release evidence. + +## Approval boundary + +Approval of this design authorizes only the detailed implementation plan. It does not claim the feature, model accuracy, language coverage, privacy controls, or cross-repository integration are shipped. Those claims require protected-branch implementation and exact-head evidence in every affected repository. \ No newline at end of file From 95617fd71500895973c8443a38f098cb7113809f Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Wed, 12 Aug 2026 09:16:12 +0900 Subject: [PATCH 03/18] docs(doctoring): ground LLM email guidance and judge calibration --- ...-writing-guidance-and-fast-mlsirm-judge.md | 205 ++++++++++++++++++ 1 file changed, 205 insertions(+) create mode 100644 docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md diff --git a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md new file mode 100644 index 000000000..d07236a0e --- /dev/null +++ b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md @@ -0,0 +1,205 @@ +# LLM email writing guidance and fast-mlsirm judge calibration + +**Status:** Proposed architecture evidence; no production accuracy or language-coverage claim is made. +**Date:** 2026-08-12 + +## Purpose + +This record supports Naruon's decision to use contextual LLM judgment for email writing guidance while restricting deterministic code to authorization, schema, revision, selector, safety, and operational integrity. It also defines why `fast-mlsirm` is used to measure and calibrate the LLM-as-a-Judge layer instead of replacing it with keyword or regular-expression rules. + +## Current CWL implementation evidence + +`ContextualWisdomLab/fast-mlsirm` pull request #733 introduced a provider-neutral `ContextualOrchestratorJudge`, strict criterion-level JSON, explicit polytomous categories, runtime-derived acceptance, `LLMJudgeResult.to_irt_row()`, multi-item response-matrix validation, and a deliberate no-keyword/no-positional-repair boundary. All judge model calls are injected through contextual-orchestrator rather than bound to one provider. + +This is useful infrastructure, not proof that an email-writing rubric is valid. Naruon must still define the construct, criteria, category anchors, evaluation cases, language profiles, human reference process, calibration policy, and consequences of false positive or false negative guidance. + +## Why keyword matching is rejected + +### Lexical presence does not establish pragmatic function + +A phrase may be a quotation, a factual incident transcript, a proper name, code, a rhetorical example, or a direct interpersonal act. The same words can therefore support different judgments. Conversely, blame, sarcasm, ambiguity, excessive deference, or technically unsound requests can be expressed without any known trigger phrase. + +A keyword detector can be useful for exact policy terms whose lexical occurrence is itself the target. It is not a valid proxy for grammar, clarity, politeness, audience pragmatics, actionability, or technical correctness. + +### Recipient and thread context matter + +Email interpretation depends on who is speaking to whom, the copied audience, prior commitments, the purpose of the reply, and the surrounding thread. The same sentence can be acceptable in a private peer exchange but read as public rebuke in a broad executive CC thread. Naruon therefore re-reads authorized email/thread context and asks a contextual model; it does not infer the judgment from recipient count alone. + +### Deterministic fallback creates hidden false certainty + +When a model is unavailable or abstains, a lexical fallback would silently change the feature from contextual review to a different, uncalibrated classifier. The architecture instead degrades to “review unavailable” while preserving editing and sending. + +## What research supports—and does not support + +### Structured LLM evaluation is useful but fallible + +G-Eval showed that rubric-driven form filling with an LLM can correlate better with human judgments than several traditional automatic metrics in the evaluated summarization setting. MT-Bench and Chatbot Arena demonstrated scalable LLM-based evaluation and documented limitations such as position, verbosity, and self-enhancement bias. + +These results support structured criteria, independent judging, and explicit evaluation. They do not prove that an LLM is a universally reliable grammar or workplace-pragmatics authority. + +### Evaluator bias requires calibration and monitoring + +Research has documented unfair preferences, position effects, artifact sensitivity, self-preference, persuasion effects, and disagreement across judge models and evaluation settings. A valid JSON response is therefore necessary for automation but not sufficient evidence of validity. + +Naruon addresses this by: + +- separating candidate generation and judgment roles; +- using criterion-level ordered categories; +- admitting abstention and adjudication; +- measuring calibration and error by language/context group; +- versioning model, rubric, prompt, and policy; +- preserving user authority over every applied replacement. + +### Multilingual performance cannot be assumed + +Multilingual LLM-judge research reports inconsistent performance across languages and lower-resource settings. Naruon can expose a language-neutral editor and API contract, but each language profile requires its own validation evidence. A model's ability to accept Korean text is not equivalent to validated Korean pragmatic guidance. + +### Psychometrics adds measurement discipline + +An LLM judge is treated as a rater/measurement device. Criterion prompts act like items, ordered categories produce responses, and model/provider/prompt configurations can act as raters or facets. fast-mlsirm can help evaluate item difficulty/discrimination, category use, latent structure, rater severity/interaction where implemented, DIF, reliability, category-count choices, and drift. + +This does not make LLM output objective truth. It makes assumptions, uncertainty, group differences, and policy changes measurable and auditable. + +## Proposed construct map + +The initial system distinguishes at least four related constructs: + +1. **Mechanical correctness** — spelling, spacing, punctuation, morphology, and syntax. +2. **Comprehensibility and structure** — clarity, concision, reference resolution, logical ordering, and redundancy. +3. **Pragmatic fitness** — audience, relationship, public/private context, face threat, firmness, and collaborative interpretation. +4. **Operational fidelity** — preservation of facts, actors, quantities, deadlines, request strength, technical claims, and actionable next steps. + +A single scalar score is inadequate because a grammatically polished replacement can still weaken accountability or distort a technical request. Criterion-level responses and mandatory preservation floors are therefore part of the admission policy. + +## Proposed criterion-response design + +Each candidate diagnostic is judged on observable criteria such as: + +```text +issue_support +span_fidelity +replacement_correctness +intent_preservation +fact_preservation +request_strength_preservation +audience_pragmatics +technical_precision +actionability +explanation_quality +``` + +The first study should compare category counts such as 2, 3, 4, 5, and 7 rather than assuming that finer scales always improve measurement. Category anchors must be explicit, ordered, and criterion-specific where necessary. + +A runtime decision is derived from the criterion pattern and a versioned calibration policy. It is not derived from the occurrence of “rude” words or a free-form model declaration. + +## Evaluation design + +### Reference judgments + +Use trained human reviewers with a documented rubric and adjudication procedure. Human agreement and disagreement are reported; adjudicated labels are reference evidence, not infallible ground truth. + +### Contrast sets + +The benchmark contains matched cases that vary one factor at a time: + +- identical wording used as quotation versus direct speech act; +- paraphrases of the same issue with no shared trigger words; +- proper names, code, URLs, file paths, and product terms resembling misspellings; +- identical drafts with different recipient roles and thread histories; +- firm legitimate requests that must not be weakened; +- polite but technically invalid requests; +- terse but acceptable operational messages; +- Korean, English, mixed-language, code-switched, and CJK/Unicode cases. + +These cases directly test the claim that semantic behavior is not keyword matching. + +### Metrics + +Report, at minimum: + +- issue/category precision, recall, macro-F1; +- smallest-sufficient-span accuracy and span intersection-over-union; +- replacement correctness; +- fact, intent, actor, deadline, and request-strength preservation; +- unsupported-claim and hallucinated-fact rate; +- user acceptance/ignore/rewrite rates with selection caveats; +- human inter-rater agreement; +- Brier score, expected calibration error, and reliability curves; +- category occupancy and sparse-category behavior; +- test-retest and cross-model agreement; +- DIF or group-specific error by language, role/context, recipient configuration, and thread depth; +- temporal/model/prompt/rubric drift; +- latency, token, cost, and orchestration-depth distributions. + +No overall “email danger” or “tone risk” score substitutes for the criterion results. + +## Multi-agent and compute allocation + +The workflow can allocate more test-time computation without lexical routing: + +- explicit incremental review uses a bounded candidate reviewer and independent judge; +- explicit deep review decomposes mechanics, discourse/actionability, pragmatics, and technical precision; +- calibrated disagreement, preservation conflict, missing context, or uncertainty can trigger adjudication; +- role-specific reasoning effort and workflow depth are recorded; +- ablations compare single-model routing, independent judge, multi-agent review, and adjudication. + +Speed is not the primary scientific criterion, but every workflow remains resource-bounded and observable. + +## Revision and annotation validity + +Inkspan's revision-scoped W3C `TextPositionSelector` contract supplies Unicode-code-point offsets and a strong document revision. The W3C model itself warns that position selectors are brittle when a resource changes. Naruon and Inkspan therefore reject stale diagnostics rather than searching for keywords or moving the suggestion to the nearest similar sentence. + +## Privacy and PII + +Pragmatic review may need names, organizational roles, dates, quantities, recipient relationships, and thread commitments. Blind masking can destroy the construct being measured. The system instead uses compensating controls: + +- approved provider/model/region policy; +- encrypted transport and credential storage; +- contractual no-training/no-secondary-use restrictions; +- no raw content in ordinary logs or telemetry; +- short-lived or zero-retention review sessions; +- opaque identifiers and policy/version provenance; +- synthetic or carefully de-identified persistent benchmark cases; +- explicit access, export, retention, and deletion controls. + +This design does not assert that every provider satisfies those requirements. Tenant configuration and procurement evidence determine which providers are eligible. + +## Standards and governance alignment + +- **NIST AI 600-1** supports lifecycle risk identification, evaluation, monitoring, and governance for generative AI systems. +- **ISO/IEC 23894:2023** provides AI risk-management guidance. +- **ISO/IEC 42001:2023** establishes an AI management-system framework for responsible development and use. +- **W3C Web Annotation Data Model** defines `TextPositionSelector` semantics and resource-change limitations. +- **AERA/APA/NCME Standards** inform validity, reliability, fairness, intended-use, and consequence evidence. Naruon does not claim to be a regulated psychological test merely because psychometric methods are used. + +## APA 7th references + +American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). *Standards for educational and psychological testing*. American Educational Research Association. + +Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). *Artificial intelligence risk management framework: Generative artificial intelligence profile* (NIST AI 600-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1 + +Chen, H., & Goldfarb-Tarrant, S. (2025). Safer or luckier? LLMs as safety evaluators are not robust to artifacts. In *Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)* (pp. 19750–19766). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.970 + +Fu, X., & Liu, W. (2025). How reliable is multilingual LLM-as-a-Judge? In *Findings of the Association for Computational Linguistics: EMNLP 2025*. Association for Computational Linguistics. https://aclanthology.org/2025.findings-emnlp.587/ + +International Organization for Standardization. (2023a). *Information technology—Artificial intelligence—Guidance on risk management* (ISO/IEC Standard No. 23894:2023). https://www.iso.org/standard/77304.html + +International Organization for Standardization. (2023b). *Information technology—Artificial intelligence—Management system* (ISO/IEC Standard No. 42001:2023). https://www.iso.org/standard/42001.html + +Liu, S., Xu, Z., Liu, Z., Yan, Y., Yu, M., Gu, Y., Chen, C., Xie, H., & Yu, G. (2026). Mitigating judgment preference bias in large language models through group-based polling. In *Findings of the Association for Computational Linguistics: ACL 2026* (pp. 1448–1464). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.71 + +Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG evaluation using GPT-4 with better human alignment. In *Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing* (pp. 2511–2522). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.153 + +Shen, C., Cheng, L., Nguyen, X.-P., You, Y., & Bing, L. (2023). Large language models are not yet human-level evaluators for abstractive summarization. In *Findings of the Association for Computational Linguistics: EMNLP 2023* (pp. 4215–4233). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.278 + +Usami, H., Hara, K., Tsuboi, A., & Matsuda, N. (2026). *LLM judges have dark current: A psychometric datasheet for LLM-as-a-Judge evaluation* [Preprint]. arXiv. https://arxiv.org/abs/2606.15610 + +Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Kong, L., Liu, Q., Liu, T., & Sui, Z. (2024). Large language models are not fair evaluators. In *Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)* (pp. 9440–9450). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.511 + +World Wide Web Consortium. (2017). *Web annotation data model*. https://www.w3.org/TR/annotation-model/ + +Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena* [Preprint]. arXiv. https://arxiv.org/abs/2306.05685 + +## Claim boundary + +The cited evidence supports structured, bias-aware, calibrated LLM evaluation and revision-bound editor integrity. It does not prove that the planned Naruon rubric, model, provider, language profile, or fast-mlsirm configuration is sufficiently valid. Those claims require the benchmark, ablation, DIF, drift, security, privacy, and user-consequence evidence defined in the design. \ No newline at end of file From 9a1a9489997e50b994ee35fe2a8ac11f6b625b70 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Wed, 12 Aug 2026 09:16:23 +0900 Subject: [PATCH 04/18] docs(adr): add Naruon decision index --- docs/adr/README.md | 21 +++++++++++++++++++++ 1 file changed, 21 insertions(+) create mode 100644 docs/adr/README.md diff --git a/docs/adr/README.md b/docs/adr/README.md new file mode 100644 index 000000000..6ef77b3b5 --- /dev/null +++ b/docs/adr/README.md @@ -0,0 +1,21 @@ +# Naruon Architecture Decision Records + +Status values: Proposed, Accepted, Superseded. + +This index records durable Naruon architectural decisions. A decision is not shipped behavior merely because its ADR exists. `Accepted` requires protected-branch implementation or process authority plus current verification evidence. + +| ADR | Status | Decision | +|---|---|---| +| [0001](0001-inkspan-backed-llm-email-writing-guidance.md) | Proposed | Inkspan-backed, LLM-native email writing guidance with fast-mlsirm judge calibration and no keyword semantic fallback | + +## Decision discipline + +- **Proposed:** design and ownership boundaries are documented, but implementation or operational acceptance evidence is incomplete. +- **Accepted:** protected `develop` contains the governing implementation or process and its tests, security evidence, documentation, and rollback contract are current. +- **Superseded:** retained for traceability but replaced by a later ADR. + +Material changes to model authority, editor ownership, keyword/heuristic fallback, PII handling, send gating, calibration, persistence, or cross-repository integration require a new or superseding ADR rather than silent edits. + +## Required ADR sections + +Every material ADR records context, alternatives, decision, consequences, failure and recovery, security and privacy, accessibility where applicable, compatibility and migration, verification, research/standards traceability, and rollback or supersession conditions. From ed07bf12336ded9e3692ce3b7d6af2ceacf89140 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Wed, 12 Aug 2026 09:21:52 +0900 Subject: [PATCH 05/18] docs(spec): reconcile feedback enum and tighten review contract --- ...kspan-llm-email-writing-guidance-design.md | 398 +++++++++--------- 1 file changed, 189 insertions(+), 209 deletions(-) diff --git a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md index b5cd6cd06..cf5f861fb 100644 --- a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md +++ b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md @@ -9,13 +9,13 @@ Replace Naruon's plain reply textarea and one-shot rewrite experience with a Grammarly-like, revision-safe authoring workflow: - the user writes directly in Inkspan; -- Naruon reviews the current reply against the source email, complete thread, recipients, reply purpose, and tenant policy; +- Naruon reviews the current reply against the authorized source email, complete relevant thread, recipients, reply purpose, and tenant policy; - an LLM produces passage-level candidate diagnostics; -- an independent LLM-as-a-Judge evaluates those candidates under a versioned rubric; -- deterministic code admits only strict, current-revision, selector-valid results; +- an independent LLM-as-a-Judge evaluates each candidate under a versioned rubric; +- deterministic code admits only exact-schema, current-revision, selector-valid results; - Inkspan displays, navigates, applies, ignores, and dismisses suggestions; - fast-mlsirm calibrates judge behavior and publishes versioned admission evidence; -- no keyword, regex, phrase dictionary, domain list, or positional repair acts as a semantic fallback. +- no keyword, regex, phrase dictionary, sender domain, recipient-count rule, language-name rule, or positional repair acts as a semantic fallback. The feature is writing guidance, not a send-risk score or mandatory gate. @@ -23,15 +23,15 @@ The feature is writing guidance, not a send-risk score or mandatory gate. ### Continuous authoring -The reply composer is an Inkspan editor in email-compatible HTML mode. The user can type, paste safe rich content, use formatting, undo, and send even when semantic review is unavailable. +The reply composer is an Inkspan editor in email-compatible HTML mode. The user can type, paste safe rich content, format, undo, and send even when semantic review is unavailable. ### Incremental review -After a bounded debounce, Naruon reviews the changed paragraph plus enough authorized document/thread context to interpret it. Incremental review is optimized for responsive feedback, but semantic correctness is still model-based. It does not run a local trigger-word detector before deciding what category to assign. +After a bounded debounce, Naruon reviews the changed paragraph or selection plus enough authorized thread context to interpret it. The semantic decision is model-based. No local trigger-word detector assigns a category before or after the model call. ### Deep review -The author can explicitly request `전체 메일 검토`. Deep review evaluates the complete draft for cross-sentence structure, repetition, actor ambiguity, request completeness, audience pragmatics, technical precision, and intent preservation. It may use contextual-orchestrator conduct mode with separate reviewer, critic, judge, and adjudicator roles. +The author can request `전체 메일 검토`. Deep review examines the complete draft for cross-sentence structure, repetition, actor ambiguity, request completeness, audience pragmatics, technical precision, and intent preservation. contextual-orchestrator may use separate reviewer, critic, judge, and adjudicator roles. ### Suggestions @@ -41,23 +41,23 @@ Each admitted diagnostic shows: - category and concise title; - evidence-based explanation; - optional replacement; -- confidence/admission provenance appropriate for the UI; +- bounded confidence and admission provenance appropriate for the UI; - Apply, Ignore, Dismiss, and Explain actions. Apply changes only the bound passage. Ignore and Dismiss do not alter the document. New asynchronous results do not steal focus. Stale results are visibly invalidated and cannot apply. ### Whole-document guidance -Some issues cannot be represented as one replacement range. The response may therefore also contain non-mutating document guidance, such as: +Issues without one safe replacement range may be returned as non-mutating document guidance: - inferred purpose summary; - likely reader interpretation; - unclear actor or missing deliverable; -- missing response deadline or channel; +- missing deadline or response channel; - structural reordering suggestion; -- unresolved factual/technical verification question. +- unresolved factual or technical verification question. -Document guidance is advisory and cannot mutate the editor without a separate, explicit author action. +Document guidance never mutates the editor without a separate explicit author action. ## Architecture @@ -79,41 +79,41 @@ flowchart TB M --> P ``` -### Component boundaries +## Component boundaries -#### `email_writing_api` +### `email_writing_api` -Owns authenticated HTTP contracts, tenant/workspace scope, rate limits, payload bounds, idempotency, and error mapping. +Owns authenticated HTTP contracts, tenant/workspace scope, rate limits, payload bounds, idempotency, and privacy-safe error mapping. -#### `email_writing_context_service` +### `email_writing_context_service` -Re-reads the source email and thread by opaque server-authorized identifiers. It creates a bounded context bundle with subject, selected source message, relevant prior messages, sender/recipient metadata, explicit reply objective, and current draft. Client-provided recipient roles or thread text are never authoritative. +Re-reads the source email and relevant thread by server-authorized identifiers. It creates a bounded context bundle containing the subject, selected source message, relevant prior messages, sender/recipient metadata, explicit reply objective, and current draft. Client-provided thread text or recipient roles are never authoritative. -#### `email_writing_review_service` +### `email_writing_review_service` -Builds untrusted-data-delimited prompts, calls contextual-orchestrator, validates the candidate response, invokes the independent judge, applies the published admission policy, and returns a structured review result. +Builds untrusted-data-delimited prompts, calls contextual-orchestrator, validates the candidate response, invokes the independent judge, applies the published admission policy, and returns a structured result. -#### `writing_review_judge_port` +### `writing_review_judge_port` -Defines the Naruon-side interface to an LLM judge. The preferred adapter consumes a released, hash-locked fast-mlsirm package and delegates every model call through contextual-orchestrator. The port allows testing and future service deployment without leaking fast-mlsirm internals into the API router. +Defines Naruon's interface to a judge. The preferred adapter consumes a released, hash-locked fast-mlsirm package and delegates every model call through contextual-orchestrator. The port prevents API routers from depending on fast-mlsirm internals. -#### `writing_diagnostic_validator` +### `writing_diagnostic_validator` -Performs only deterministic structural and revision validation. It does not infer semantics. +Performs deterministic structural, authorization-adjacent, revision, projection, Unicode, grapheme, selector, overlap, and safety validation. It does not infer semantics. -#### `InkspanReplyEditor` +### `InkspanReplyEditor` -A small Naruon `'use client'` boundary around the released Inkspan component. It owns editor refs, revision capture, diagnostic props, current-review state, feedback callbacks, email serialization, focus, and send integration. +A small Naruon `'use client'` boundary around the released Inkspan editor. It owns editor refs, revision capture, diagnostic props, current-review state, feedback callbacks, email serialization, focus, and send integration. -#### `judge_policy_registry` +### `judge_policy_registry` -Provides signed/versioned policy artifacts: rubric version, category anchors, accepted model/provider set, calibration scope, language profiles, thresholds or posterior admission rules, expiry, and rollback version. The first implementation may load a checked-in bounded policy artifact; long-term persistence belongs in a two-word snake_case registry object. +Provides integrity-bound policy artifacts: rubric version, category anchors, accepted model/provider set, calibration scope, language profiles, admission rules, expiry, and rollback version. Runtime consumes a published policy; it does not fit a model. ## API contracts ### `POST /api/email-writing/reviews` -Creates a bounded semantic review for the current exact editor revision. +Creates a bounded review for one exact editor revision. ```json { @@ -135,12 +135,12 @@ Creates a bounded semantic review for the current exact editor revision. Rules: -- `source_email_id` is scoped and re-read server-side. -- `document_revision`, projection, and draft must describe one exact Inkspan snapshot. -- `changed_selector` is required for incremental mode and omitted for deep mode. -- `reply_objective` is user guidance, not authority to override system policy. -- unexpected fields are rejected. -- raw email/thread data is not accepted from the browser as authoritative context. +- `source_email_id` is scoped and re-read server-side; +- revision, projection, and draft describe one exact Inkspan snapshot; +- `changed_selector` is required in incremental mode and omitted in deep mode; +- `reply_objective` is untrusted user guidance, not a policy override; +- unexpected fields are rejected; +- raw email or thread content is not accepted from the browser as authoritative context. Response: @@ -179,10 +179,7 @@ Response: "document_guidance": { "purpose_summary": "자료 범위와 일정 확인", "reader_interpretation": "여러 확인 질문과 전문성 방어가 섞여 핵심 요청이 흐려질 수 있음", - "missing_requests": [ - "수행 주체", - "회신 가능 예정일" - ], + "missing_requests": ["수행 주체", "회신 가능 예정일"], "structure_suggestion": "목적, 확인 항목, 요청 산출물, 일정 순으로 정리" }, "provenance": { @@ -195,11 +192,11 @@ Response: } ``` -The API does not expose raw candidate or judge output by default. +Raw candidate and judge output are not returned to the browser by default. ### `POST /api/email-writing/reviews/{review_session_id}/feedback` -Records an explicit author response. +Records one explicit author response. ```json { @@ -210,177 +207,172 @@ Records an explicit author response. } ``` -Allowed actions initially: +Exact initial enum values: ```text applied ignored - dismissed +dismissed requested_explanation stale conflict ``` -Implementation must normalize the accidental whitespace in documentation and use exact enum values; the canonical value is `dismissed`. - -Feedback stores opaque IDs, category, policy version, action, timing bucket, and revision references by default. It does not duplicate source text or replacement text into generic telemetry. +Feedback stores opaque IDs, category, policy version, action, timing bucket, and revision references by default. It does not duplicate source or replacement text into ordinary telemetry. ## Structured model contracts ### Candidate reviewer output -The reviewer returns an exact JSON object with: +The reviewer returns one exact JSON object with: -- `diagnostics` array; -- `document_guidance` object; -- `context_limitations` array; +- `diagnostics`; +- `document_guidance`; +- `context_limitations`; - `review_language`; -- `abstained_claims` array. +- `abstained_claims`. -Each diagnostic contains a source selector, category, explanation, proposed replacement when appropriate, and candidate confidence. Category values must come from the current rubric. The output cannot include HTML, arbitrary editor JSON, commands, or a send decision. +A diagnostic contains a source selector, category, explanation, optional replacement, and candidate confidence. It cannot contain HTML, arbitrary editor JSON, executable commands, tool requests, or a send decision. ### Independent judge input -Each candidate becomes one evaluation task. The judge sees: +Each candidate becomes one judge task containing only bounded required context: -- bounded source/thread context required for the criterion; +- source/thread evidence; - current draft and selected span; -- the claimed issue; +- claimed issue; - proposed replacement; - exact criterion descriptions and ordered category anchors; -- explicit instruction that mail/document content is untrusted data; -- no candidate model self-reported chain of thought. +- explicit instruction that all mail/document content is untrusted data; +- no candidate-model chain of thought. ### Independent judge output -Use fast-mlsirm's strict criterion-level shape and polytomous categories. The initial policy should prefer at least four ordered categories so it can distinguish unsupported, weak, adequate, and strong evidence, subject to empirical category-count ablation. +Use fast-mlsirm's strict criterion-level shape and explicit polytomous categories. The first study compares category counts rather than assuming that a larger scale is better. Duplicate keys, missing IDs, non-integral categories, extra fields, excessive depth, invalid scores, or malformed JSON fail closed. -No parser repairs a malformed response by extracting words such as “pass,” “polite,” or “incorrect.” Duplicate keys, missing criterion IDs, non-integral categories, extra fields, invalid depths, and invalid scores fail closed. +No parser extracts words such as “pass,” “polite,” or “incorrect” to repair a response. -### Admission policy +## Admission policy -An admitted diagnostic satisfies all of the following: +A diagnostic is admitted only when: -1. candidate output schema is valid; -2. source selector targets the current reviewed projection; +1. candidate schema is exact; +2. selector targets the reviewed projection; 3. every mandatory judge criterion is present; -4. the signed current policy accepts the category pattern or calibrated posterior evidence; +4. the current integrity-bound policy accepts the criterion pattern or calibrated evidence; 5. no mandatory preservation criterion falls below its floor; -6. model/rubric/language profile is within the policy's validated scope; -7. no independent-adjudication requirement remains unresolved; -8. replacement passes Inkspan-safe content policy; -9. the document revision remains current when returned/applied. +6. model, rubric, and language profile are within validated scope; +7. required adjudication is complete; +8. replacement passes Inkspan content safety; +9. document revision remains current when returned and applied. -The policy is versioned and observable. It is not a hidden weighted keyword score. +The policy is versioned and observable. It is not a weighted word list. ## Review modes and compute allocation ### Incremental mode -- input: changed paragraph/range plus bounded thread context; -- default workflow: one candidate reviewer plus one independent judge; -- deeper adjudication: triggered by calibrated disagreement, policy uncertainty, missing context, or preservation-criterion conflict; -- goal: responsive feedback without sacrificing explicit judge validation. +- changed paragraph or selection plus bounded thread context; +- one candidate reviewer and one independent judge by default; +- deeper adjudication for calibrated disagreement, uncertainty, missing context, or preservation conflict; +- bounded responsiveness without dropping judge validation. ### Deep mode -- input: full draft, selected source email, relevant thread history, recipient roles, and reply objective; -- workflow: decomposed reviewer roles for mechanics, discourse/actionability, pragmatics, and technical precision; independent judge; optional adjudicator; -- output: passage diagnostics plus whole-document guidance; -- speed is secondary to quality, but payload and step bounds remain explicit. - -### Operational mode selection +- complete draft and relevant authorized thread context; +- decomposed mechanics, discourse/actionability, pragmatics, and technical-precision roles; +- independent judge and optional adjudicator; +- passage diagnostics plus whole-document guidance; +- quality prioritized over speed, with explicit step and payload bounds. -The explicit user action chooses incremental or deep review. Document length and provider availability may alter batching. Semantic escalation depends on model/judge evidence and calibrated uncertainty, not lexical triggers. +Explicit user mode, document size, and provider health may alter batching. Semantic escalation depends on model/judge evidence and calibrated uncertainty, never lexical triggers. ## fast-mlsirm calibration plan ### Measurement unit -A criterion on a candidate diagnostic is an evaluation item. A model/provider/prompt configuration acts as a rater. Human expert decisions provide reference evidence, not assumed infallible truth. +One criterion on one candidate diagnostic is an evaluation item. Model/provider/prompt configurations act as raters or facets. Human expert judgments are reference evidence, not assumed infallible truth. ### Response matrix -For each diagnostic candidate, collect ordered category responses for multiple criteria and, where feasible, repeated raters/models/prompts. Use fast-mlsirm's response-matrix validation before fitting. +Collect ordered category responses across criteria and, where feasible, repeated models, providers, prompts, and reasoning settings. Validate every response matrix through fast-mlsirm before fitting. ### Analyses -- item/category frequency and sparse-category checks; +- category occupancy and sparse-category checks; - criterion difficulty and discrimination; -- model/rater severity and interaction where supported; +- model/rater severity and interactions where supported; - latent writing-quality and preservation dimensions; - reliability and prompt/model test-retest; -- Brier score and calibration curves for acceptance probabilities; -- differential item functioning by language, organization role, recipient configuration, thread depth, document length, and review mode; +- Brier score and calibration curves; +- DIF by language, organizational role, recipient configuration, thread depth, document length, and review mode; - temporal drift across model, prompt, rubric, and policy releases; - category-count ablation; - reasoning-effort and single-model versus multi-agent ablation; -- human consequence analysis for overcorrection and missed issues. +- consequences of overcorrection and missed issues. ### Policy publication -A calibration run emits a signed or integrity-bound `judge_policy_artifact` with: +A calibration run emits an integrity-bound `judge_policy_artifact` containing: -- policy version and creation/expiry times; +- version, creation time, expiry, and rollback version; - compatible fast-mlsirm, Naruon, Inkspan, and orchestrator contract versions; - approved model/provider/rubric/language profiles; -- category anchors; -- admission and preservation rules; +- category anchors and admission rules; - calibration and DIF summary; -- sample and dataset provenance hashes; -- known limitations and rollback policy. +- dataset and source hashes; +- known limitations. Naruon runtime consumes the artifact but does not refit it. ## Benchmark design -### Human-authored cases +### Long-lived data boundary + +Use consented, de-identified, or synthetic reconstruction cases. Do not persist real confidential mail bodies merely because Naruon can read them. -Build a consented, de-identified or synthetic reconstruction corpus covering: +### Required domains -- internal status requests; +- internal status and deadline requests; - vendor/client scope disputes; -- deadline and responsibility clarification; +- responsibility and deliverable clarification; - incident reporting; - meeting coordination; - technical review; - executive updates; -- apology/correction; +- apology and correction; - Korean, English, mixed-language, and code-switched mail. -Do not persist real confidential mail bodies merely because they are available in Naruon. +### Contrast sets proving semantic behavior -### Contrast sets proving semantic—not keyword—behavior - -1. **Same words, different meaning:** a phrase appears inside a quotation, a neutral incident transcript, and a direct interpersonal rebuke. -2. **Same issue, different words:** multiple paraphrases express public blame without sharing trigger terms. -3. **Proper name/code protection:** a suspected spelling form appears as a product name, identifier, URL, file path, quotation, or code sample. -4. **Context shift:** an identical draft is evaluated with a one-to-one peer recipient, a large executive CC list, and an external customer thread. -5. **Intent preservation:** a forceful but legitimate deadline request must not be weakened merely to sound polite. -6. **Technical precision:** a confident but unsuitable metric request receives technical guidance even without any “rude” wording. -7. **Non-issue negative controls:** terse but acceptable operational messages should not be expanded gratuitously. +1. identical words in quotation, neutral transcript, and direct rebuke; +2. the same issue paraphrased without trigger words; +3. apparent misspellings used as product names, identifiers, URLs, file paths, quotations, or code; +4. identical draft under peer, executive-CC, and external-customer contexts; +5. a firm legitimate deadline request that must not be weakened; +6. polite but technically unsuitable metric or method requests; +7. terse acceptable operational messages that must not be expanded gratuitously. ### Metrics -- issue-level and category-level precision, recall, macro-F1; -- selector span intersection-over-union and exact/smallest-sufficient-span rate; -- replacement grammaticality and correctness; +- issue/category precision, recall, macro-F1; +- selector exactness, smallest-sufficient-span rate, and span intersection-over-union; +- replacement correctness; - intent, fact, actor, deadline, request-strength, and technical-claim preservation; - unsupported-claim and hallucinated-fact rate; - accepted-suggestion precision and ignore rate; -- human inter-rater agreement and adjudicated disagreement; -- Brier score, expected calibration error, reliability diagrams; -- judge/category response validity and sparse-category behavior; -- DIF and drift evidence; -- latency, token, step, and cost distributions by mode; -- stale-result and conflict rates; +- human inter-rater agreement; +- Brier score, expected calibration error, and reliability curves; +- category occupancy, DIF, and drift; +- latency, token, step, and cost distributions; +- stale/conflict rates; - accessibility task success. -No single scalar “email risk score” is a release criterion. +No scalar “email risk score” is a release criterion. ## Data model -Phase 1 can keep review state ephemeral. If persistence is introduced, objects use two-or-more-word `snake_case` names: +Phase 1 may keep review state ephemeral. Persistent objects, if later required, use two-or-more-word `snake_case` names: ```text email_review_session @@ -393,37 +385,32 @@ review_rubric_version review_evaluation_case ``` -`email_review_session` stores opaque IDs, owner scope, source email reference, document revision, model/rubric/policy versions, status, and timestamps. Raw source email/draft text is not duplicated by default. - -`writing_diagnostic_record` is optional and encrypted if policy requires persistence. It never becomes canonical email content. +Raw source email or draft text is not duplicated by default. Persisted diagnostic content requires encryption, retention, access, and deletion policy and never becomes canonical email content. ## Privacy and PII controls -Naruon does not blindly mask names, roles, organizations, dates, quantities, or relationships required to judge pragmatics and actionability. It instead uses: +Naruon does not blindly mask names, roles, organizations, dates, quantities, or relationships required to judge pragmatics and actionability. It uses compensating controls: - tenant policy and explicit provider approval; -- encryption in transit and at rest where persisted; -- encrypted credential registry; +- encrypted transport and credential registry; - region/provider/model restrictions; - contractual no-training/no-secondary-use controls; - short-lived review context and no ordinary raw-content logging; - opaque IDs and privacy-minimized feedback telemetry; -- synthetic/de-identified long-lived evaluation corpora; -- access-controlled export and deletion procedures; -- audit events that record policy/model outcomes without copying mail text. +- synthetic or de-identified long-lived evaluation corpora; +- access-controlled export and deletion; +- audit outcomes without copied mail text. -A tenant can disable remote review or require an approved local provider. If no allowed model is available, the feature is unavailable; no lexical fallback runs. +A tenant may disable remote review or require an approved local provider. If no allowed model exists, semantic review is unavailable; no lexical fallback runs. ## Prompt-injection boundary -Every source email, quoted thread, draft, recipient display name, signature, attachment-derived text, and user reply objective is untrusted content. The prompt uses explicit delimited JSON/data blocks and tells every role that content cannot change the rubric, tool access, output schema, provider policy, or system instruction. +Source email, quoted thread, draft, display names, signatures, attachment-derived text, and reply objective are untrusted data. Prompts use delimited JSON/data blocks. Content cannot change the rubric, tool access, output schema, provider policy, or system instruction. -Candidate and judge roles receive only the minimum required context. They have no arbitrary tool access. The response parser accepts only the exact schema and rejects duplicate keys, extra fields, unsupported identifiers, excessive nesting, and unbounded strings. +Candidate and judge roles receive minimum required context and no arbitrary tool access. Parsers reject duplicate keys, extra fields, unsupported identifiers, excessive nesting, and unbounded strings. ## Error contract -Recommended HTTP/application statuses: - ```text review_completed review_abstained @@ -436,143 +423,136 @@ provider_unavailable judge_disagreement ``` -External responses do not expose raw provider errors, prompts, credentials, server URLs, or email text. Server logs retain typed error codes and trace IDs, not raw source content. +External responses omit raw provider errors, prompts, credentials, server URLs, and email text. Logs retain typed codes and trace IDs, not source content. ## Observability Allowed ordinary metrics: -- review count/status; -- mode; -- document/changed-range length bucket; -- diagnostic count/category counts; -- policy/rubric/model profile identifiers; +- review status and mode; +- document/range length bucket; +- diagnostic/category counts; +- policy/rubric/model profile IDs; - latency, token, step, retry, and cost buckets; -- admission/abstention/disagreement counts; -- feedback action counts; -- stale/conflict counts; +- admission, abstention, disagreement, feedback, stale, and conflict counts; - provider health. Disallowed by default: -- source email/draft text; -- selected span text; -- replacement text; +- email or draft text; +- selected span or replacement; - full explanation; -- prompts and raw responses; -- participant addresses/names; +- prompts or raw responses; +- participant identities; - document envelopes; - secrets. ## Accessibility -The integration must preserve Inkspan's diagnostic accessibility contract: +The integration preserves Inkspan's accessible diagnostic contract: -- keyboard navigation among suggestions; -- accessible category, count, and explanation; +- keyboard navigation; +- accessible category, count, passage, and explanation; - non-color-only range indication; -- predictable focus when opening, applying, ignoring, or closing a card; -- polite status for review completion and stale invalidation; -- no focus theft from asynchronous arrival; -- screen-reader equivalent to pointer/hover behavior; -- mobile/touch actions with sufficiently large targets. +- predictable focus after open/apply/ignore/close; +- polite completion and stale-status announcements; +- no asynchronous focus theft; +- screen-reader equivalence to pointer behavior; +- touch targets suitable for mobile use. ## Security and threat cases Tests include: -- source email instructing the model to ignore the system or approve every phrase; +- source content instructing the model to ignore policy or approve everything; - quoted JSON pretending to be judge output; -- duplicate-key and deep-nesting model responses; -- oversized thread, draft, explanation, replacement, and diagnostic arrays; -- hostile Unicode, bidi controls, combining marks, grapheme splits, HTML, links, and inline images; -- cross-tenant source email IDs; +- duplicate-key, deep-nesting, oversized, and hostile-Unicode output; +- cross-tenant source IDs and forged feedback IDs; - stale revision and selector reuse; -- feedback ID forgery; -- provider base URL SSRF/DNS-rebinding attempts under existing Naruon controls; +- unsafe HTML/link/image replacements; +- provider URL SSRF and DNS-rebinding under existing Naruon controls; - raw-content leakage through logs, metrics, errors, or traces; -- judge self-preference and candidate/judge same-model ablation; -- malicious replacement attempting unsafe markup or credential exfiltration. +- judge self-preference and candidate/judge same-model ablations; +- credential-exfiltration attempts. ## Testing layers -### Backend contract tests +### Backend -- request/response schemas and extra-field rejection; +- exact schemas and extra-field rejection; - auth and source re-read; - prompt/data delimiting; -- exact candidate and judge parser behavior; -- no keyword/positional repair; +- strict candidate and judge parsing; +- no keyword or positional repair; - policy admission and abstention; -- provider errors and stale results; -- privacy-safe errors/telemetry. +- provider failure and stale result; +- privacy-safe telemetry. -### fast-mlsirm integration tests +### fast-mlsirm integration -- released adapter import and version compatibility; +- immutable released adapter and version compatibility; - criterion/category contract; -- response-matrix conversion and validation; +- `to_irt_row()` and response-matrix validation; - policy artifact parsing; - offline calibration fixtures and deterministic seeds; -- scheduled live NIM evaluation with `NVIDIA_NIM_API_KEY`. +- scheduled live evaluation with `NVIDIA_NIM_API_KEY`. -### Frontend/Inkspan integration tests +### Frontend and Inkspan - real Inkspan component instead of textarea; - revision capture and review request; -- diagnostics rendering and action callbacks; -- stale invalidation after typing; -- Apply/Ignore/Dismiss/Explain; -- undo, focus, SSR, accessibility, email serialization, and send behavior; -- review outage does not disable editing/sending. +- diagnostic rendering and callbacks; +- invalidation after editing; +- Apply, Ignore, Dismiss, Explain, and undo; +- focus, SSR, accessibility, email serialization, and sending; +- review outage does not disable editing or sending. -### End-to-end tests +### End to end -- import/read an actual test email fixture; +- import/read a test email fixture; - author a response; -- receive model-backed diagnostics through a controlled provider fixture; -- apply one suggestion and ignore another; -- verify the serialized outgoing message and thread headers; -- assert no raw mail content enters captured logs/metrics; -- run a live-model scheduled evaluation separately from deterministic CI. +- receive controlled model-backed diagnostics; +- apply one and ignore another; +- verify outgoing serialization and thread headers; +- assert no raw content in captured logs/metrics; +- run live-model evidence separately from deterministic CI. -## Release and migration plan +## Release sequence -1. Inkspan design/ADR PR accepted. -2. Inkspan implementation plan and TDD implementation. -3. Inkspan release containing the public diagnostic contract. -4. fast-mlsirm immutable release containing the strict judge/IRT contract used by Naruon. -5. Naruon backend adapter, benchmark, and policy artifact. -6. Naruon reply editor migration to the released Inkspan package. -7. current-head CI/security/coverage/review evidence. -8. CHANGELOG/version/release evaluation. +1. accept Inkspan design/ADR; +2. write and execute Inkspan TDD plan; +3. release Inkspan diagnostic contract; +4. release the compatible strict fast-mlsirm judge/IRT contract; +5. implement Naruon backend adapter, benchmark, and policy artifact; +6. migrate Naruon reply composer to released Inkspan; +7. obtain exact-head CI, security, coverage, accessibility, package, and review evidence; +8. update CHANGELOG/version and evaluate release. -No unreleased branch or mutable source archive is a production dependency. Naruon's hash-locked dependency policy remains authoritative. +No mutable branch or source archive becomes a production dependency. ## Rollback -- disable semantic review feature flag or tenant policy; -- preserve Inkspan authoring and sending; -- keep the last safe composer value and normal draft state; -- stop emitting diagnostic requests; -- retain only bounded audit evidence required by policy; -- revert to one-shot draft generation if explicitly enabled, without claiming Grammarly-like review; -- no canonical email/database migration is required to remove ephemeral review state. +- disable semantic review by tenant/product configuration; +- preserve Inkspan authoring, drafts, and sending; +- stop review requests; +- retain only policy-required bounded audit evidence; +- optionally retain the separately described one-shot drafting path without claiming contextual inline review; +- require no canonical email or database migration to remove ephemeral state. -## Documentation and traceability updates required during implementation +## Documentation and traceability required during implementation -- Naruon ADR index and architecture; +- ADR index and architecture; - API contract; - threat model; - test strategy; -- operability and data-retention guidance; +- operability and retention; - product-event dictionary; -- Inkspan integration/version contract; -- fast-mlsirm judge policy/calibration traceability; +- Inkspan version contract; +- fast-mlsirm calibration traceability; - contextual-orchestrator workflow/role contract; - APA 7th doctoring; - CHANGELOG and release evidence. ## Approval boundary -Approval of this design authorizes only the detailed implementation plan. It does not claim the feature, model accuracy, language coverage, privacy controls, or cross-repository integration are shipped. Those claims require protected-branch implementation and exact-head evidence in every affected repository. \ No newline at end of file +Approval of this design authorizes only a detailed implementation plan. It does not claim the feature, model accuracy, language coverage, privacy controls, or cross-repository integration are shipped. Those claims require protected-branch implementation and exact-head evidence in every affected repository. From 83ee7cecf10bca5306cc87111fa40cecc6e52cf5 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Wed, 12 Aug 2026 09:41:45 +0900 Subject: [PATCH 06/18] docs(plan): add LLM email writing implementation plan --- ...m-email-writing-guidance-implementation.md | 705 ++++++++++++++++++ 1 file changed, 705 insertions(+) create mode 100644 docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md diff --git a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md new file mode 100644 index 000000000..06e2be5d6 --- /dev/null +++ b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md @@ -0,0 +1,705 @@ +# Inkspan-Based LLM Email Writing Guidance Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Replace Naruon's plain reply textarea and unquestioned one-shot rewrite with a revision-safe Inkspan composer whose passage-level writing guidance is generated by contextual-orchestrator, independently judged through fast-mlsirm, admitted by a versioned calibration policy, and explicitly accepted or rejected by the author. + +**Architecture:** Naruon re-reads the selected source email and thread under server-side authorization, then sends a bounded untrusted-data bundle to a candidate-review workflow in contextual-orchestrator. Candidate JSON is strictly parsed and each candidate is evaluated by fast-mlsirm's `ContextualOrchestratorJudge` through a separate orchestrator call. A signed or integrity-bound policy artifact admits, escalates, or abstains. Deterministic backend code validates schema, policy scope, revision, projection, selector, and safe replacement boundaries. The frontend binds the response to one exact Inkspan revision, renders Inkspan's generic diagnostic UI, and sends only explicit user actions as privacy-minimized feedback. Model fitting, DIF, drift, and policy publication remain offline or nearline and never block the synchronous mail-send path. + +**Tech Stack:** Python 3.12-3.14, FastAPI, Pydantic v2, SQLAlchemy/Alembic/PostgreSQL, httpx, contextual-orchestrator's authenticated OpenAI-compatible API, an immutable fast-mlsirm release, Next.js 16, React 19, TypeScript 6, pnpm, an immutable Inkspan npm release, Vitest, Playwright, OpenTelemetry, GitHub Actions, and scheduled NVIDIA NIM live evaluation. + +## Global Constraints + +- Every semantic claim about spelling, grammar, clarity, concision, discourse, workplace pragmatics, audience fit, technical precision, actionability, or intent preservation originates from an LLM workflow. +- Production code must not create semantic judgments from keywords, regular expressions, phrase dictionaries, sender domains, recipient counts, language names, token positions, nearest-text search, or static sentiment/tone tables. +- Regular expressions remain permitted only for deterministic transport, identifier, schema, and resource-bound validation. +- A provider failure, malformed candidate, malformed Judge result, unsupported language/profile, expired policy, unresolved disagreement, stale revision, invalid selector, or unsafe replacement produces abstention or review unavailability. It never invokes a lexical fallback. +- All semantic calls route through contextual-orchestrator or a contract-compatible provider-neutral port. Direct provider SDK calls are not added to the new review service. +- Candidate reviewer and independent Judge are separate roles. The same model/profile may be used only when the current approved policy explicitly permits it; otherwise the result is escalated or withheld. +- fast-mlsirm is the measurement/Judge layer, not the mail client, editor, source-context authority, send gate, or synchronous model-fitting service. +- Runtime consumes an immutable released fast-mlsirm package and an approved policy artifact. It does not fit IRT or multilevel models during an email review request. +- Naruon consumes an immutable released Inkspan npm package. A mutable branch, git URL, source archive, workspace path, or copied editor fork is prohibited in production dependencies. +- The browser supplies `source_email_id`, the exact Inkspan revision/projection, the current bounded draft, review mode, changed selector when applicable, language tag, and optional reply objective. It does not supply authoritative thread text, recipients, sender roles, or tenant policy. +- Source email, thread, sender, recipients, and account scope are re-read server-side through `Email.owner_filters()` and the authenticated tenant/workspace context. +- Names, roles, dates, relationships, and technical identifiers needed for valid pragmatic judgment are not destructively masked. Compensating controls are mandatory: approved provider/region, encrypted credentials, no-training/no-secondary-use terms, bounded retention, no raw-content ordinary logs, and access-controlled evaluation datasets. +- Raw email body, draft, selected passage, replacement, explanation, prompt, model answer, Judge answer, or complete orchestration trace is not stored in ordinary product telemetry. +- Review records persist opaque IDs, hashes, selector coordinates, category, policy/model/rubric versions, status, timing buckets, and explicit feedback only. +- Diagnostic guidance is advisory. Editing and sending remain available when review is disabled, unavailable, stale, rejected, or incomplete. +- Remaining diagnostics do not implicitly block send. Any future send policy requires a separate ADR, explicit authorization, user-visible rule, and independent acceptance evidence. +- Initial validated product claims are limited to language/model/rubric profiles that pass the benchmark and DIF gates. Architecture support does not constitute a language-validity claim. +- The first production policy uses ordered polytomous criterion categories; the category count is selected by empirical ablation rather than convenience. +- Scheduled live tests use `NVIDIA_NIM_API_KEY`. `COPILOT_GITHUB_TOKEN` is never used for model calls. +- Python 3.14 is a required CI and package-compatibility lane. +- Production statement and branch coverage remains 100%; public backend APIs and shipped Python modules/classes/functions receive complete docstrings. +- Database objects use two-or-more-word `snake_case` names. +- Feature work remains under `Unreleased`; version and release publication occur only after exact-head product, calibration, security, review, and rollback evidence exists. + +--- + +## Task 1: Establish Immutable Cross-Repository Dependency Gates + +**Files:** +- Modify: `backend/pyproject.toml` +- Modify: `backend/requirements.txt` +- Modify: `backend/requirements-hashes.txt` +- Create: `backend/tests/test_email_writing_dependency_contract.py` +- Modify: `frontend/package.json` +- Modify: `frontend/pnpm-lock.yaml` +- Create: `frontend/src/lib/email-writing-dependency-contract.test.ts` +- Modify: `docs/TRACEABILITY.md` + +- [ ] Wait for Inkspan's writing-diagnostics implementation and release-only PR to publish one immutable npm version, tarball integrity, source commit, and package manifest. +- [ ] Wait for fast-mlsirm to publish one immutable package version containing `ContextualOrchestratorJudge`, `JudgeCriterion`, `JudgeFormatError`, `LLMJudgeResult`, `validate_irt_response_matrix`, and the polytomous response contract. +- [ ] Record both exact released versions and integrity/source identifiers literally in `docs/TRACEABILITY.md` before adding runtime imports. +- [ ] Add fast-mlsirm as an exact hashed backend dependency and regenerate the repository's canonical requirements hashes under Python 3.12 and 3.14 compatibility checks. +- [ ] Add `@contextualwisdomlab/cwl-editor` as an exact npm dependency and commit the canonical pnpm lockfile. +- [ ] Do not add contextual-orchestrator as an in-process mutable source dependency. Naruon integrates through its authenticated versioned HTTP contract and a local port. +- [ ] Add tests that fail for git URLs, branch names, workspace links, local paths, source archives, missing hashes, incompatible versions, or a package whose public symbols differ from the recorded contract. +- [ ] Add a packed npm smoke fixture that imports the Inkspan root, styles, and `writing-diagnostics` subpath from Naruon's actual lockfile installation. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_dependency_contract.py + +cd ../frontend +pnpm exec vitest run src/lib/email-writing-dependency-contract.test.ts +pnpm typecheck +``` + +- [ ] Commit: + +```bash +git add backend/pyproject.toml backend/requirements.txt backend/requirements-hashes.txt backend/tests/test_email_writing_dependency_contract.py frontend/package.json frontend/pnpm-lock.yaml frontend/src/lib/email-writing-dependency-contract.test.ts docs/TRACEABILITY.md +git commit -m "build(email-writing): pin Inkspan and Judge contracts" +``` + +## Task 2: Define Strict Review, Diagnostic, Guidance, and Feedback Contracts + +**Files:** +- Create: `backend/services/email_writing_contracts.py` +- Create: `backend/tests/test_email_writing_contracts.py` +- Create: `frontend/src/lib/email-writing.ts` +- Create: `frontend/src/lib/email-writing.test.ts` + +- [ ] Write failing Pydantic tests for exact fields, extra-field rejection, invalid types, non-integral selector offsets, unsupported projection, invalid BCP 47 tags, incremental/deep conditional fields, oversized draft/objective, malformed revisions, invalid statuses, duplicate diagnostic IDs, and unsafe enum values. +- [ ] Define `EmailWritingReviewRequest` with: + - `source_email_id`; + - strong `document_revision`; + - literal projection name `inkspan-prosemirror-text` and version `1`; + - bounded `draft_plain_text`; + - `language_tag`; + - `review_mode` of `incremental` or `deep`; + - required `changed_selector` only for incremental mode; + - optional bounded `reply_objective`. +- [ ] Define exact response models for `review_session_id`, reviewed revision/projection, `review_status`, diagnostics, document guidance, context limitations, abstained claims, and redacted provenance. +- [ ] Define feedback actions exactly as: + +```text +applied +ignored +dismissed +requested_explanation +stale +conflict +``` + +- [ ] Define stable application error codes without provider messages, prompt fragments, URLs, credentials, source text, or raw model output. +- [ ] Keep API JSON in `snake_case`; add explicit frontend adapters to Inkspan's camelCase diagnostic types rather than leaking transport naming into the editor package. +- [ ] Reject non-finite numeric values, duplicate JSON object keys, excessive nesting, oversized arrays, invalid Unicode boundaries, and unexpected fields. +- [ ] Add frontend runtime parsers that treat every API response as unknown input and fail closed before passing diagnostics to Inkspan. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_contracts.py + +cd ../frontend +pnpm exec vitest run src/lib/email-writing.test.ts +pnpm typecheck +``` + +- [ ] Commit: + +```bash +git add backend/services/email_writing_contracts.py backend/tests/test_email_writing_contracts.py frontend/src/lib/email-writing.ts frontend/src/lib/email-writing.test.ts +git commit -m "feat(email-writing): define strict review contracts" +``` + +## Task 3: Add Privacy-Minimized Review Evidence Tables + +**Files:** +- Modify: `backend/db/models.py` +- Create: `backend/alembic/versions/20260812_0001_add_email_writing_review_evidence.py` +- Create: `backend/tests/test_email_writing_models.py` +- Create: `backend/tests/test_email_writing_migration.py` + +- [ ] Write failing model and migration tests before adding tables. +- [ ] Add `email_review_session` with generated UUID, owner scope, source email reference, reviewed document revision/projection, review mode, language profile, status, workflow/model/rubric/policy versions, prompt hash, timing/cost buckets, creation/expiry timestamps, and no raw source/draft content. +- [ ] Add `writing_diagnostic_record` with session reference, opaque diagnostic ID, category, priority, selector start/end, candidate hash, replacement hash, explanation hash, criterion-category JSON, Judge score, admission status/reason, and no replacement/explanation text. +- [ ] Add `diagnostic_feedback_event` with diagnostic reference, owner scope, action, reviewed/resulting revisions, conflict/stale reason, event time, and idempotency key. +- [ ] Use two-or-more-word `snake_case` for every table, column, index, constraint, enum/check identifier, and relationship name. +- [ ] Add unique constraints for one diagnostic ID per session and one idempotent feedback event per user action key. +- [ ] Add owner-scope indexes and bounded retention fields so an operator can delete expired evidence without scanning raw mail data. +- [ ] Add database checks for selector order, valid status/action enums, non-negative timing buckets, and scores in `[0, 1]`. +- [ ] Prove migration upgrade/downgrade on PostgreSQL and SQLite test paths used by the repository, without dropping unrelated objects. +- [ ] Prove ORM serialization and logs never include source email body, draft, replacement, explanation, prompt, raw output, provider token, or complete trace. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_models.py tests/test_email_writing_migration.py +``` + +- [ ] Commit: + +```bash +git add backend/db/models.py backend/alembic/versions/20260812_0001_add_email_writing_review_evidence.py backend/tests/test_email_writing_models.py backend/tests/test_email_writing_migration.py +git commit -m "feat(email-writing): persist privacy-minimized review evidence" +``` + +## Task 4: Build the Server-Authoritative Email and Thread Context Service + +**Files:** +- Create: `backend/services/email_writing_context_service.py` +- Create: `backend/tests/test_email_writing_context_service.py` + +- [ ] Write failing tests for valid ownership, cross-user access, cross-organization access, missing email, deleted email, forged browser recipient data, malformed thread IDs, thread ordering, duplicate messages, oversized threads, quotations, signatures, and recipient-role derivation. +- [ ] Re-read the selected `Email` with `Email.owner_filters(auth_context.user_id, auth_context.organization_id)`. +- [ ] Re-read relevant thread messages through canonical thread keys and the same owner scope; never trust browser-supplied thread text or participants. +- [ ] Produce one immutable `EmailWritingContextBundle` containing bounded subject, selected source message, relevant chronological messages, sender/reply-to/recipient metadata, source timestamps, explicit reply objective, current draft, and declared language tag. +- [ ] Apply deterministic content bounds by complete semantic unit: do not silently cut a UTF-8 code point, grapheme cluster, JSON string, or selected source span. +- [ ] Preserve quoted material and sender attribution needed for interpretation while marking every field as untrusted content in the prompt bundle. +- [ ] Record context limitations when older thread messages are omitted by a documented budget; do not present a truncated context as complete. +- [ ] Do not run a keyword selector to decide which messages, roles, or categories matter. Selection uses server-authoritative thread membership, chronology, explicit review mode, and bounded recency/size policy. +- [ ] Return typed `context_insufficient` or authorization errors without exposing whether another tenant owns a referenced email. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_context_service.py +``` + +- [ ] Commit: + +```bash +git add backend/services/email_writing_context_service.py backend/tests/test_email_writing_context_service.py +git commit -m "feat(email-writing): build authorized thread context" +``` + +## Task 5: Implement the Authenticated contextual-orchestrator Port + +**Files:** +- Create: `backend/services/contextual_orchestrator_client.py` +- Create: `backend/services/email_writing_orchestrator_port.py` +- Create: `backend/tests/test_contextual_orchestrator_client.py` +- Modify: `backend/services/tenant_config_scope.py` +- Modify: `backend/api/tenant_config.py` +- Modify: `backend/tests/test_tenant_config_scope.py` + +- [ ] Define a narrow port with asynchronous candidate completion and a synchronous `complete(messages, mode=...)` shape compatible with fast-mlsirm's Judge. +- [ ] Store the orchestrator endpoint/profile reference and encrypted inference credential through the existing tenant-scoped configuration boundary; never accept an arbitrary endpoint or bearer token from the review request. +- [ ] Validate the configured endpoint with the repository's outbound URL/SSRF policy: HTTPS only, explicit operator allowlist, no loopback/private/link-local/reserved/multicast destination, DNS re-resolution checks, fixed request path, and no redirect to an unapproved host. +- [ ] POST only to `/v1/chat/completions` with exact allowed fields, bounded messages, `mode` of `route` or `conduct`, and `include_orchestration_trace` enabled only for the trusted internal call. +- [ ] Normalize the OpenAI-compatible response into: + +```python +{ + "answer": "strict model output", + "mode": "route | conduct", + "trace": [ + {"usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0}} + ], +} +``` + +- [ ] Retain only redacted token/step/model/profile evidence. Do not persist trace messages, prompts, answers, URLs, or credentials. +- [ ] Enforce bounded connect/read/write/pool timeouts, retry only transient errors, apply circuit breaking, and propagate stable typed outcomes for unauthorized, rate-limited, saturated, unavailable, malformed, and policy-rejected responses. +- [ ] Run fast-mlsirm's synchronous Judge call inside a capacity-limited worker lane rather than blocking the FastAPI event loop or using an unbounded default executor. +- [ ] Add mock-server tests for redirect attacks, DNS rebinding, oversized bodies, duplicate JSON keys, missing choices, invalid mode, leaked provider details, retry classification, cancellation, and client close. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_contextual_orchestrator_client.py tests/test_tenant_config_scope.py +``` + +- [ ] Commit: + +```bash +git add backend/services/contextual_orchestrator_client.py backend/services/email_writing_orchestrator_port.py backend/tests/test_contextual_orchestrator_client.py backend/services/tenant_config_scope.py backend/api/tenant_config.py backend/tests/test_tenant_config_scope.py +git commit -m "feat(email-writing): add secure orchestration port" +``` + +## Task 6: Create the Candidate Reviewer Prompt and Strict Output Parser + +**Files:** +- Create: `backend/services/email_writing_candidate_review.py` +- Create: `backend/services/email_writing_prompt.py` +- Create: `backend/tests/test_email_writing_candidate_review.py` +- Create: `backend/tests/fixtures/email_writing/candidate_outputs.json` + +- [ ] Write failing tests for exact valid JSON, duplicate keys, markdown fences, surrounding prose, extra/missing fields, invalid categories, invalid confidence, overlapping/empty/out-of-range selectors, unsafe replacements, excessive nesting, and prompt injection inside every input field. +- [ ] Define the exact candidate output object: + +```text +diagnostics +document_guidance +context_limitations +review_language +abstained_claims +``` + +- [ ] Define each candidate diagnostic with selector, category, priority, title, explanation, optional plain-text replacement, candidate confidence, and candidate evidence IDs. It must not contain HTML, editor JSON, commands, a send decision, tool request, or chain-of-thought. +- [ ] Delimit the complete context bundle as canonical JSON under explicit untrusted-data markers. Mail, thread, recipient display names, signatures, quoted instructions, attachment-derived text, draft, and reply objective cannot modify the system rubric or output schema. +- [ ] Candidate categories come from the versioned review rubric, not runtime keyword matching. +- [ ] Require the reviewer to abstain from unsupported factual/technical claims and list the missing context. +- [ ] Validate source selectors against the submitted draft's Unicode-code-point projection before any Judge call. +- [ ] Permit whole-document guidance only in the non-mutating guidance object. It cannot carry a replacement transaction. +- [ ] Hash the normalized prompt template and candidate payload for provenance without logging either plaintext value. +- [ ] Add same-words/different-context and same-issue/different-words fixtures; parser behavior must be identical because semantics come from model output, not the parser. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_candidate_review.py +``` + +- [ ] Commit: + +```bash +git add backend/services/email_writing_candidate_review.py backend/services/email_writing_prompt.py backend/tests/test_email_writing_candidate_review.py backend/tests/fixtures/email_writing/candidate_outputs.json +git commit -m "feat(email-writing): parse contextual review candidates" +``` + +## Task 7: Integrate fast-mlsirm's Independent Criterion Judge + +**Files:** +- Create: `backend/services/email_writing_judge.py` +- Create: `backend/tests/test_email_writing_judge.py` +- Create: `backend/tests/fixtures/email_writing/judge_outputs.json` + +- [ ] Import the released `ContextualOrchestratorJudge`, `JudgeCriterion`, `JudgeFormatError`, `LLMJudgeResult`, and response-matrix validator from fast-mlsirm. +- [ ] Define independently observable criteria with two-or-more-word snake_case IDs: + - `issue_support`; + - `span_fidelity`; + - `replacement_correctness`; + - `intent_preservation`; + - `fact_preservation`; + - `request_strength_preservation`; + - `audience_pragmatics`; + - `technical_precision`; + - `actionability_support`; + - `explanation_quality`. +- [ ] Make the required criterion subset depend on candidate type without changing IDs or anchors. A no-replacement diagnostic does not fabricate `replacement_correctness`; the policy schema declares which criteria are mandatory for each candidate kind. +- [ ] Use explicit ordered categories and exact anchors. The initial category count remains an evaluation parameter until Task 14 ablation selects and publishes an approved value. +- [ ] Pass bounded task, candidate answer, reference context, and rubric as untrusted data to fast-mlsirm; never parse free-form `pass`, `polite`, or `correct` tokens. +- [ ] Derive final acceptance from fast-mlsirm's validated criterion result plus the external policy. Do not trust candidate confidence or the Judge's advisory boolean. +- [ ] Add tests for duplicate JSON keys, non-integral categories, reversed anchors, missing criterion IDs, extra criterion IDs, NaN/infinity, score/category disagreement, same-model policy violation, worker-lane saturation, cancellation, and redacted errors. +- [ ] Convert repeated Judge results into response rows and validate the matrix before any calibration export. +- [ ] Verify no Judge raw output, task text, answer text, reference text, or full trace reaches ordinary logs or persistent product tables. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_judge.py +``` + +- [ ] Commit: + +```bash +git add backend/services/email_writing_judge.py backend/tests/test_email_writing_judge.py backend/tests/fixtures/email_writing/judge_outputs.json +git commit -m "feat(email-writing): add independent fast-mlsirm Judge" +``` + +## Task 8: Implement the Versioned Judge Policy Registry and Admission Gate + +**Files:** +- Create: `backend/policies/email_writing_judge_policy.schema.json` +- Create: `backend/policies/email_writing_judge_evaluation_only_v1.json` +- Create: `backend/policies/email_writing_policy_manifest.json` +- Create: `backend/services/email_writing_policy.py` +- Create: `backend/tests/test_email_writing_policy.py` + +- [ ] Define a strict JSON Schema for policy ID/version, status, creation/expiry, compatible Naruon/Inkspan/fast-mlsirm/orchestrator contracts, approved model/provider/rubric/language profiles, category count/anchors, mandatory criterion floors, posterior/score admission rules, adjudication conditions, calibration/DIF summary, dataset/provenance hashes, limitations, and rollback version. +- [ ] Check in an integrity-bound `evaluation_only` policy that can exercise the pipeline but cannot emit user-facing diagnostics in production. +- [ ] Add a manifest containing the literal SHA-256 of every allowed policy artifact; reject modified, unknown, expired, future-dated, revoked, incompatible, or non-approved artifacts. +- [ ] Implement outcomes: + +```text +admit +withhold +adjudicate +unsupported_profile +policy_unavailable +``` + +- [ ] Require every mandatory preservation criterion to meet its published floor. A high weighted average cannot compensate for failed fact, intent, actor, deadline, or request-strength preservation. +- [ ] Require an approved profile for language, model/provider pair, candidate/Judge separation, review mode, and rubric version. +- [ ] Do not hard-code unvalidated production thresholds in Python. Approved threshold values enter only through a Task 14 policy artifact with calibration evidence. +- [ ] Add rollback tests proving a revoked current policy can atomically fall back only to a compatible, non-expired, explicitly listed rollback policy; otherwise semantic review disables. +- [ ] Add policy tests for field smuggling, signature/hash mismatch, downgrade, category-anchor reversal, impossible floors, hidden keyword tables, and unexpected executable content. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_policy.py +``` + +- [ ] Commit: + +```bash +git add backend/policies/email_writing_judge_policy.schema.json backend/policies/email_writing_judge_evaluation_only_v1.json backend/policies/email_writing_policy_manifest.json backend/services/email_writing_policy.py backend/tests/test_email_writing_policy.py +git commit -m "feat(email-writing): add calibrated admission policy boundary" +``` + +## Task 9: Compose the Review Service with Bounded Compute and Fail-Closed Recovery + +**Files:** +- Create: `backend/services/email_writing_review_service.py` +- Create: `backend/tests/test_email_writing_review_service.py` + +- [ ] Write failing service tests for incremental review, deep review, no candidates, multiple candidates, invalid candidate, Judge rejection, adjudication, unsupported profile, stale request metadata, provider outage, partial Judge failure, cancellation, timeouts, and persistence rollback. +- [ ] Orchestrate this exact order: + +```text +context build +-> candidate reviewer +-> strict candidate validation +-> independent criterion Judge +-> policy admission/adjudication +-> deterministic selector/replacement validation +-> privacy-minimized evidence transaction +-> response +``` + +- [ ] Incremental mode reviews the declared changed selector plus bounded authorized thread context; deep mode reviews the complete bounded draft and may use contextual-orchestrator `conduct` mode with decomposed reviewer roles. +- [ ] Semantic escalation is based on structured Judge disagreement, policy uncertainty, missing context, preservation-criterion conflict, or approved compute policy. It is never triggered by lexical words or phrase counts. +- [ ] Bound candidate count, Judge calls, concurrency, context bytes, draft bytes, orchestration steps, retries, total wall time, and cost/tokens by policy. +- [ ] Use one database transaction for review session and admitted/withheld diagnostic metadata; roll back on persistence failure without returning unrecorded diagnostics. +- [ ] Return no raw candidate or Judge output. Return admitted diagnostics, non-mutating document guidance, limitations, abstentions, and redacted provenance only. +- [ ] If one candidate fails, apply the policy's explicit atomicity rule. Do not silently drop failed candidates while claiming a complete review. +- [ ] Ensure review cancellation closes provider requests, releases capacity limiters, and leaves a typed cancelled/abstained session rather than a running record. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_review_service.py +``` + +- [ ] Commit: + +```bash +git add backend/services/email_writing_review_service.py backend/tests/test_email_writing_review_service.py +git commit -m "feat(email-writing): orchestrate judged writing reviews" +``` + +## Task 10: Add Authenticated Review and Feedback APIs + +**Files:** +- Create: `backend/api/email_writing.py` +- Create: `backend/tests/test_email_writing_api.py` +- Modify: `backend/main.py` +- Modify: `backend/tests/test_main.py` + +- [ ] Write failing API tests for authentication, owner scope, exact request/response schema, CSRF/origin middleware, no-store headers, idempotency, review statuses, stale/conflict feedback, forged session/diagnostic IDs, cross-tenant feedback, provider errors, rate limits, and error redaction. +- [ ] Add `POST /api/email-writing/reviews` with an `Idempotency-Key` header bound to owner, source email, revision, review mode, and request hash. +- [ ] Add `POST /api/email-writing/reviews/{review_session_id}/feedback` with exact action and revision contracts. +- [ ] Re-read session/diagnostic ownership before feedback; never trust browser category, selector, policy, or provenance values. +- [ ] Reject feedback for expired/revoked sessions, mismatched document revision, duplicate incompatible idempotency keys, or nonexistent diagnostic IDs. +- [ ] Map outcomes to stable HTTP/application responses: + +```text +200 review_completed +200 review_abstained +409 review_stale +409 judge_disagreement +422 review_rejected +422 context_insufficient +503 policy_unavailable +503 provider_unavailable +``` + +- [ ] Preserve editing and send APIs independently; the review API cannot mutate or send email. +- [ ] Include the router under the existing private API dependency boundary in `backend/main.py`. +- [ ] Add OpenAPI examples containing synthetic text only. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_api.py tests/test_main.py +``` + +- [ ] Commit: + +```bash +git add backend/api/email_writing.py backend/tests/test_email_writing_api.py backend/main.py backend/tests/test_main.py +git commit -m "feat(email-writing): expose review and feedback APIs" +``` + +## Task 11: Extract the Reply Composer and Introduce the Released Inkspan Editor + +**Files:** +- Create: `frontend/src/components/email/InkspanReplyEditor.tsx` +- Create: `frontend/src/components/email/InkspanReplyEditor.test.tsx` +- Create: `frontend/src/components/email/EmailReplyComposer.tsx` +- Create: `frontend/src/components/email/EmailReplyComposer.test.tsx` +- Modify: `frontend/src/components/EmailDetail.tsx` +- Modify: `frontend/src/components/EmailDetail.test.tsx` +- Modify: `frontend/src/app/layout.tsx` +- Modify: `frontend/src/app/layout.test.tsx` + +- [ ] Write failing tests that prove the existing `Textarea` is removed from the reply workflow and a real released Inkspan component is rendered. +- [ ] Import Inkspan's stylesheet once from the root layout and add a layout regression preventing duplicate or conditional style imports. +- [ ] Create a small `'use client'` `InkspanReplyEditor` adapter that owns the editor ref, current envelope/revision capture, canonical plain-text projection, diagnostics props, action callbacks, read-only/review state, and email serialization. +- [ ] Extract reply-specific state and controls from the oversized `EmailDetail.tsx` into `EmailReplyComposer`; leave mail/thread/synthesis responsibilities in `EmailDetail`. +- [ ] Preserve the existing one-shot `/api/llm/draft` action during migration, but insert its result as an explicit generated proposal into Inkspan rather than replacing the editor invisibly. +- [ ] Keep manual authoring available before, during, and after review outages. +- [ ] Preserve current send semantics first: derive a deterministic plain-text body from the exact Inkspan document and pass it through the existing `buildReplyPayload()` contract. Rich HTML sending is a separate reviewed slice because the current backend `EmailMessageParams` is plain-text only. +- [ ] Preserve reply recipient, subject, `In-Reply-To`, `References`, simulation status, event IDs, and send error behavior. +- [ ] Add tests for typing, generated proposal insertion, clear, undo, disabled send on empty draft, send during review outage, focus, SSR/hydration, and switching selected emails while async work is pending. +- [ ] Run: + +```bash +cd frontend +pnpm exec vitest run src/components/email/InkspanReplyEditor.test.tsx src/components/email/EmailReplyComposer.test.tsx src/components/EmailDetail.test.tsx src/app/layout.test.tsx +pnpm typecheck +``` + +- [ ] Commit: + +```bash +git add frontend/src/components/email/InkspanReplyEditor.tsx frontend/src/components/email/InkspanReplyEditor.test.tsx frontend/src/components/email/EmailReplyComposer.tsx frontend/src/components/email/EmailReplyComposer.test.tsx frontend/src/components/EmailDetail.tsx frontend/src/components/EmailDetail.test.tsx frontend/src/app/layout.tsx frontend/src/app/layout.test.tsx +git commit -m "feat(email-writing): replace reply textarea with Inkspan" +``` + +## Task 12: Add Revision-Safe Incremental and Deep Review Control + +**Files:** +- Create: `frontend/src/hooks/useEmailWritingReview.ts` +- Create: `frontend/src/hooks/useEmailWritingReview.test.tsx` +- Modify: `frontend/src/components/email/EmailReplyComposer.tsx` +- Modify: `frontend/src/components/email/EmailReplyComposer.test.tsx` + +- [ ] Write failing hook tests for bounded debounce, exact revision capture, changed selector, deep review, aborted requests, out-of-order responses, source email switching, stale results, policy unavailable, Judge disagreement, retry, and unmount. +- [ ] For incremental review, wait for a bounded idle interval, capture one immutable Inkspan snapshot/revision/projection, and send the changed paragraph selector plus authorized source email ID. +- [ ] Do not run a local keyword detector to decide whether or what to review. +- [ ] For deep review, require explicit `전체 메일 검토` activation and send the complete bounded draft snapshot with no changed selector. +- [ ] Use `AbortController`, monotonic request IDs, email ID, editor identity, and document revision together. A response is accepted only when all still match. +- [ ] Pass only runtime-validated API diagnostics through the transport-to-Inkspan adapter. +- [ ] When the user edits after a response, let Inkspan invalidate diagnostics immediately and record one stale feedback event without blocking further typing. +- [ ] Display review states distinctly: idle, scheduled, reviewing, completed, abstained, stale, unavailable, disagreement, and rejected. +- [ ] Do not convert `unavailable`, `abstained`, or an empty admitted list into “문제가 없음.” Use explicit evidence language. +- [ ] Add manual `다시 검토` and `전체 메일 검토` controls with keyboard and screen-reader support; asynchronous completion must not steal focus. +- [ ] Run: + +```bash +cd frontend +pnpm exec vitest run src/hooks/useEmailWritingReview.test.tsx src/components/email/EmailReplyComposer.test.tsx +pnpm typecheck +``` + +- [ ] Commit: + +```bash +git add frontend/src/hooks/useEmailWritingReview.ts frontend/src/hooks/useEmailWritingReview.test.tsx frontend/src/components/email/EmailReplyComposer.tsx frontend/src/components/email/EmailReplyComposer.test.tsx +git commit -m "feat(email-writing): add revision-safe review control" +``` + +## Task 13: Wire Explicit Diagnostic Actions, Guidance, and Privacy-Safe Product Events + +**Files:** +- Modify: `frontend/src/components/email/InkspanReplyEditor.tsx` +- Modify: `frontend/src/components/email/EmailReplyComposer.tsx` +- Create: `frontend/src/components/email/EmailWritingGuidance.tsx` +- Create: `frontend/src/components/email/EmailWritingGuidance.test.tsx` +- Modify: `frontend/src/lib/product-events.ts` +- Modify: `frontend/src/lib/product-events.test.ts` +- Modify: `backend/services/email_writing_review_service.py` +- Modify: `backend/tests/test_email_writing_review_service.py` + +- [ ] Map Inkspan Apply/Ignore/Dismiss/Explain/Stale/Conflict callbacks to the feedback API with exact reviewed and resulting revisions. +- [ ] Do not optimistically claim feedback persistence. Show success only after the server accepts the event; preserve the local document result if telemetry fails and expose a retryable non-blocking status. +- [ ] Render whole-document guidance outside the editor as non-mutating content: purpose summary, likely reader interpretation, missing requests, structure suggestion, context limitations, and unsupported claims. +- [ ] Never add a one-click whole-document replacement in the first release. +- [ ] Add product events for review requested/completed/abstained/stale, diagnostic viewed/applied/ignored/dismissed/explanation requested, feedback persistence, and review outage. +- [ ] Event payloads may contain opaque review/diagnostic IDs, category, mode, policy/rubric/model profile, count, latency/cost buckets, and action. They must not contain source/draft/span/replacement/explanation text, prompt, raw output, participant data, or document envelopes. +- [ ] Add source tests that reject telemetry keys associated with raw content and dynamic event payload tests that inspect captured events. +- [ ] Ensure an Apply action preserves author intent/fact/request strength according to the admitted Judge evidence but remains undoable and explicitly user-authorized. +- [ ] Run: + +```bash +cd frontend +pnpm exec vitest run src/components/email/EmailWritingGuidance.test.tsx src/lib/product-events.test.ts + +cd ../backend +python -m pytest -q tests/test_email_writing_review_service.py +``` + +- [ ] Commit: + +```bash +git add frontend/src/components/email/InkspanReplyEditor.tsx frontend/src/components/email/EmailReplyComposer.tsx frontend/src/components/email/EmailWritingGuidance.tsx frontend/src/components/email/EmailWritingGuidance.test.tsx frontend/src/lib/product-events.ts frontend/src/lib/product-events.test.ts backend/services/email_writing_review_service.py backend/tests/test_email_writing_review_service.py +git commit -m "feat(email-writing): connect guidance actions and evidence" +``` + +## Task 14: Build the Human-Grounded Contrast Benchmark and Publish the First Approved Policy + +**Files:** +- Create: `evaluation/email_writing/README.md` +- Create: `evaluation/email_writing/case_schema.json` +- Create: `evaluation/email_writing/synthetic_contrast_cases.jsonl` +- Create: `evaluation/email_writing/run_evaluation.py` +- Create: `evaluation/email_writing/analyze_judge_responses.py` +- Create: `backend/tests/test_email_writing_evaluation_contract.py` +- Create: `.github/workflows/email-writing-live-evaluation.yml` +- Modify: `backend/policies/email_writing_policy_manifest.json` +- Create after passing evidence: `backend/policies/email_writing_judge_approved_v1.json` + +- [ ] Define a consented/de-identified or synthetic case schema with source context, recipient roles, draft, gold spans, categories, acceptable replacements, preservation constraints, annotator IDs, adjudication, and consequence severity. +- [ ] Include contrast families: + - same words in quotation, neutral incident report, and direct rebuke; + - same pragmatic issue expressed through unrelated paraphrases; + - product names, identifiers, URLs, file paths, code, and quotations that resemble spelling errors; + - identical draft under peer, executive-CC, and external-customer contexts; + - legitimate firm deadline/request that must not be softened; + - technically unsuitable request with no rude wording; + - terse but acceptable negative controls; + - Korean, English, mixed-language, CJK, emoji, combining-mark, and hostile-Unicode cases. +- [ ] Collect independent human annotations and report inter-rater agreement plus adjudicated disagreement. Human labels are reference evidence, not unquestioned truth. +- [ ] Measure issue/category precision, recall, macro-F1, span IoU/exactness, replacement correctness, intent/fact/actor/deadline/request-strength preservation, unsupported-claim rate, accepted-suggestion precision, Brier score, calibration error, and human ignore rate. +- [ ] Use fast-mlsirm response matrices to analyze criterion/category behavior, item/rater effects where supported, reliability, prompt/model test-retest, category sparsity, category-count ablation, and DIF by language, recipient configuration, role/hierarchy, thread depth, document length, review mode, model, prompt, and time. +- [ ] Compare single-model, separated candidate/Judge, and deeper multi-agent/adjudicator workflows with role-specific reasoning effort. Speed is recorded but does not override validity. +- [ ] Use deterministic offline fixtures in required CI. Run live scheduled/manual evaluation with `NVIDIA_NIM_API_KEY`; do not expose the secret to pull-request code from untrusted forks and do not use `COPILOT_GITHUB_TOKEN`. +- [ ] Fail the policy publication job when sample size, calibration, preservation, DIF, drift, or consequence thresholds are not met. Do not publish a policy merely because average accuracy is high. +- [ ] Generate the approved policy artifact from measured outputs, review its literal values, add its SHA-256 to the manifest, and prove production rejects the earlier `evaluation_only` artifact for user-facing output. +- [ ] Store reports, response matrices, aggregate metrics, manifests, and source hashes without confidential raw company email. +- [ ] Run: + +```bash +python evaluation/email_writing/run_evaluation.py --fixture evaluation/email_writing/synthetic_contrast_cases.jsonl --offline +python evaluation/email_writing/analyze_judge_responses.py --input evaluation/email_writing/results/offline.jsonl +cd backend +python -m pytest -q tests/test_email_writing_evaluation_contract.py tests/test_email_writing_policy.py +``` + +- [ ] Commit: + +```bash +git add evaluation/email_writing .github/workflows/email-writing-live-evaluation.yml backend/tests/test_email_writing_evaluation_contract.py backend/policies/email_writing_policy_manifest.json backend/policies/email_writing_judge_approved_v1.json +git commit -m "test(email-writing): publish calibrated Judge policy evidence" +``` + +## Task 15: Add End-to-End, Security, Accessibility, and Degraded-Mode Verification + +**Files:** +- Create: `frontend/e2e/email-writing-guidance.spec.ts` +- Create: `backend/tests/test_email_writing_security.py` +- Create: `frontend/src/components/email/emailWritingAccessibility.test.tsx` +- Modify: `frontend/scripts/full-product-ui-smoke.mjs` +- Modify: `frontend/scripts/pilot-ui-smoke.mjs` +- Modify: `docs/TEST_STRATEGY.md` + +- [ ] Add an end-to-end fixture that imports/reads a test email, authors a reply in Inkspan, receives controlled model-backed candidates, admits one diagnostic through the Judge/policy fixture, applies one, ignores another, sends the final plain-text reply, and verifies thread headers. +- [ ] Capture logs/metrics/traces in the test and assert no raw source email, draft, selected span, replacement, explanation, prompt, raw output, participant address, or credential appears. +- [ ] Add prompt-injection cases from source mail, quoted thread, display names, signatures, draft, replacement, and fake Judge JSON. +- [ ] Add security cases for duplicate keys, excessive nesting, oversized payloads, hostile Unicode/bidi, selector reuse, feedback forgery, cross-tenant IDs, endpoint SSRF/DNS rebinding, trace leakage, policy downgrade, model self-preference, and unsafe markup replacements. +- [ ] Add accessibility tests for diagnostic discovery, count, keyboard navigation, focus return, action announcements, non-color-only rendering, mobile target sizes, screen-reader names, review arrival without focus theft, and review-unavailable messaging. +- [ ] Add degraded-mode tests proving editor typing, one-shot draft proposal, clear, undo, and send still work with orchestrator down, Judge down, policy missing, review disabled, or calibration unsupported. +- [ ] Extend full-product and pilot smoke scripts with the reviewed happy path and explicit no-keyword-fallback negative controls. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_security.py + +cd ../frontend +pnpm exec vitest run src/components/email/emailWritingAccessibility.test.tsx +pnpm test:e2e -- e2e/email-writing-guidance.spec.ts +pnpm full:smoke +pnpm pilot:smoke +``` + +- [ ] Commit: + +```bash +git add frontend/e2e/email-writing-guidance.spec.ts backend/tests/test_email_writing_security.py frontend/src/components/email/emailWritingAccessibility.test.tsx frontend/scripts/full-product-ui-smoke.mjs frontend/scripts/pilot-ui-smoke.mjs docs/TEST_STRATEGY.md +git commit -m "test(email-writing): verify secure contextual guidance end to end" +``` + +## Task 16: Reconcile Product, Architecture, API, Data, Threat, Operability, and Doctoring Documents + +**Files:** +- Modify: `README.md` +- Modify: `AGENTS.md` +- Modify: `ARCHITECTURE.md` +- Modify: `CHANGELOG.md` +- Modify: `docs/architecture/naruon-product-spec.md` +- Modify: `docs/PRD.md` +- Modify: `docs/TRD.md` +- Modify: `docs/API_CONTRACT.md` +- Modify: `docs/DATA_MODEL.md` +- Modify: `docs/THREAT_MODEL.md` +- Modify: `docs/TEST_STRATEGY.md` +- Modify: `docs/OPERABILITY.md` +- Modify: `docs/TRACEABILITY.md` +- Modify: `docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md` +- Modify: `docs/adr/README.md` +- Modify: `docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md` +- Create: `backend/tests/test_email_writing_documentation_contract.py` + +- [ ] Reconcile canonical documents with the implemented ownership, data flow, API, evidence tables, retention, feature flag, errors, rollback, and release dependencies. +- [ ] Add Mermaid component, sequence, state, deployment, and evidence-lineage diagrams that match actual module and table names. +- [ ] Document PII/confidential-data controls without claiming that destructive masking is safe for pragmatics or actionability. +- [ ] Document language/model/rubric validation scope and explicitly distinguish architecture support from empirical validity. +- [ ] Document candidate/Judge/adjudicator separation, compute allocation, reasoning-effort ablations, and why speed is not the primary acceptance criterion. +- [ ] Add an operational runbook for orchestrator outage, policy expiry/revocation, Judge disagreement, capacity saturation, high stale rate, calibration drift, model retirement, feedback backlog, retention deletion, rollback, and incident evidence. +- [ ] Add traceability from every ADR requirement to source, test, benchmark, policy artifact, workflow, database constraint, frontend state, and release gate. +- [ ] Keep ADR 0001 `Proposed` until protected `develop` contains the runtime implementation and exact-head acceptance evidence; promote only during release reconciliation. +- [ ] Update `CHANGELOG.md` under `Unreleased` without claiming a validated language/model profile before Task 14 passes. +- [ ] Add documentation contract tests that fail if keyword fallback, direct provider calls, unjudged candidate mutation, implicit send gating, raw-content telemetry, mutable dependencies, or synchronous model fitting reappear. +- [ ] Ensure all research and standards citations in doctoring use APA 7th and are tied to explicit product decisions and limitations. +- [ ] Run: + +```bash +cd backend +python -m pytest -q tests/test_email_writing_documentation_contract.py +``` + +- [ ] Commit: + +```bash +git add README.md AGENTS.md ARCHITECTURE.md CHANGELOG.md docs/architecture/naruon-product-spec.md docs/PRD.md docs/TRD.md docs/API_CONTRACT.md docs/DATA_MODEL.md docs/THREAT_MODEL.md docs/TEST_STRATEGY.md docs/OPERABILITY.md docs/TRACEABILITY.md docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md docs/adr/README.md docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md backend/tests/test_email_writing_documentation_contract.py +git commit -m "docs(email-writing): reconcile product and assurance contracts" +``` + +## Task 17: Exact-Head Repository Acceptance and Merge + +- [ ] Rebase or merge the latest protected `develop` while preserving valid concurrent changes. +- [ ] Run all backend tests on Python 3.12, 3.13, and 3.14 as defined by repository policy. +- [ ] Run backend compile, lint, type/static, docstring, statement coverage, and branch coverage gates; production coverage must remain 100%. +- [ ] Run all frontend unit tests, typecheck, lint, coverage, production build, full smoke, pilot smoke, and Playwright end-to-end tests. +- [ ] Build and validate backend, frontend, and combined container images. +- [ ] Run migration upgrade/downgrade against the supported PostgreSQL matrix. +- [ ] Run dependency review, CodeQL, Semgrep/SAST, secret scan, OSV, Trivy, Scorecard, SBOM, provenance, image validation, coverage evidence, OpenCode review, and every current required check on the exact final head. +- [ ] Review every current-head CodeRabbit, GitHub Advanced Security, Dependabot, OpenCode, Noema, Strix, human, and other applicable finding. +- [ ] Resolve every valid review thread and rerun affected direct and CI tests. +- [ ] Confirm zero valid unresolved review threads. +- [ ] Obtain qualifying non-author current-head approval. +- [ ] Move the PR from Draft to Ready only after implementation, calibration policy, documentation, direct validation, and exact-head gates are complete. +- [ ] Merge without bypass only when all protected checks and approvals pass. +- [ ] Refetch protected `develop`, open PRs, migrations, package locks, policy manifest, and release metadata after merge. + +## Task 18: Versioned Release, Rollback Rehearsal, and Post-Release Drift Monitoring + +- [ ] Open a separate release-only PR after feature merge. +- [ ] Promote ADR 0001 to `Accepted` only with protected-branch implementation and exact-head evidence. +- [ ] Increase backend/frontend/repository versions consistently and create the final CHANGELOG release section. +- [ ] Bind the release to exact Inkspan, fast-mlsirm, contextual-orchestrator API, rubric, prompt, approved policy, dataset, and source commit versions. +- [ ] Generate immutable artifacts, SBOM, provenance, migration manifest, policy manifest, benchmark report, rollback instructions, and operator acceptance packet. +- [ ] Rehearse rollback by disabling semantic review, reverting to the prior approved policy, preserving Inkspan authoring/send, applying database downgrade only when safe, and proving no canonical email content migration is required. +- [ ] Publish only from the exact reviewed release head. +- [ ] Run a production-like canary with no raw-content telemetry and explicit tenant opt-in. +- [ ] Schedule drift monitoring for model/provider/prompt/rubric/policy/language profiles using fast-mlsirm response evidence and human feedback. +- [ ] Automatically revoke or withhold a policy profile when drift, preservation failure, DIF, unsupported-claim rate, security regression, or consequence thresholds cross the approved boundary. +- [ ] Do not claim universal language support, perfect grammar correction, objective tone truth, compliance certification, or send safety from this release. From 41cd2be90f92788c4b545d133a512cc2df2e955f Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Sat, 15 Aug 2026 20:36:39 +0900 Subject: [PATCH 07/18] ci(email-writing): refresh ADR branch on current develop --- .../refresh-email-writing-adr-base.yml | 94 +++++++++++++++++++ 1 file changed, 94 insertions(+) create mode 100644 .github/workflows/refresh-email-writing-adr-base.yml diff --git a/.github/workflows/refresh-email-writing-adr-base.yml b/.github/workflows/refresh-email-writing-adr-base.yml new file mode 100644 index 000000000..e18fa8553 --- /dev/null +++ b/.github/workflows/refresh-email-writing-adr-base.yml @@ -0,0 +1,94 @@ +name: Refresh Email Writing ADR Base + +on: + push: + branches: + - feat/inkspan-email-writing-guide + workflow_dispatch: + +permissions: + contents: write + +concurrency: + group: refresh-email-writing-adr-base + cancel-in-progress: true + +env: + FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: true + PYTHONWARNINGS: error + DISABLE_BACKGROUND_WORKERS: "1" + +jobs: + refresh: + runs-on: ubuntu-24.04 + timeout-minutes: 20 + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + ref: feat/inkspan-email-writing-guide + persist-credentials: true + fetch-depth: 0 + - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 + with: + python-version: "3.14" + cache: pip + cache-dependency-path: backend/requirements-hashes.txt + - name: Merge the current protected base + run: | + set -euo pipefail + git fetch origin develop + git config user.name "CWL Email Writing ADR Refresh" + git config user.email "actions@users.noreply.github.com" + git merge --no-edit --no-ff origin/develop + - name: Install hash-locked backend dependencies + run: python -m pip install --disable-pip-version-check --require-hashes -r backend/requirements-hashes.txt + - name: Reproduce the former backend blocker + run: | + cd backend + python -m pytest -q \ + 'tests/test_text_safety.py::test_strip_html_markup_never_returns_raw_tag_like_payloads[-->-]' + - name: Validate ADR and implementation-plan structure + run: | + python - <<'PY' + from pathlib import Path + + required = { + Path("docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md"): ( + "## Decision", + "## Security and privacy impact", + "## Verification and acceptance evidence", + "## Rollback or supersession", + ), + Path( + "docs/superpowers/plans/" + "2026-08-12-llm-email-writing-guidance-implementation.md" + ): ( + "Task 1", + "Task 18", + "NVIDIA_NIM_API_KEY", + ), + Path( + "docs/doctoring/" + "llm-email-writing-guidance-and-fast-mlsirm-judge.md" + ): ( + "## References", + "APA", + ), + } + for path, markers in required.items(): + text = path.read_text(encoding="utf-8") + missing = [marker for marker in markers if marker not in text] + if missing: + raise SystemExit(f"{path}: missing {missing}") + if text.count("```") % 2: + raise SystemExit(f"{path}: unbalanced Markdown fences") + print(f"Validated {len(required)} decision artifacts") + PY + git diff --check + - name: Remove the one-shot workflow and publish the refresh + run: | + set -euo pipefail + rm .github/workflows/refresh-email-writing-adr-base.yml + git add .github/workflows/refresh-email-writing-adr-base.yml + git commit -m "ci(email-writing): remove ADR base refresh workflow" + git push origin HEAD:feat/inkspan-email-writing-guide From bed7c216ac74078edfcae76e7b273f24324dd3f0 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Sat, 15 Aug 2026 20:39:53 +0900 Subject: [PATCH 08/18] ci(email-writing): align ADR doctoring marker --- .github/workflows/refresh-email-writing-adr-base.yml | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/.github/workflows/refresh-email-writing-adr-base.yml b/.github/workflows/refresh-email-writing-adr-base.yml index e18fa8553..e1edd2a89 100644 --- a/.github/workflows/refresh-email-writing-adr-base.yml +++ b/.github/workflows/refresh-email-writing-adr-base.yml @@ -71,8 +71,9 @@ jobs: "docs/doctoring/" "llm-email-writing-guidance-and-fast-mlsirm-judge.md" ): ( - "## References", - "APA", + "## APA 7th references", + "NIST AI 600-1", + "ISO/IEC 42001:2023", ), } for path, markers in required.items(): From 0dd28a3d4df5faa3c9007ec33785f73a16ded835 Mon Sep 17 00:00:00 2001 From: CWL Email Writing ADR Refresh Date: Sat, 15 Aug 2026 11:40:47 +0000 Subject: [PATCH 09/18] ci(email-writing): remove ADR base refresh workflow --- .../refresh-email-writing-adr-base.yml | 95 ------------------- 1 file changed, 95 deletions(-) delete mode 100644 .github/workflows/refresh-email-writing-adr-base.yml diff --git a/.github/workflows/refresh-email-writing-adr-base.yml b/.github/workflows/refresh-email-writing-adr-base.yml deleted file mode 100644 index e1edd2a89..000000000 --- a/.github/workflows/refresh-email-writing-adr-base.yml +++ /dev/null @@ -1,95 +0,0 @@ -name: Refresh Email Writing ADR Base - -on: - push: - branches: - - feat/inkspan-email-writing-guide - workflow_dispatch: - -permissions: - contents: write - -concurrency: - group: refresh-email-writing-adr-base - cancel-in-progress: true - -env: - FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: true - PYTHONWARNINGS: error - DISABLE_BACKGROUND_WORKERS: "1" - -jobs: - refresh: - runs-on: ubuntu-24.04 - timeout-minutes: 20 - steps: - - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - with: - ref: feat/inkspan-email-writing-guide - persist-credentials: true - fetch-depth: 0 - - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 - with: - python-version: "3.14" - cache: pip - cache-dependency-path: backend/requirements-hashes.txt - - name: Merge the current protected base - run: | - set -euo pipefail - git fetch origin develop - git config user.name "CWL Email Writing ADR Refresh" - git config user.email "actions@users.noreply.github.com" - git merge --no-edit --no-ff origin/develop - - name: Install hash-locked backend dependencies - run: python -m pip install --disable-pip-version-check --require-hashes -r backend/requirements-hashes.txt - - name: Reproduce the former backend blocker - run: | - cd backend - python -m pytest -q \ - 'tests/test_text_safety.py::test_strip_html_markup_never_returns_raw_tag_like_payloads[-->-]' - - name: Validate ADR and implementation-plan structure - run: | - python - <<'PY' - from pathlib import Path - - required = { - Path("docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md"): ( - "## Decision", - "## Security and privacy impact", - "## Verification and acceptance evidence", - "## Rollback or supersession", - ), - Path( - "docs/superpowers/plans/" - "2026-08-12-llm-email-writing-guidance-implementation.md" - ): ( - "Task 1", - "Task 18", - "NVIDIA_NIM_API_KEY", - ), - Path( - "docs/doctoring/" - "llm-email-writing-guidance-and-fast-mlsirm-judge.md" - ): ( - "## APA 7th references", - "NIST AI 600-1", - "ISO/IEC 42001:2023", - ), - } - for path, markers in required.items(): - text = path.read_text(encoding="utf-8") - missing = [marker for marker in markers if marker not in text] - if missing: - raise SystemExit(f"{path}: missing {missing}") - if text.count("```") % 2: - raise SystemExit(f"{path}: unbalanced Markdown fences") - print(f"Validated {len(required)} decision artifacts") - PY - git diff --check - - name: Remove the one-shot workflow and publish the refresh - run: | - set -euo pipefail - rm .github/workflows/refresh-email-writing-adr-base.yml - git add .github/workflows/refresh-email-writing-adr-base.yml - git commit -m "ci(email-writing): remove ADR base refresh workflow" - git push origin HEAD:feat/inkspan-email-writing-guide From beb6ce14a7ea5922a2da481e7bab55d10847678f Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Sat, 15 Aug 2026 21:05:56 +0900 Subject: [PATCH 10/18] ci(email-writing): reconcile ADR numbering with current develop --- .../reconcile-email-writing-adr-number.yml | 178 ++++++++++++++++++ 1 file changed, 178 insertions(+) create mode 100644 .github/workflows/reconcile-email-writing-adr-number.yml diff --git a/.github/workflows/reconcile-email-writing-adr-number.yml b/.github/workflows/reconcile-email-writing-adr-number.yml new file mode 100644 index 000000000..1a06aff6c --- /dev/null +++ b/.github/workflows/reconcile-email-writing-adr-number.yml @@ -0,0 +1,178 @@ +name: Reconcile Email Writing ADR Number + +on: + push: + branches: + - feat/inkspan-email-writing-guide + workflow_dispatch: + +permissions: + contents: write + +concurrency: + group: reconcile-email-writing-adr-number + cancel-in-progress: true + +env: + FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: true + PYTHONWARNINGS: error + DISABLE_BACKGROUND_WORKERS: "1" + +jobs: + reconcile: + runs-on: ubuntu-24.04 + timeout-minutes: 25 + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + ref: feat/inkspan-email-writing-guide + persist-credentials: true + fetch-depth: 0 + - name: Merge current develop and resolve only the ADR index conflict + run: | + set -euo pipefail + git fetch origin develop + git config user.name "CWL Email Writing ADR Reconciliation" + git config user.email "actions@users.noreply.github.com" + set +e + git merge --no-commit --no-ff origin/develop + merge_status=$? + set -e + conflicts="$(git diff --name-only --diff-filter=U)" + if [[ "$merge_status" -ne 0 && "$conflicts" != "docs/adr/README.md" ]]; then + printf 'Unexpected merge conflicts:\n%s\n' "$conflicts" >&2 + exit 1 + fi + cat >docs/adr/README.md <<'EOF' + # Naruon Architecture Decision Records + + This index records cross-cutting Naruon decisions that must survive beyond an + individual pull request, implementation plan, or chat. `Accepted` means only that + the Naruon decision governs its stated local scope; it does not transfer authority + to an external service or mean that a future integration is implemented. + `Proposed` records a discoverable target for later review and does not govern + implementation. + + | ADR | Decision | Status | Capability effect | + |---|---|---|---| + | [ADR-0001](0001-topic-measurement-authority.md) | Consume structural topic measurement only from scientific authority, never a keyword or label heuristic | Accepted | `ACCEPTED-NARUON-POLICY`; no runtime promotion | + | [ADR-0002](0002-fitted-topic-artifact-consumption.md) | Conditionally consume only a versioned fitted topic artifact through a fail-closed adapter | Proposed | Target `PLANNED`; runtime `BLOCKED-UPSTREAM` | + | [ADR-0003](0003-separate-topic-measurement-from-agenda-generation.md) | Keep statistical topic measurement separate from agenda generation | Proposed | Future capability `PLANNED`; no implementation authorization | + | [ADR-0004](0004-inkspan-backed-llm-email-writing-guidance.md) | Compose Inkspan, contextual-orchestrator, and fast-mlsirm for LLM-native email guidance without lexical semantic fallback | Proposed | Design and TDD plan only; runtime remains unshipped | + + The complete topic-intelligence requirements, architecture, contract, UML, + conceptual ERD, security, test, and operability graph is indexed at + [`docs/topic-intelligence/README.md`](../topic-intelligence/README.md). + Its [canonical digest inventory](../topic-intelligence/README.md#canonical-digest-inventory) + is the single cross-document list for the planned adapter profile. + + ## Decision discipline + + - **Proposed:** design and ownership boundaries are documented, but implementation + or operational acceptance evidence is incomplete. + - **Accepted:** protected `develop` contains the governing implementation or + process and its tests, security evidence, documentation, and rollback contract + are current. + - **Superseded:** retained for traceability but replaced by a later ADR. + + ## Required ADR sections + + Every material ADR records context, alternatives, decision, consequences, failure + and recovery, security and privacy, accessibility where applicable, compatibility + and migration, verification, research or standards traceability, and rollback or + supersession conditions. + + ## Change rule + + Create or update an ADR when a Naruon change adopts or declines an external service + contract, introduces a scientific or statistical inference contract, changes + persistence or tenant authority, changes model or credential trust boundaries, or + replaces a fail-closed product capability with a different production dependency. + A Naruon ADR records Naruon's decision only; it cannot assign authority to, or + accept a decision for, another service. + + Every implementing PR must keep the corresponding source, tests, doctoring, + architecture and operability contract, and CHANGELOG maturity truthful. An active + PR, accepted local policy, or proposed target must not be described as protected- + branch implementation before it is integrated and independently verified. + EOF + git add docs/adr/README.md + - name: Assign the next non-conflicting ADR identifier + run: | + set -euo pipefail + git mv \ + docs/adr/0001-inkspan-backed-llm-email-writing-guidance.md \ + docs/adr/0004-inkspan-backed-llm-email-writing-guidance.md + python - <<'PY' + from pathlib import Path + + path = Path( + "docs/adr/0004-inkspan-backed-llm-email-writing-guidance.md" + ) + text = path.read_text(encoding="utf-8") + old = "# ADR 0001: Inkspan-backed, LLM-native email writing guidance\n" + new = "# ADR 0004: Inkspan-backed, LLM-native email writing guidance\n" + if text.count(old) != 1: + raise SystemExit("email-writing ADR heading anchor changed") + path.write_text(text.replace(old, new, 1), encoding="utf-8") + PY + - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 + with: + python-version: "3.14" + cache: pip + cache-dependency-path: backend/requirements-hashes.txt + - name: Install hash-locked backend dependencies + run: python -m pip install --disable-pip-version-check --require-hashes -r backend/requirements-hashes.txt + - name: Verify merged documentation and former base regression + run: | + python - <<'PY' + from pathlib import Path + import re + + adr_dir = Path("docs/adr") + numbered = [ + path + for path in adr_dir.glob("[0-9][0-9][0-9][0-9]-*.md") + ] + numbers = [path.name[:4] for path in numbered] + duplicates = sorted({number for number in numbers if numbers.count(number) > 1}) + if duplicates: + raise SystemExit(f"Duplicate ADR identifiers: {duplicates}") + + index = (adr_dir / "README.md").read_text(encoding="utf-8") + links = re.findall(r"\((\d{4}-[^)]+\.md)\)", index) + missing = [name for name in links if not (adr_dir / name).is_file()] + if missing: + raise SystemExit(f"Missing indexed ADR files: {missing}") + + email_adr = ( + adr_dir / "0004-inkspan-backed-llm-email-writing-guidance.md" + ).read_text(encoding="utf-8") + required = ( + "# ADR 0004:", + "## Decision", + "## Security and privacy impact", + "## Verification and acceptance evidence", + "## Rollback or supersession", + ) + absent = [marker for marker in required if marker not in email_adr] + if absent: + raise SystemExit(f"Email-writing ADR missing markers: {absent}") + if email_adr.count("```") % 2: + raise SystemExit("Email-writing ADR has unbalanced fences") + PY + cd backend + python -m pytest -q \ + 'tests/test_text_safety.py::test_strip_html_markup_never_returns_raw_tag_like_payloads[-->-]' + cd .. + git diff --check + - name: Commit the reconciled merge and remove the one-shot workflow + run: | + set -euo pipefail + rm .github/workflows/reconcile-email-writing-adr-number.yml + git add \ + .github/workflows/reconcile-email-writing-adr-number.yml \ + docs/adr/README.md \ + docs/adr/0004-inkspan-backed-llm-email-writing-guidance.md + git commit -m "docs(adr): reconcile email-writing decision numbering" + git push origin HEAD:feat/inkspan-email-writing-guide From c5b242e676cf2e3494b15bc073dac04ddb73a824 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Fri, 21 Aug 2026 05:53:36 +0900 Subject: [PATCH 11/18] docs(email-writing): preregister judge policy evidence --- ...kspan-backed-llm-email-writing-guidance.md | 28 ++++++++++++++++--- ...-writing-guidance-and-fast-mlsirm-judge.md | 24 ++++++++++++++-- ...m-email-writing-guidance-implementation.md | 8 ++++-- ...kspan-llm-email-writing-guidance-design.md | 12 +++++++- 4 files changed, 63 insertions(+), 9 deletions(-) diff --git a/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md b/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md index 7b4550c74..f85937c2b 100644 --- a/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md +++ b/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md @@ -163,9 +163,26 @@ Naruon will use Inkspan's diagnostic surface rather than a hover-only overlay. U The existing draft endpoint may remain for one-shot generation during migration, but the reply composer will move from the plain textarea to the first released Inkspan version that contains the accepted writing-diagnostic contract. Naruon must consume an immutable released package with its lockfile and package-verification evidence; it will not depend on an unreleased mutable branch in production. -The first implementation is additive and does not require a database migration if review sessions remain ephemeral. Any later persistent tables must use two-or-more-word `snake_case` names such as `email_review_session`, `writing_diagnostic_record`, `diagnostic_feedback_event`, and `judge_policy_artifact`. - -Rollback disables the review endpoint and diagnostic props while retaining the Inkspan editor and existing mail-send path. No canonical email content migration is required. +The first release persists privacy-minimized review evidence. Migration +`20260812_0001_add_email_writing_review_evidence.py` adds +`email_review_session`, `writing_diagnostic_record`, and +`diagnostic_feedback_event`; these objects contain opaque identifiers, hashes, +scope, policy provenance, bounded timing/cost buckets, selectors, statuses, and +explicit feedback only. Raw email, draft, replacement, explanation, prompt, +model output, and provider credentials are excluded. Evidence expires after 30 +days by default, and an owner-scoped deletion job removes expired sessions and +their dependent records without scanning or rewriting canonical email content. + +The migration must prove PostgreSQL upgrade/downgrade and the repository's +SQLite contract. Downgrade is allowed only after semantic review is disabled and +dependent evidence is deleted; it drops only the three feature-owned tables and +never alters canonical email tables or mail content. A failed evidence +transaction rolls back the review session and diagnostics together, so the API +never returns unrecorded evidence. + +Rollback disables the review endpoint and diagnostic props while retaining the +Inkspan editor and existing mail-send path. No canonical email content +migration is required. ## Verification and acceptance evidence @@ -179,6 +196,9 @@ Implementation cannot be accepted without: - gold human span/category/replacement data with inter-rater evidence; - span precision/recall, category macro-F1, accepted-suggestion precision, intent/fact/request-strength preservation, Brier score, calibration error, and unsupported-claim rate; - fast-mlsirm criterion response-matrix validation, category-count ablation, item/rater behavior, DIF, and drift studies where sample size supports them; +- a pre-registered evaluation protocol with fixed splits, publication thresholds, + and a locked human-labeled holdout; the protocol and holdout hashes must be + recorded in the published policy artifact before calibration begins; - independent-model or role ablations for candidate reviewer, judge, and adjudicator; - prompt-injection, malformed JSON, duplicate keys, oversized payload, stale revision, overlap, and provider failure tests; - no-keyword-fallback source and contract tests; @@ -194,4 +214,4 @@ The accompanying doctoring record covers LLM-as-a-Judge reliability, bias, multi ## Rollback or supersession -Rollback removes or disables semantic review while preserving the editor, mail data, and send path. Supersession requires a new ADR if Naruon is proposed to use keyword-based semantic classification, allow unjudged model output to mutate drafts, move email semantics into Inkspan, put fast-mlsirm fitting in the synchronous critical path, mask context needed for valid interpretation, or make suggestions an implicit send gate. \ No newline at end of file +Rollback removes or disables semantic review while preserving the editor, mail data, and send path. Supersession requires a new ADR if Naruon is proposed to use keyword-based semantic classification, allow unjudged model output to mutate drafts, move email semantics into Inkspan, put fast-mlsirm fitting in the synchronous critical path, mask context needed for valid interpretation, or make suggestions an implicit send gate. diff --git a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md index d07236a0e..ac9c090af 100644 --- a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md +++ b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md @@ -9,7 +9,7 @@ This record supports Naruon's decision to use contextual LLM judgment for email ## Current CWL implementation evidence -`ContextualWisdomLab/fast-mlsirm` pull request #733 introduced a provider-neutral `ContextualOrchestratorJudge`, strict criterion-level JSON, explicit polytomous categories, runtime-derived acceptance, `LLMJudgeResult.to_irt_row()`, multi-item response-matrix validation, and a deliberate no-keyword/no-positional-repair boundary. All judge model calls are injected through contextual-orchestrator rather than bound to one provider. +The verified source is [`ContextualWisdomLab/fast-mlsirm` at tag `v0.6.0`](https://github.com/ContextualWisdomLab/fast-mlsirm/tree/v0.6.0), whose tag resolves to commit `1bde00930502296ecc35963c2bdcb834ca15f87d1`. The related merged change is [PR #733](https://github.com/ContextualWisdomLab/fast-mlsirm/pull/733), merge commit `914127ba227d3e02d0564aeeb4f27d76137610f9`; the tag archive SHA-256 is `73415e3cd8d233df7c561ac4eb98c122ebf40e0ac7312d9775266d8830beca4e`. The verified public surface includes the provider-neutral `ContextualOrchestratorJudge`, strict criterion-level JSON, explicit polytomous categories, runtime-derived acceptance, `LLMJudgeResult.to_irt_row()`, multi-item response-matrix validation, and a deliberate no-keyword/no-positional-repair boundary. All judge model calls are injected through contextual-orchestrator rather than bound to one provider. This is useful infrastructure, not proof that an email-writing rubric is valid. Naruon must still define the construct, criteria, category anchors, evaluation cases, language profiles, human reference process, calibration policy, and consequences of false positive or false negative guidance. @@ -133,6 +133,26 @@ Report, at minimum: No overall “email danger” or “tone risk” score substitutes for the criterion results. +### Pre-registered publication protocol + +Before calibration or threshold selection, the evaluation runner writes a +canonical `email_writing_policy_protocol_v1` document and records its SHA-256. +The split is fixed by contrast family, not by individual rows: 60% calibration, +20% development, and a 20% human-labeled locked holdout. The holdout is read +only until the final publish/no-publish decision and its case-manifest SHA-256 is +stored beside the protocol hash. The initial publication gates are fixed in the +protocol: holdout macro-F1 at least 0.80, every mandatory preservation criterion +at least 0.90, unsupported-claim rate at most 0.02, expected calibration error +at most 0.05, and no prespecified language/role/context slice with macro-F1 +below 0.70 when that slice has the preregistered minimum sample size. + +The policy artifact must contain `protocol_id`, `protocol_hash`, +`calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, the +literal thresholds, and `publish_decision` (`publish` or `withhold`). Changing +any split, threshold, or holdout requires a new protocol version and a new +calibration run; post-hoc threshold tuning against the locked holdout is not a +publication decision. + ## Multi-agent and compute allocation The workflow can allocate more test-time computation without lexical routing: @@ -202,4 +222,4 @@ Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li ## Claim boundary -The cited evidence supports structured, bias-aware, calibrated LLM evaluation and revision-bound editor integrity. It does not prove that the planned Naruon rubric, model, provider, language profile, or fast-mlsirm configuration is sufficiently valid. Those claims require the benchmark, ablation, DIF, drift, security, privacy, and user-consequence evidence defined in the design. \ No newline at end of file +The cited evidence supports structured, bias-aware, calibrated LLM evaluation and revision-bound editor integrity. It does not prove that the planned Naruon rubric, model, provider, language profile, or fast-mlsirm configuration is sufficiently valid. Those claims require the benchmark, ablation, DIF, drift, security, privacy, and user-consequence evidence defined in the design. diff --git a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md index c198e9148..08f7ebdbb 100644 --- a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md +++ b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md @@ -289,7 +289,7 @@ git commit -m "feat(email-writing): parse contextual review candidates" - `request_strength_preservation`; - `audience_pragmatics`; - `technical_precision`; - - `actionability_support`; + - `actionability`; - `explanation_quality`. - [ ] Make the required criterion subset depend on candidate type without changing IDs or anchors. A no-replacement diagnostic does not fabricate `replacement_correctness`; the policy schema declares which criteria are mandatory for each candidate kind. - [ ] Use explicit ordered categories and exact anchors. The initial category count remains an evaluation parameter until Task 14 ablation selects and publishes an approved value. @@ -376,7 +376,8 @@ context build - [ ] Incremental mode reviews the declared changed selector plus bounded authorized thread context; deep mode reviews the complete bounded draft and may use contextual-orchestrator `conduct` mode with decomposed reviewer roles. - [ ] Semantic escalation is based on structured Judge disagreement, policy uncertainty, missing context, preservation-criterion conflict, or approved compute policy. It is never triggered by lexical words or phrase counts. - [ ] Bound candidate count, Judge calls, concurrency, context bytes, draft bytes, orchestration steps, retries, total wall time, and cost/tokens by policy. -- [ ] Use one database transaction for review session and admitted/withheld diagnostic metadata; roll back on persistence failure without returning unrecorded diagnostics. +- [ ] Persist the session and admitted/withheld diagnostic metadata in the Task 3 feature-owned tables through one database transaction; apply the default 30-day expiry and owner-scoped deletion contract without storing raw content. +- [ ] Roll back the transaction on persistence failure without returning unrecorded diagnostics. During rollback, disable semantic review first and downgrade only after dependent evidence is deleted; never alter canonical email tables or mail content. - [ ] Return no raw candidate or Judge output. Return admitted diagnostics, non-mutating document guidance, limitations, abstentions, and redacted provenance only. - [ ] If one candidate fails, apply the policy's explicit atomicity rule. Do not silently drop failed candidates while claiming a complete review. - [ ] Ensure review cancellation closes provider requests, releases capacity limiters, and leaves a typed cancelled/abstained session rather than a running record. @@ -557,6 +558,7 @@ git commit -m "feat(email-writing): connect guidance actions and evidence" - Create after passing evidence: `backend/policies/email_writing_judge_approved_v1.json` - [ ] Define a consented/de-identified or synthetic case schema with source context, recipient roles, draft, gold spans, categories, acceptable replacements, preservation constraints, annotator IDs, adjudication, and consequence severity. +- [ ] Before calibration, publish the canonical `email_writing_policy_protocol_v1` with a SHA-256, fixed contrast-family splits of 60% calibration, 20% development, and 20% locked human-labeled holdout, plus the literal publication gates: holdout macro-F1 >= 0.80, every mandatory preservation criterion >= 0.90, unsupported-claim rate <= 0.02, expected calibration error <= 0.05, and no prespecified slice below macro-F1 0.70 when its minimum sample size is met. - [ ] Include contrast families: - same words in quotation, neutral incident report, and direct rebuke; - same pragmatic issue expressed through unrelated paraphrases; @@ -567,12 +569,14 @@ git commit -m "feat(email-writing): connect guidance actions and evidence" - terse but acceptable negative controls; - Korean, English, mixed-language, CJK, emoji, combining-mark, and hostile-Unicode cases. - [ ] Collect independent human annotations and report inter-rater agreement plus adjudicated disagreement. Human labels are reference evidence, not unquestioned truth. +- [ ] Freeze the holdout manifest before any calibration or threshold selection; record calibration, development, and locked-holdout hashes in every result and keep the holdout read-only until the final publish/no-publish decision. - [ ] Measure issue/category precision, recall, macro-F1, span IoU/exactness, replacement correctness, intent/fact/actor/deadline/request-strength preservation, unsupported-claim rate, accepted-suggestion precision, Brier score, calibration error, and human ignore rate. - [ ] Use fast-mlsirm response matrices to analyze criterion/category behavior, item/rater effects where supported, reliability, prompt/model test-retest, category sparsity, category-count ablation, and DIF by language, recipient configuration, role/hierarchy, thread depth, document length, review mode, model, prompt, and time. - [ ] Compare single-model, separated candidate/Judge, and deeper multi-agent/adjudicator workflows with role-specific reasoning effort. Speed is recorded but does not override validity. - [ ] Use deterministic offline fixtures in required CI. Run live scheduled/manual evaluation with `NVIDIA_NIM_API_KEY`; do not expose the secret to pull-request code from untrusted forks and do not use `COPILOT_GITHUB_TOKEN`. - [ ] Fail the policy publication job when sample size, calibration, preservation, DIF, drift, or consequence thresholds are not met. Do not publish a policy merely because average accuracy is high. - [ ] Generate the approved policy artifact from measured outputs, review its literal values, add its SHA-256 to the manifest, and prove production rejects the earlier `evaluation_only` artifact for user-facing output. +- [ ] Include `protocol_id`, `protocol_hash`, `calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, literal thresholds, and `publish_decision` (`publish` or `withhold`) in the policy artifact. Any protocol, split, threshold, or holdout change creates a new protocol version and a new run; do not tune against the locked holdout post hoc. - [ ] Store reports, response matrices, aggregate metrics, manifests, and source hashes without confidential raw company email. - [ ] Run: diff --git a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md index cf5f861fb..f02b75459 100644 --- a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md +++ b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md @@ -321,9 +321,19 @@ A calibration run emits an integrity-bound `judge_policy_artifact` containing: - category anchors and admission rules; - calibration and DIF summary; - dataset and source hashes; +- a pre-registered `protocol_id` and `protocol_hash`; +- `calibration_split_hash`, `development_split_hash`, and + `locked_holdout_hash`, with the locked holdout reserved for the final + publish/no-publish decision; +- literal publication thresholds for holdout macro-F1, mandatory preservation, + unsupported-claim rate, expected calibration error, and prespecified slice + performance; +- a `publish_decision` of `publish` or `withhold`; - known limitations. -Naruon runtime consumes the artifact but does not refit it. +Naruon runtime consumes the artifact but does not refit it. Changing a split, +threshold, protocol, or holdout requires a new protocol version and a new +artifact; the locked holdout cannot be used for post-hoc threshold tuning. ## Benchmark design From afec4189ba2113e01605bf76e8bd9b9c67af9743 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Fri, 21 Aug 2026 08:39:14 +0900 Subject: [PATCH 12/18] docs(email-writing): close policy review gaps --- ...5-inkspan-backed-llm-email-writing-guidance.md | 15 +++++++++------ ...mail-writing-guidance-and-fast-mlsirm-judge.md | 10 ++++++++-- ...2-llm-email-writing-guidance-implementation.md | 10 ++++++---- ...2-inkspan-llm-email-writing-guidance-design.md | 12 +++++++++--- 4 files changed, 32 insertions(+), 15 deletions(-) diff --git a/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md b/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md index f85937c2b..4270614de 100644 --- a/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md +++ b/docs/adr/0005-inkspan-backed-llm-email-writing-guidance.md @@ -177,12 +177,15 @@ The migration must prove PostgreSQL upgrade/downgrade and the repository's SQLite contract. Downgrade is allowed only after semantic review is disabled and dependent evidence is deleted; it drops only the three feature-owned tables and never alters canonical email tables or mail content. A failed evidence -transaction rolls back the review session and diagnostics together, so the API -never returns unrecorded evidence. - -Rollback disables the review endpoint and diagnostic props while retaining the -Inkspan editor and existing mail-send path. No canonical email content -migration is required. +transaction rolls back the review session and diagnostics together, returns a +typed failure or abstention, and never returns unrecorded evidence. That +transaction failure does not itself disable the feature or downgrade its schema. + +Feature rollback is a separate, explicitly authorized operation: it disables +the review endpoint and diagnostic props, deletes dependent feature evidence, +and applies the documented downgrade only after those conditions are proven, +while retaining the Inkspan editor and existing mail-send path. No canonical +email content migration is required. ## Verification and acceptance evidence diff --git a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md index ac9c090af..c6c66f9fe 100644 --- a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md +++ b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md @@ -146,12 +146,18 @@ at least 0.90, unsupported-claim rate at most 0.02, expected calibration error at most 0.05, and no prespecified language/role/context slice with macro-F1 below 0.70 when that slice has the preregistered minimum sample size. +The preregistered `minimum_slice_sample_size` is 30 labeled cases per slice. A +slice below that count is reported as underpowered rather than selectively +excluded after seeing its result. + The policy artifact must contain `protocol_id`, `protocol_hash`, `calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, the literal thresholds, and `publish_decision` (`publish` or `withhold`). Changing any split, threshold, or holdout requires a new protocol version and a new -calibration run; post-hoc threshold tuning against the locked holdout is not a -publication decision. +calibration run. Thresholds must be frozen before the locked-holdout labels are +accessed. A run that tunes thresholds against the locked holdout is invalid for +publication and requires a new locked holdout; it is not a publication +decision. ## Multi-agent and compute allocation diff --git a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md index 08f7ebdbb..eab1d1f23 100644 --- a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md +++ b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md @@ -323,6 +323,7 @@ git commit -m "feat(email-writing): add independent fast-mlsirm Judge" - [ ] Define a strict JSON Schema for policy ID/version, status, creation/expiry, compatible Naruon/Inkspan/fast-mlsirm/orchestrator contracts, approved model/provider/rubric/language profiles, category count/anchors, mandatory criterion floors, posterior/score admission rules, adjudication conditions, calibration/DIF summary, dataset/provenance hashes, limitations, and rollback version. - [ ] Check in an integrity-bound `evaluation_only` policy that can exercise the pipeline but cannot emit user-facing diagnostics in production. +- [ ] Require runtime admission to assert `publish_decision == "publish"`; `withhold` and `evaluation_only` artifacts remain audit/pipeline evidence only and cannot produce user-facing diagnostics. - [ ] Add a manifest containing the literal SHA-256 of every allowed policy artifact; reject modified, unknown, expired, future-dated, revoked, incompatible, or non-approved artifacts. - [ ] Implement outcomes: @@ -377,7 +378,8 @@ context build - [ ] Semantic escalation is based on structured Judge disagreement, policy uncertainty, missing context, preservation-criterion conflict, or approved compute policy. It is never triggered by lexical words or phrase counts. - [ ] Bound candidate count, Judge calls, concurrency, context bytes, draft bytes, orchestration steps, retries, total wall time, and cost/tokens by policy. - [ ] Persist the session and admitted/withheld diagnostic metadata in the Task 3 feature-owned tables through one database transaction; apply the default 30-day expiry and owner-scoped deletion contract without storing raw content. -- [ ] Roll back the transaction on persistence failure without returning unrecorded diagnostics. During rollback, disable semantic review first and downgrade only after dependent evidence is deleted; never alter canonical email tables or mail content. +- [ ] Roll back the evidence transaction on persistence failure without returning unrecorded diagnostics. The transaction failure returns a typed failure or abstention; it does not silently disable the feature or downgrade its schema. +- [ ] Treat feature rollback as a separate, explicitly authorized operation: disable semantic review, delete dependent feature evidence, and downgrade only when the ADR rollback conditions are satisfied; never alter canonical email tables or mail content. - [ ] Return no raw candidate or Judge output. Return admitted diagnostics, non-mutating document guidance, limitations, abstentions, and redacted provenance only. - [ ] If one candidate fails, apply the policy's explicit atomicity rule. Do not silently drop failed candidates while claiming a complete review. - [ ] Ensure review cancellation closes provider requests, releases capacity limiters, and leaves a typed cancelled/abstained session rather than a running record. @@ -558,7 +560,7 @@ git commit -m "feat(email-writing): connect guidance actions and evidence" - Create after passing evidence: `backend/policies/email_writing_judge_approved_v1.json` - [ ] Define a consented/de-identified or synthetic case schema with source context, recipient roles, draft, gold spans, categories, acceptable replacements, preservation constraints, annotator IDs, adjudication, and consequence severity. -- [ ] Before calibration, publish the canonical `email_writing_policy_protocol_v1` with a SHA-256, fixed contrast-family splits of 60% calibration, 20% development, and 20% locked human-labeled holdout, plus the literal publication gates: holdout macro-F1 >= 0.80, every mandatory preservation criterion >= 0.90, unsupported-claim rate <= 0.02, expected calibration error <= 0.05, and no prespecified slice below macro-F1 0.70 when its minimum sample size is met. +- [ ] Before calibration, publish the canonical `email_writing_policy_protocol_v1` with a SHA-256, fixed contrast-family splits of 60% calibration, 20% development, and 20% locked human-labeled holdout, a fixed `minimum_slice_sample_size` of 30 cases per prespecified slice, plus the literal publication gates: holdout macro-F1 >= 0.80, every mandatory preservation criterion >= 0.90, unsupported-claim rate <= 0.02, expected calibration error <= 0.05, and no prespecified slice below macro-F1 0.70 when its minimum sample size is met. - [ ] Include contrast families: - same words in quotation, neutral incident report, and direct rebuke; - same pragmatic issue expressed through unrelated paraphrases; @@ -569,14 +571,14 @@ git commit -m "feat(email-writing): connect guidance actions and evidence" - terse but acceptable negative controls; - Korean, English, mixed-language, CJK, emoji, combining-mark, and hostile-Unicode cases. - [ ] Collect independent human annotations and report inter-rater agreement plus adjudicated disagreement. Human labels are reference evidence, not unquestioned truth. -- [ ] Freeze the holdout manifest before any calibration or threshold selection; record calibration, development, and locked-holdout hashes in every result and keep the holdout read-only until the final publish/no-publish decision. +- [ ] Freeze the holdout manifest before any calibration or threshold selection; record calibration, development, locked-holdout, and protocol hashes in every result and keep the holdout read-only until the final publish/no-publish decision. - [ ] Measure issue/category precision, recall, macro-F1, span IoU/exactness, replacement correctness, intent/fact/actor/deadline/request-strength preservation, unsupported-claim rate, accepted-suggestion precision, Brier score, calibration error, and human ignore rate. - [ ] Use fast-mlsirm response matrices to analyze criterion/category behavior, item/rater effects where supported, reliability, prompt/model test-retest, category sparsity, category-count ablation, and DIF by language, recipient configuration, role/hierarchy, thread depth, document length, review mode, model, prompt, and time. - [ ] Compare single-model, separated candidate/Judge, and deeper multi-agent/adjudicator workflows with role-specific reasoning effort. Speed is recorded but does not override validity. - [ ] Use deterministic offline fixtures in required CI. Run live scheduled/manual evaluation with `NVIDIA_NIM_API_KEY`; do not expose the secret to pull-request code from untrusted forks and do not use `COPILOT_GITHUB_TOKEN`. - [ ] Fail the policy publication job when sample size, calibration, preservation, DIF, drift, or consequence thresholds are not met. Do not publish a policy merely because average accuracy is high. - [ ] Generate the approved policy artifact from measured outputs, review its literal values, add its SHA-256 to the manifest, and prove production rejects the earlier `evaluation_only` artifact for user-facing output. -- [ ] Include `protocol_id`, `protocol_hash`, `calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, literal thresholds, and `publish_decision` (`publish` or `withhold`) in the policy artifact. Any protocol, split, threshold, or holdout change creates a new protocol version and a new run; do not tune against the locked holdout post hoc. +- [ ] Include `protocol_id`, `protocol_hash`, `calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, literal `minimum_slice_sample_size: 30`, literal thresholds, and `publish_decision` (`publish` or `withhold`) in the policy artifact. Freeze thresholds before reading locked-holdout labels; a run that tunes thresholds against that holdout is invalid for publication and requires a new locked holdout. Any protocol, split, threshold, or holdout change creates a new protocol version and a new run. - [ ] Store reports, response matrices, aggregate metrics, manifests, and source hashes without confidential raw company email. - [ ] Run: diff --git a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md index f02b75459..387c4eb65 100644 --- a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md +++ b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md @@ -325,15 +325,21 @@ A calibration run emits an integrity-bound `judge_policy_artifact` containing: - `calibration_split_hash`, `development_split_hash`, and `locked_holdout_hash`, with the locked holdout reserved for the final publish/no-publish decision; +- a fixed `minimum_slice_sample_size` of 30 labeled cases per prespecified + slice; - literal publication thresholds for holdout macro-F1, mandatory preservation, unsupported-claim rate, expected calibration error, and prespecified slice performance; - a `publish_decision` of `publish` or `withhold`; - known limitations. -Naruon runtime consumes the artifact but does not refit it. Changing a split, -threshold, protocol, or holdout requires a new protocol version and a new -artifact; the locked holdout cannot be used for post-hoc threshold tuning. +Naruon runtime consumes the artifact but does not refit it, and admits it for +user-facing diagnostics only when `publish_decision == "publish"`. `withhold` +and `evaluation_only` artifacts remain audit/pipeline evidence and cannot +produce user-facing output. Changing a split, threshold, protocol, or holdout +requires a new protocol version and a new artifact. Thresholds are frozen before +locked-holdout labels are accessed; tuning against that holdout invalidates the +run for publication and requires a new locked holdout. ## Benchmark design From ec0e1d367a0fb03f4431286260897e010421ef0c Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 1 Sep 2026 20:46:57 +0900 Subject: [PATCH 13/18] docs(email-writing): refresh immutable judge evidence --- ...lm-email-writing-guidance-and-fast-mlsirm-judge.md | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md index c6c66f9fe..06b19899d 100644 --- a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md +++ b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md @@ -1,7 +1,8 @@ # LLM email writing guidance and fast-mlsirm judge calibration **Status:** Proposed architecture evidence; no production accuracy or language-coverage claim is made. -**Date:** 2026-08-12 +**Date:** 2026-08-12 +**Evidence refreshed:** 2026-09-01 ## Purpose @@ -9,9 +10,13 @@ This record supports Naruon's decision to use contextual LLM judgment for email ## Current CWL implementation evidence -The verified source is [`ContextualWisdomLab/fast-mlsirm` at tag `v0.6.0`](https://github.com/ContextualWisdomLab/fast-mlsirm/tree/v0.6.0), whose tag resolves to commit `1bde00930502296ecc35963c2bdcb834ca15f87d1`. The related merged change is [PR #733](https://github.com/ContextualWisdomLab/fast-mlsirm/pull/733), merge commit `914127ba227d3e02d0564aeeb4f27d76137610f9`; the tag archive SHA-256 is `73415e3cd8d233df7c561ac4eb98c122ebf40e0ac7312d9775266d8830beca4e`. The verified public surface includes the provider-neutral `ContextualOrchestratorJudge`, strict criterion-level JSON, explicit polytomous categories, runtime-derived acceptance, `LLMJudgeResult.to_irt_row()`, multi-item response-matrix validation, and a deliberate no-keyword/no-positional-repair boundary. All judge model calls are injected through contextual-orchestrator rather than bound to one provider. +The current immutable `fast-mlsirm` GitHub release is [`v0.9.1`](https://github.com/ContextualWisdomLab/fast-mlsirm/releases/tag/v0.9.1), published 2026-08-26. The `v0.9.1` tag resolves to source commit `09f762ded35786dd1078222a4577ff09d649816f`. Its tagged package source exports the provider-neutral `ContextualOrchestratorJudge`, `JudgeCriterion`, `JudgeFormatError`, `LLMJudgeResult`, and `validate_irt_response_matrix`; package metadata declares version `0.9.1` and Python `>=3.12`. The GitHub release currently has no attached distributable assets. -This is useful infrastructure, not proof that an email-writing rubric is valid. Naruon must still define the construct, criteria, category anchors, evaluation cases, language profiles, human reference process, calibration policy, and consequences of false positive or false negative guidance. +That distinction is material. The released source contract is sufficient to retire the obsolete `v0.6.0` source-surface assumption, but it is not yet sufficient for Naruon runtime import. Before Naruon consumes `fast-mlsirm`, the dependency gate must verify an approved immutable distributable package source, exact artifact integrity hash, provenance back to source commit `09f762ded35786dd1078222a4577ff09d649816f`, Python 3.14 compatibility, and Naruon's immutable hash lock. Mutable branches, Git URLs, source copies, local stubs, and workspace paths are not acceptable production dependencies. Package-unavailable behavior must remain testable by dependency injection rather than by assuming the package is absent from the environment. + +The verified public surface includes strict criterion-level structured output, explicit dichotomous/polytomous category semantics, deterministic response-matrix validation, and provider-neutral contextual-orchestrator injection. This is useful infrastructure, not proof that an email-writing rubric is valid. Naruon must still define the construct, criteria, category anchors, evaluation cases, language profiles, human reference process, calibration policy, and consequences of false-positive or false-negative guidance. + +Inkspan is a separate immutable dependency gate. Its current immutable GitHub release is `v0.3.1`; that release does not contain the open writing-diagnostics stack. Naruon therefore must not consume Inkspan writing diagnostics from an open branch or Draft PR. Runtime integration waits for a released public package subpath with version, integrity, source provenance, browser/package compatibility, and the revision-bound diagnostic contract required by this design. ## Why keyword matching is rejected From d8ed5f301a29ca3af898f856b5011d5727b6e798 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 1 Sep 2026 21:48:30 +0900 Subject: [PATCH 14/18] docs(email-writing): tighten policy publication evidence --- ...-writing-guidance-and-fast-mlsirm-judge.md | 61 ++++++++++++++++--- 1 file changed, 52 insertions(+), 9 deletions(-) diff --git a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md index 06b19899d..8bb318648 100644 --- a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md +++ b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md @@ -16,7 +16,7 @@ That distinction is material. The released source contract is sufficient to reti The verified public surface includes strict criterion-level structured output, explicit dichotomous/polytomous category semantics, deterministic response-matrix validation, and provider-neutral contextual-orchestrator injection. This is useful infrastructure, not proof that an email-writing rubric is valid. Naruon must still define the construct, criteria, category anchors, evaluation cases, language profiles, human reference process, calibration policy, and consequences of false-positive or false-negative guidance. -Inkspan is a separate immutable dependency gate. Its current immutable GitHub release is `v0.3.1`; that release does not contain the open writing-diagnostics stack. Naruon therefore must not consume Inkspan writing diagnostics from an open branch or Draft PR. Runtime integration waits for a released public package subpath with version, integrity, source provenance, browser/package compatibility, and the revision-bound diagnostic contract required by this design. +Inkspan is a separate immutable dependency gate. The current immutable release is [`v0.3.1`](https://github.com/ContextualWisdomLab/inkspan/releases/tag/v0.3.1) from `ContextualWisdomLab/inkspan`; tag `v0.3.1` resolves to source commit `67afc7099cc0e5711a9cc9476bf3be5bb820e229`, and the tagged npm manifest identifies `@contextualwisdomlab/cwl-editor` version `0.3.1`. The GitHub release contains no attached assets, so there is no release-asset SHA-256 to record or invent. Its public exports cover the editor root, collaboration, converter, styles, and fonts but not a `writing-diagnostics` subpath. Naruon therefore must not consume the open writing-diagnostics stack from a branch or Draft PR. Runtime integration waits for a future immutable package artifact that exposes the required writing-diagnostics public subpath and supplies registry/tarball integrity, source provenance, browser/package compatibility, and the revision-bound diagnostic contract required by this design. ## Why keyword matching is rejected @@ -155,14 +155,52 @@ The preregistered `minimum_slice_sample_size` is 30 labeled cases per slice. A slice below that count is reported as underpowered rather than selectively excluded after seeing its result. +The same protocol preregisters three additional release-risk gates before any +locked-holdout labels are accessed: + +- **DIF/fairness:** `maximum_large_dif_flags = 0`. A “large” flag must come from + the exact released fast-mlsirm DIF estimator and its method-specific, + preregistered effect-size classification. The protocol also caps the absolute + macro-F1 gap between every supported prespecified slice and the pooled + holdout at `0.10`, and the absolute mandatory-preservation-rate gap at `0.05`. + If the released estimator cannot produce a compatible effect-size + classification for a requested profile, that profile is withheld rather than + declared DIF-clean. +- **Temporal drift:** relative to the last published baseline on the same frozen + recurrent benchmark, macro-F1 may fall by at most `0.05` and expected + calibration error may increase by at most `0.02`. Crossing either bound + withholds the affected profile pending a new preregistered evaluation. +- **Critical consequences:** for cases whose adjudicated reference marks a + fact/actor/deadline/request-strength distortion as critical, + `maximum_observed_critical_consequence_errors = 0` and the one-sided 95% + exact-binomial upper confidence bound for the critical-consequence error rate + must be at most `0.05`. An underpowered critical-consequence set therefore + withholds publication rather than converting zero observed errors into a + safety claim. + +These numeric cutoffs are conservative Naruon product-admission tolerances, not +universal psychometric, fairness, or AI-safety constants. The AERA/APA/NCME +Standards support intended-use, fairness, reliability, and consequence evidence; +DIF literature supports combining statistical evidence with effect magnitude; +and NIST AI RMF guidance supports comparing production behavior with +pre-deployment metrics and monitoring drift. None of those sources mandates +these particular Naruon cutoffs. Any threshold change requires a new protocol +version before the new locked holdout is opened. + The policy artifact must contain `protocol_id`, `protocol_hash`, -`calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, the -literal thresholds, and `publish_decision` (`publish` or `withhold`). Changing -any split, threshold, or holdout requires a new protocol version and a new -calibration run. Thresholds must be frozen before the locked-holdout labels are -accessed. A run that tunes thresholds against the locked holdout is invalid for -publication and requires a new locked holdout; it is not a publication -decision. +`calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, +literal `minimum_slice_sample_size: 30`, every literal publication threshold and +decision rule above, a lifecycle `status`, and `publish_decision` (`publish` or +`withhold`). An `evaluation_only` artifact is represented as +`status: evaluation_only` plus `publish_decision: withhold`; the two fields are +orthogonal and the combination is enforced by schema. Every artifact consumer +must have access to the immutable protocol identified by `protocol_hash` and +must reject an artifact whose duplicated literal values disagree with that +protocol. Changing any split, threshold, protocol, or holdout requires a new +protocol version and a new calibration run. Thresholds must be frozen before +the locked-holdout labels are accessed. A run that tunes thresholds against the +locked holdout is invalid for publication and requires a new locked holdout; it +is not a publication decision. ## Multi-agent and compute allocation @@ -198,6 +236,7 @@ This design does not assert that every provider satisfies those requirements. Te ## Standards and governance alignment - **NIST AI 600-1** supports lifecycle risk identification, evaluation, monitoring, and governance for generative AI systems. +- **NIST AI 100-1** supports use-case-specific measurement and ongoing comparison of deployed behavior with pre-deployment evidence rather than universal thresholds. - **ISO/IEC 23894:2023** provides AI risk-management guidance. - **ISO/IEC 42001:2023** establishes an AI management-system framework for responsible development and use. - **W3C Web Annotation Data Model** defines `TextPositionSelector` semantics and resource-change limitations. @@ -211,6 +250,8 @@ Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, Chen, H., & Goldfarb-Tarrant, S. (2025). Safer or luckier? LLMs as safety evaluators are not robust to artifacts. In *Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)* (pp. 19750–19766). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.970 +Gómez-Benito, J., Hidalgo, M. D., & Zumbo, B. D. (2013). Effectiveness of combining statistical tests and effect sizes when using logistic discriminant function regression to detect differential item functioning for polytomous items. *Educational and Psychological Measurement, 73*(5), 875–897. https://doi.org/10.1177/0013164413492419 + Fu, X., & Liu, W. (2025). How reliable is multilingual LLM-as-a-Judge? In *Findings of the Association for Computational Linguistics: EMNLP 2025*. Association for Computational Linguistics. https://aclanthology.org/2025.findings-emnlp.587/ International Organization for Standardization. (2023a). *Information technology—Artificial intelligence—Guidance on risk management* (ISO/IEC Standard No. 23894:2023). https://www.iso.org/standard/77304.html @@ -223,6 +264,8 @@ Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG evalu Shen, C., Cheng, L., Nguyen, X.-P., You, Y., & Bing, L. (2023). Large language models are not yet human-level evaluators for abstractive summarization. In *Findings of the Association for Computational Linguistics: EMNLP 2023* (pp. 4215–4233). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.278 +Tabassi, E. (2023). *Artificial intelligence risk management framework (AI RMF 1.0)* (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1 + Usami, H., Hara, K., Tsuboi, A., & Matsuda, N. (2026). *LLM judges have dark current: A psychometric datasheet for LLM-as-a-Judge evaluation* [Preprint]. arXiv. https://arxiv.org/abs/2606.15610 Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Kong, L., Liu, Q., Liu, T., & Sui, Z. (2024). Large language models are not fair evaluators. In *Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)* (pp. 9440–9450). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.511 @@ -233,4 +276,4 @@ Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li ## Claim boundary -The cited evidence supports structured, bias-aware, calibrated LLM evaluation and revision-bound editor integrity. It does not prove that the planned Naruon rubric, model, provider, language profile, or fast-mlsirm configuration is sufficiently valid. Those claims require the benchmark, ablation, DIF, drift, security, privacy, and user-consequence evidence defined in the design. +The cited evidence supports structured, bias-aware, calibrated LLM evaluation and revision-bound editor integrity. It does not prove that the planned Naruon rubric, model, provider, language profile, or fast-mlsirm configuration is sufficiently valid. The numeric admission thresholds above are product policy, not literature-derived universal constants. Product validity claims require the benchmark, ablation, DIF, drift, security, privacy, and user-consequence evidence defined in the design. From 1e2498964f0294728b3ca8f3c23984b4b553bd67 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 1 Sep 2026 21:50:54 +0900 Subject: [PATCH 15/18] docs(email-writing): enforce preregistered review policy --- ...m-email-writing-guidance-implementation.md | 25 +++++++++++-------- 1 file changed, 14 insertions(+), 11 deletions(-) diff --git a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md index eab1d1f23..a10c7ef54 100644 --- a/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md +++ b/docs/superpowers/plans/2026-08-12-llm-email-writing-guidance-implementation.md @@ -279,8 +279,9 @@ git commit -m "feat(email-writing): parse contextual review candidates" - Create: `backend/tests/test_email_writing_judge.py` - Create: `backend/tests/fixtures/email_writing/judge_outputs.json` +- [ ] Write failing Judge-contract tests first for duplicate JSON keys, non-integral categories, reversed anchors, missing criterion IDs, extra criterion IDs, NaN/infinity, score/category disagreement, same-model policy violation, worker-lane saturation, cancellation, redacted errors, unavailable-package dependency injection, and mixed response-matrix semantics. - [ ] Import the released `ContextualOrchestratorJudge`, `JudgeCriterion`, `JudgeFormatError`, `LLMJudgeResult`, and response-matrix validator from fast-mlsirm. -- [ ] Define independently observable criteria with two-or-more-word snake_case IDs: +- [ ] Define independently observable criteria with stable snake_case IDs: - `issue_support`; - `span_fidelity`; - `replacement_correctness`; @@ -295,9 +296,9 @@ git commit -m "feat(email-writing): parse contextual review candidates" - [ ] Use explicit ordered categories and exact anchors. The initial category count remains an evaluation parameter until Task 14 ablation selects and publishes an approved value. - [ ] Pass bounded task, candidate answer, reference context, and rubric as untrusted data to fast-mlsirm; never parse free-form `pass`, `polite`, or `correct` tokens. - [ ] Derive final acceptance from fast-mlsirm's validated criterion result plus the external policy. Do not trust candidate confidence or the Judge's advisory boolean. -- [ ] Add tests for duplicate JSON keys, non-integral categories, reversed anchors, missing criterion IDs, extra criterion IDs, NaN/infinity, score/category disagreement, same-model policy violation, worker-lane saturation, cancellation, and redacted errors. -- [ ] Convert repeated Judge results into response rows and validate the matrix before any calibration export. +- [ ] Convert repeated Judge results into response rows and validate the matrix before any calibration export; reject mixed criterion identity/order/category semantics rather than coercing rows. - [ ] Verify no Judge raw output, task text, answer text, reference text, or full trace reaches ordinary logs or persistent product tables. +- [ ] Keep package-unavailable behavior testable through dependency injection; do not define the test by assuming fast-mlsirm is absent from the environment. - [ ] Run: ```bash @@ -321,9 +322,10 @@ git commit -m "feat(email-writing): add independent fast-mlsirm Judge" - Create: `backend/services/email_writing_policy.py` - Create: `backend/tests/test_email_writing_policy.py` -- [ ] Define a strict JSON Schema for policy ID/version, status, creation/expiry, compatible Naruon/Inkspan/fast-mlsirm/orchestrator contracts, approved model/provider/rubric/language profiles, category count/anchors, mandatory criterion floors, posterior/score admission rules, adjudication conditions, calibration/DIF summary, dataset/provenance hashes, limitations, and rollback version. -- [ ] Check in an integrity-bound `evaluation_only` policy that can exercise the pipeline but cannot emit user-facing diagnostics in production. -- [ ] Require runtime admission to assert `publish_decision == "publish"`; `withhold` and `evaluation_only` artifacts remain audit/pipeline evidence only and cannot produce user-facing diagnostics. +- [ ] Write failing policy-schema tests before implementation, including missing/invalid `publish_decision`, invalid lifecycle `status`, an `evaluation_only` artifact whose decision is not `withhold`, unknown fields, malformed hashes, incompatible profiles, and runtime attempts to admit non-published artifacts. +- [ ] Define a strict JSON Schema for policy ID/version, lifecycle `status`, required `publish_decision` constrained exactly to `publish | withhold`, creation/expiry, compatible Naruon/Inkspan/fast-mlsirm/orchestrator contracts, approved model/provider/rubric/language profiles, category count/anchors, mandatory criterion floors, posterior/score admission rules, adjudication conditions, calibration/DIF summary, dataset/provenance hashes, limitations, and rollback version. +- [ ] Check in an integrity-bound `evaluation_only` policy with `status: evaluation_only` and `publish_decision: withhold`; it can exercise the pipeline but cannot emit user-facing diagnostics in production. Keep lifecycle status and publication decision as separate fields and enforce their legal combinations. +- [ ] Require runtime admission to assert `publish_decision == "publish"` and an otherwise admissible published lifecycle status; `withhold`, `evaluation_only`, superseded, revoked, and incompatible artifacts remain audit/pipeline evidence only and cannot produce user-facing diagnostics. - [ ] Add a manifest containing the literal SHA-256 of every allowed policy artifact; reject modified, unknown, expired, future-dated, revoked, incompatible, or non-approved artifacts. - [ ] Implement outcomes: @@ -339,7 +341,7 @@ policy_unavailable - [ ] Require an approved profile for language, model/provider pair, candidate/Judge separation, review mode, and rubric version. - [ ] Do not hard-code unvalidated production thresholds in Python. Approved threshold values enter only through a Task 14 policy artifact with calibration evidence. - [ ] Add rollback tests proving a revoked current policy can atomically fall back only to a compatible, non-expired, explicitly listed rollback policy; otherwise semantic review disables. -- [ ] Add policy tests for field smuggling, signature/hash mismatch, downgrade, category-anchor reversal, impossible floors, hidden keyword tables, and unexpected executable content. +- [ ] Add policy tests for field smuggling, signature/hash mismatch, downgrade, category-anchor reversal, impossible floors, hidden keyword tables, unexpected executable content, and `evaluation_only` preservation without user-facing admission. - [ ] Run: ```bash @@ -560,7 +562,8 @@ git commit -m "feat(email-writing): connect guidance actions and evidence" - Create after passing evidence: `backend/policies/email_writing_judge_approved_v1.json` - [ ] Define a consented/de-identified or synthetic case schema with source context, recipient roles, draft, gold spans, categories, acceptable replacements, preservation constraints, annotator IDs, adjudication, and consequence severity. -- [ ] Before calibration, publish the canonical `email_writing_policy_protocol_v1` with a SHA-256, fixed contrast-family splits of 60% calibration, 20% development, and 20% locked human-labeled holdout, a fixed `minimum_slice_sample_size` of 30 cases per prespecified slice, plus the literal publication gates: holdout macro-F1 >= 0.80, every mandatory preservation criterion >= 0.90, unsupported-claim rate <= 0.02, expected calibration error <= 0.05, and no prespecified slice below macro-F1 0.70 when its minimum sample size is met. +- [ ] Before calibration, publish the canonical `email_writing_policy_protocol_v1` with a SHA-256, fixed contrast-family splits of 60% calibration, 20% development, and 20% locked human-labeled holdout, and fixed `minimum_slice_sample_size: 30`. Record literal initial publication gates: holdout macro-F1 >= 0.80; every mandatory preservation criterion >= 0.90; unsupported-claim rate <= 0.02; expected calibration error <= 0.05; no supported prespecified slice below macro-F1 0.70; `maximum_large_dif_flags = 0` using the released fast-mlsirm estimator's preregistered method-specific effect classification; supported-slice versus pooled-holdout absolute macro-F1 gap <= 0.10 and mandatory-preservation-rate gap <= 0.05; temporal macro-F1 drop <= 0.05 and ECE increase <= 0.02 versus the last published baseline on the same frozen recurrent benchmark; `maximum_observed_critical_consequence_errors = 0`; and a one-sided 95% exact-binomial upper bound for the critical-consequence error rate <= 0.05. Treat these numbers as conservative Naruon product-admission tolerances, not universal scientific constants. +- [ ] If the released fast-mlsirm DIF estimator cannot provide a compatible method-specific effect classification for a requested profile, mark that profile's DIF evidence insufficient and withhold it rather than substituting a Naruon lexical/statistical shortcut. - [ ] Include contrast families: - same words in quotation, neutral incident report, and direct rebuke; - same pragmatic issue expressed through unrelated paraphrases; @@ -572,13 +575,13 @@ git commit -m "feat(email-writing): connect guidance actions and evidence" - Korean, English, mixed-language, CJK, emoji, combining-mark, and hostile-Unicode cases. - [ ] Collect independent human annotations and report inter-rater agreement plus adjudicated disagreement. Human labels are reference evidence, not unquestioned truth. - [ ] Freeze the holdout manifest before any calibration or threshold selection; record calibration, development, locked-holdout, and protocol hashes in every result and keep the holdout read-only until the final publish/no-publish decision. -- [ ] Measure issue/category precision, recall, macro-F1, span IoU/exactness, replacement correctness, intent/fact/actor/deadline/request-strength preservation, unsupported-claim rate, accepted-suggestion precision, Brier score, calibration error, and human ignore rate. +- [ ] Measure issue/category precision, recall, macro-F1, span IoU/exactness, replacement correctness, intent/fact/actor/deadline/request-strength preservation, unsupported-claim rate, accepted-suggestion precision, Brier score, calibration error, human ignore rate, preregistered supported-slice gaps, critical-consequence errors and confidence bounds. - [ ] Use fast-mlsirm response matrices to analyze criterion/category behavior, item/rater effects where supported, reliability, prompt/model test-retest, category sparsity, category-count ablation, and DIF by language, recipient configuration, role/hierarchy, thread depth, document length, review mode, model, prompt, and time. - [ ] Compare single-model, separated candidate/Judge, and deeper multi-agent/adjudicator workflows with role-specific reasoning effort. Speed is recorded but does not override validity. - [ ] Use deterministic offline fixtures in required CI. Run live scheduled/manual evaluation with `NVIDIA_NIM_API_KEY`; do not expose the secret to pull-request code from untrusted forks and do not use `COPILOT_GITHUB_TOKEN`. -- [ ] Fail the policy publication job when sample size, calibration, preservation, DIF, drift, or consequence thresholds are not met. Do not publish a policy merely because average accuracy is high. +- [ ] Fail the policy publication job when sample size, calibration, preservation, DIF, supported-slice gap, drift, or consequence gates are not met. Do not publish a policy merely because average accuracy is high. - [ ] Generate the approved policy artifact from measured outputs, review its literal values, add its SHA-256 to the manifest, and prove production rejects the earlier `evaluation_only` artifact for user-facing output. -- [ ] Include `protocol_id`, `protocol_hash`, `calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, literal `minimum_slice_sample_size: 30`, literal thresholds, and `publish_decision` (`publish` or `withhold`) in the policy artifact. Freeze thresholds before reading locked-holdout labels; a run that tunes thresholds against that holdout is invalid for publication and requires a new locked holdout. Any protocol, split, threshold, or holdout change creates a new protocol version and a new run. +- [ ] Include `protocol_id`, `protocol_hash`, `calibration_split_hash`, `development_split_hash`, `locked_holdout_hash`, literal `minimum_slice_sample_size: 30`, every literal threshold/decision rule, lifecycle `status`, and required `publish_decision` (`publish` or `withhold`) in the policy artifact. Represent evaluation-only evidence as `status: evaluation_only` plus `publish_decision: withhold`. Require every duplicated literal value to match the immutable protocol identified by `protocol_hash`. Freeze thresholds before reading locked-holdout labels; a run that tunes thresholds against that holdout is invalid for publication and requires a new locked holdout. Any protocol, split, threshold, or holdout change creates a new protocol version and a new run. - [ ] Store reports, response matrices, aggregate metrics, manifests, and source hashes without confidential raw company email. - [ ] Run: From 5cb21487ccdf292681952d1d574c86b6efd91b9f Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 1 Sep 2026 21:52:15 +0900 Subject: [PATCH 16/18] docs(email-writing): align evaluation-only policy contract --- ...kspan-llm-email-writing-guidance-design.md | 34 +++++++++++++------ 1 file changed, 23 insertions(+), 11 deletions(-) diff --git a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md index 387c4eb65..e0bfd21f8 100644 --- a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md +++ b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md @@ -316,6 +316,7 @@ Collect ordered category responses across criteria and, where feasible, repeated A calibration run emits an integrity-bound `judge_policy_artifact` containing: - version, creation time, expiry, and rollback version; +- lifecycle `status`; - compatible fast-mlsirm, Naruon, Inkspan, and orchestrator contract versions; - approved model/provider/rubric/language profiles; - category anchors and admission rules; @@ -327,19 +328,29 @@ A calibration run emits an integrity-bound `judge_policy_artifact` containing: publish/no-publish decision; - a fixed `minimum_slice_sample_size` of 30 labeled cases per prespecified slice; -- literal publication thresholds for holdout macro-F1, mandatory preservation, - unsupported-claim rate, expected calibration error, and prespecified slice - performance; -- a `publish_decision` of `publish` or `withhold`; +- literal publication thresholds and decision rules for holdout macro-F1, + mandatory preservation, unsupported-claim rate, expected calibration error, + prespecified slice performance, DIF/fairness gaps, temporal drift, and + critical consequences; +- a required `publish_decision` of `publish` or `withhold`; - known limitations. -Naruon runtime consumes the artifact but does not refit it, and admits it for -user-facing diagnostics only when `publish_decision == "publish"`. `withhold` -and `evaluation_only` artifacts remain audit/pipeline evidence and cannot -produce user-facing output. Changing a split, threshold, protocol, or holdout -requires a new protocol version and a new artifact. Thresholds are frozen before -locked-holdout labels are accessed; tuning against that holdout invalidates the -run for publication and requires a new locked holdout. +Lifecycle and publication decision are separate dimensions. An evaluation-only +artifact is represented exactly as `status: evaluation_only` and +`publish_decision: withhold`. Schema validation rejects +`status: evaluation_only` with `publish_decision: publish`, and the runtime +rejects every artifact from user-facing output unless its lifecycle status is +admissible for publication and `publish_decision == "publish"`. `withhold`, +`evaluation_only`, superseded, revoked, and incompatible artifacts remain +available only as bounded audit/pipeline evidence. + +Every producer and consumer validates the duplicated literal policy fields +against the immutable protocol named by `protocol_hash`, including +`minimum_slice_sample_size: 30` and all threshold/decision-rule values. Changing +a split, threshold, protocol, or holdout requires a new protocol version and a +new artifact. Thresholds are frozen before locked-holdout labels are accessed; +tuning against that holdout invalidates the run for publication and requires a +new locked holdout. ## Benchmark design @@ -380,6 +391,7 @@ Use consented, de-identified, or synthetic reconstruction cases. Do not persist - human inter-rater agreement; - Brier score, expected calibration error, and reliability curves; - category occupancy, DIF, and drift; +- supported-slice performance gaps and critical-consequence confidence bounds; - latency, token, step, and cost distributions; - stale/conflict rates; - accessibility task success. From 8c5566c0e62ab3fe4d7b6a329621e6fd9e93b2de Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 1 Sep 2026 23:16:57 +0900 Subject: [PATCH 17/18] docs(email-writing): reconcile source identity and persistence contract --- ...kspan-llm-email-writing-guidance-design.md | 30 +++++++++---------- 1 file changed, 15 insertions(+), 15 deletions(-) diff --git a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md index e0bfd21f8..c59ea0aba 100644 --- a/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md +++ b/docs/superpowers/specs/2026-08-12-inkspan-llm-email-writing-guidance-design.md @@ -117,7 +117,7 @@ Creates a bounded review for one exact editor revision. ```json { - "source_email_id": 184, + "source_email_id": "", "document_revision": "\"sha256-BASE64URL\"", "projection_name": "inkspan-prosemirror-text", "projection_version": 1, @@ -135,7 +135,7 @@ Creates a bounded review for one exact editor revision. Rules: -- `source_email_id` is scoped and re-read server-side; +- `source_email_id` is the opaque source-message identifier already exposed by Naruon mail contracts (the message `message_id`), never the sequential/internal database `email_id`; it is owner/tenant scoped and re-read server-side; - revision, projection, and draft describe one exact Inkspan snapshot; - `changed_selector` is required in incremental mode and omitted in deep mode; - `reply_objective` is untrusted user guidance, not a policy override; @@ -400,20 +400,19 @@ No scalar “email risk score” is a release criterion. ## Data model -Phase 1 may keep review state ephemeral. Persistent objects, if later required, use two-or-more-word `snake_case` names: +The first release persists exactly three privacy-minimized, feature-owned evidence tables through migration `20260812_0001_add_email_writing_review_evidence.py`: ```text email_review_session writing_diagnostic_record diagnostic_feedback_event -judge_policy_artifact -judge_calibration_run -review_model_profile -review_rubric_version -review_evaluation_case ``` -Raw source email or draft text is not duplicated by default. Persisted diagnostic content requires encryption, retention, access, and deletion policy and never becomes canonical email content. +`email_review_session` owns the review scope, opaque source-message reference, reviewed revision/projection, policy/workflow/rubric provenance, status, bounded timing/cost buckets, and expiry. `writing_diagnostic_record` owns selector coordinates, hashes, criterion categories, admission metadata, and the session relationship without replacement or explanation plaintext. `diagnostic_feedback_event` owns explicit author feedback, revision references, stale/conflict reason, event time, and an idempotency key. Session and admitted/withheld diagnostic evidence are committed atomically; a persistence failure rolls back the evidence transaction and returns a typed failure or abstention without disabling the feature or changing schema. + +The default evidence-retention period is 30 days. Owner-scoped expiry/deletion removes a session and its dependent diagnostic and feedback rows without scanning, rewriting, or deleting canonical email content. Raw source email, draft, replacement, explanation, prompt, model/Judge output, provider credentials, and complete traces are not stored in these tables. PostgreSQL and the repository's SQLite compatibility path must both prove upgrade and downgrade behavior. Every table, column, index, constraint, and relationship uses descriptive two-or-more-word `snake_case` names. + +Later calibration or registry persistence such as `judge_policy_artifact`, `judge_calibration_run`, `review_model_profile`, `review_rubric_version`, or `review_evaluation_case` requires its own explicit migration/retention contract before it becomes durable state; naming those concepts here does not authorize an unversioned first-release table. ## Privacy and PII controls @@ -560,12 +559,13 @@ No mutable branch or source archive becomes a production dependency. ## Rollback -- disable semantic review by tenant/product configuration; -- preserve Inkspan authoring, drafts, and sending; -- stop review requests; -- retain only policy-required bounded audit evidence; -- optionally retain the separately described one-shot drafting path without claiming contextual inline review; -- require no canonical email or database migration to remove ephemeral state. +- disable semantic review by tenant/product configuration before removing any evidence contract; +- preserve Inkspan authoring, existing drafts, and the mail-send path; +- stop new review requests and writes to the three feature-owned evidence tables; +- delete dependent `diagnostic_feedback_event` and `writing_diagnostic_record` rows before their owning `email_review_session` rows under the authorized retention/deletion operation; +- run the documented downgrade only after semantic review is disabled and dependent feature evidence is removed; downgrade drops only `email_review_session`, `writing_diagnostic_record`, and `diagnostic_feedback_event` and never alters canonical email tables or mail content; +- treat an individual evidence-transaction failure as a typed review failure/abstention only; it never implicitly disables the feature or triggers schema downgrade; +- optionally retain the separately described one-shot drafting path without claiming contextual inline review. ## Documentation and traceability required during implementation From 9f1836d09e6b4db97855d701d8220268e9fb4d87 Mon Sep 17 00:00:00 2001 From: Seongho Bae Date: Tue, 1 Sep 2026 23:18:16 +0900 Subject: [PATCH 18/18] docs(email-writing): seal locked holdout until final evaluation --- ...il-writing-guidance-and-fast-mlsirm-judge.md | 17 ++++++++++------- 1 file changed, 10 insertions(+), 7 deletions(-) diff --git a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md index 8bb318648..aed3e491f 100644 --- a/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md +++ b/docs/doctoring/llm-email-writing-guidance-and-fast-mlsirm-judge.md @@ -143,13 +143,16 @@ No overall “email danger” or “tone risk” score substitutes for the crite Before calibration or threshold selection, the evaluation runner writes a canonical `email_writing_policy_protocol_v1` document and records its SHA-256. The split is fixed by contrast family, not by individual rows: 60% calibration, -20% development, and a 20% human-labeled locked holdout. The holdout is read -only until the final publish/no-publish decision and its case-manifest SHA-256 is -stored beside the protocol hash. The initial publication gates are fixed in the -protocol: holdout macro-F1 at least 0.80, every mandatory preservation criterion -at least 0.90, unsupported-claim rate at most 0.02, expected calibration error -at most 0.05, and no prespecified language/role/context slice with macro-F1 -below 0.70 when that slice has the preregistered minimum sample size. +20% development, and a 20% human-labeled locked holdout. Locked-holdout labels +remain sealed and inaccessible to calibration, development, threshold selection, +and policy fitting; they are opened only for the preregistered final +publish/no-publish evaluation after all thresholds and decision rules are frozen. +The holdout case-manifest SHA-256 is stored beside the protocol hash before that +label access. The initial publication gates are fixed in the protocol: holdout +macro-F1 at least 0.80, every mandatory preservation criterion at least 0.90, +unsupported-claim rate at most 0.02, expected calibration error at most 0.05, +and no prespecified language/role/context slice with macro-F1 below 0.70 when +that slice has the preregistered minimum sample size. The preregistered `minimum_slice_sample_size` is 30 labeled cases per slice. A slice below that count is reported as underpowered rather than selectively