diff --git a/CHANGELOG.md b/CHANGELOG.md index 926b0fd0c..3dc91f640 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,12 @@ - Added fail-closed shape, finite-value, likelihood-trace, parameter-count, connectedness, and design-fingerprint replay checks plus non-suppressible human-review routing for non-converged or disconnected fits. - Added complete public documentation, APA 7th equation-to-source traceability, privacy guarantees, deterministic fixtures, Rust-delegation tests, and statement/branch coverage for the new reporting boundary. +#### Accessible standalone essay score report artifacts + +- Added `render_essay_score_report_html`, which replay-verifies one governed `EssayScoreReport` and emits a deterministic, source-text-free, script-free standalone HTML audit artifact. +- The artifact exposes exact report, assessment, rubric, task-revision, engine, request, result, observation, criterion, trigger, and evidence-reference identities through semantic landmarks, keyboard-accessible exact-value tables, and canonical JSON. +- A restrictive meta-delivered Content Security Policy and output encoding reduce content-injection impact. Review routing remains an audit signal only and does not establish scoring validity, fairness, reliability, interchangeability, accessibility conformance, security certification, or authorization for consequential deployment. + #### Provenance-bound essay score reports - Added a provider-neutral, content-addressed `EssayScoreReport` adapter over the existing governed essay request, shared scoring result, and engine descriptor. diff --git a/docs/automated_essay_score_reports.md b/docs/automated_essay_score_reports.md index b94ad9e2f..3e231c07b 100644 --- a/docs/automated_essay_score_reports.md +++ b/docs/automated_essay_score_reports.md @@ -6,7 +6,9 @@ `EssayScoreReport` from an exact governed essay request, scoring result, and engine descriptor. The adapter is a reporting and human-review routing boundary. It does not score essays, combine analytic criteria, estimate psychometric -parameters, or establish validity. +parameters, or establish validity. Criterion-level many-facet estimator output +uses the separate `EssayFacetsCalibrationReport` boundary and is not embedded in +this operational score report. ```python from fast_mlsirm.scoring.essay import build_essay_score_report @@ -29,6 +31,40 @@ The report embeds the canonical serialized forms and fingerprints of: - the complete shared scoring result and criterion observations; and - transparent review triggers and audit metadata. +## Standalone HTML audit artifact + +The same governed report can be rendered as a self-contained, source-text-free +HTML artifact: + +```python +from fast_mlsirm.scoring.essay import render_essay_score_report_html + +render_essay_score_report_html( + report, + "artifacts/essay_score_report.html", + title="Pilot Essay Score Audit", +) +``` + +The renderer first rebuilds the report through the governed request, observation, +result, and report factories and rejects canonical replay differences. The HTML +contains exact report, assessment, rubric, task-revision, engine, request, result, +observation, criterion, trigger, and evidence-reference identities. It does not +contain prompt text, response text, or source text. + +The artifact uses semantic landmarks, a keyboard-accessible skip link, labelled +exact-value tables, focusable overflow regions, and a focusable canonical JSON +section. It has no script or external resource dependency. A restrictive +meta-delivered Content Security Policy is included as defense in depth; output +encoding and governed content validation remain required because CSP is not a +replacement for either control. + +These implementation choices are informed by WCAG 2.2 and the current Content +Security Policy Level 3 working draft. They support audit usability and reduce +content-injection impact, but they do not constitute a blanket conformance or +security certification. Accessibility conformance still requires full-page, +assistive-technology, and human evaluation in the buyer's deployment context. + ## Non-suppressible review triggers The builder derives structural triggers that callers cannot remove: @@ -84,3 +120,9 @@ handbook of automated essay evaluation*. Routledge. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. *Educational Measurement: Issues and Practice, 31*(1), 2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x + +World Wide Web Consortium. (2023, October 5). *Web Content Accessibility +Guidelines (WCAG) 2.2*. https://www.w3.org/TR/WCAG22/ + +World Wide Web Consortium. (2026, May 5). *Content Security Policy Level 3* +(W3C Working Draft). https://www.w3.org/TR/CSP3/ diff --git a/docs/changelog.d/essay-score-report-html.md b/docs/changelog.d/essay-score-report-html.md new file mode 100644 index 000000000..571f3f7bb --- /dev/null +++ b/docs/changelog.d/essay-score-report-html.md @@ -0,0 +1,7 @@ +# Accessible standalone essay score report artifacts + +## Added + +- Added `render_essay_score_report_html`, which replay-verifies one governed `EssayScoreReport` and emits a deterministic, source-text-free, script-free standalone HTML audit artifact. +- The artifact exposes exact report, assessment, rubric, task-revision, engine, request, result, observation, criterion, trigger, and evidence-reference identities through semantic landmarks, keyboard-accessible exact-value tables, and canonical JSON. +- A restrictive meta-delivered Content Security Policy and output encoding reduce content-injection impact. Review routing remains an audit signal only and does not establish scoring validity, fairness, reliability, interchangeability, accessibility conformance, security certification, or authorization for consequential deployment. diff --git a/python/fast_mlsirm/scoring/essay/__init__.py b/python/fast_mlsirm/scoring/essay/__init__.py index 1bdb05cef..38bd1191f 100644 --- a/python/fast_mlsirm/scoring/essay/__init__.py +++ b/python/fast_mlsirm/scoring/essay/__init__.py @@ -33,6 +33,9 @@ from .contracts import build_essay_scoring_request as build_essay_scoring_request from .contracts import build_essay_submission as build_essay_submission from .contracts import score_essay_request as score_essay_request +from .report_html import ( + render_essay_score_report_html as render_essay_score_report_html, +) from .reporting import EssayScoreReport as EssayScoreReport from .reporting import ( MAX_ESSAY_REPORT_REVIEW_TRIGGERS as MAX_ESSAY_REPORT_REVIEW_TRIGGERS, @@ -61,5 +64,6 @@ "build_essay_scoring_request", "build_essay_submission", "fit_essay_facets_calibration_report", + "render_essay_score_report_html", "score_essay_request", ] diff --git a/python/fast_mlsirm/scoring/essay/report_html.py b/python/fast_mlsirm/scoring/essay/report_html.py new file mode 100644 index 000000000..469ed247d --- /dev/null +++ b/python/fast_mlsirm/scoring/essay/report_html.py @@ -0,0 +1,320 @@ +"""Accessible standalone HTML rendering for governed essay score reports. + +The renderer emits a source-text-free, script-free audit artifact from one exact +:class:`~fast_mlsirm.scoring.essay.reporting.EssayScoreReport`. It performs no +scoring, aggregation, psychometric estimation, validity inference, or deployment +authorization. +""" + +from __future__ import annotations + +import json +from html import escape +from pathlib import Path + +from .._validation import assessment_error +from .reporting import EssayScoreReport, build_essay_score_report + +_DEFAULT_TITLE = "Governed Automated-Essay Score Report" +_VALIDITY_NOTICE = ( + "Review routing is an audit signal only. Absence of a trigger is not " + "evidence of scoring validity, fairness, reliability, interchangeability, " + "or authorization for consequential deployment." +) + + +def _validated_report(report: EssayScoreReport) -> EssayScoreReport: + """Replay one report through governed factories before serialization.""" + if not isinstance(report, EssayScoreReport): + raise assessment_error( + "invalid_essay_score_report", + "$.report", + "report must be an EssayScoreReport", + ) + replayed = build_essay_score_report( + report_id=report.report_id, + request=report.essay_request, + result=report.scoring_result, + engine=report.engine_descriptor, + additional_review_trigger_ids=report.review_trigger_ids, + metadata=report.metadata, + ) + if replayed.report_fingerprint != report.report_fingerprint: + raise assessment_error( + "essay_score_report_replay_mismatch", + "$.report", + "report content does not match a freshly validated report", + ) + return replayed + + +def _content_security_policy() -> str: + """Return a restrictive meta-delivered policy for the standalone artifact.""" + return "; ".join( + ( + "default-src 'none'", + "style-src 'unsafe-inline'", + "img-src data:", + "object-src 'none'", + "base-uri 'none'", + "form-action 'none'", + ) + ) + + +def _display(value: object | None) -> str: + """Return one escaped exact display value with an explicit missing marker.""" + return "Not applicable" if value is None else escape(str(value)) + + +def _definition_rows(rows: tuple[tuple[str, object], ...]) -> str: + """Render exact key-value provenance as a semantic definition list.""" + items = [] + for label, value in rows: + items.extend((f"
{escape(label)}
", f"
{_display(value)}
")) + return "\n".join(("
", *items, "
")) + + +def _table( + *, + caption: str, + headers: tuple[str, ...], + rows: tuple[tuple[object | None, ...], ...], + empty_message: str, +) -> str: + """Render an accessible exact-value table or one explicit empty state.""" + if not rows: + return f'

{escape(empty_message)}

' + heading = "".join(f'{escape(header)}' for header in headers) + body = [] + for row in rows: + cells = "".join(f"{_display(value)}" for value in row) + body.append(f"{cells}") + return "\n".join( + ( + '
', + "", + f"", + f"{heading}", + f"{''.join(body)}", + "
{escape(caption)}
", + "
", + ) + ) + + +def _criterion_rows(report: EssayScoreReport) -> tuple[tuple[object | None, ...], ...]: + """Return criterion outcomes without averaging or interpreting scores.""" + return tuple( + ( + observation.criterion_id, + observation.status.value, + observation.score_category, + observation.reason_code, + len(observation.evidence_references), + observation.observation_fingerprint, + ) + for observation in report.scoring_result.observations + ) + + +def _evidence_rows(report: EssayScoreReport) -> tuple[tuple[object | None, ...], ...]: + """Return source-text-free evidence identities for every observation.""" + return tuple( + ( + observation.criterion_id, + evidence.source_id, + evidence.span_id, + evidence.evidence_role.value, + evidence.content_fingerprint, + evidence.evidence_fingerprint, + ) + for observation in report.scoring_result.observations + for evidence in observation.evidence_references + ) + + +def _trigger_section(report: EssayScoreReport) -> str: + """Render every transparent review trigger or an explicit empty state.""" + if not report.review_trigger_ids: + return '

No structural review trigger was emitted.

' + items = "".join( + f"
  • {escape(trigger_id)}
  • " + for trigger_id in report.review_trigger_ids + ) + return f'' + + +def _canonical_json(report: EssayScoreReport) -> str: + """Return escaped deterministic JSON for exact audit reconstruction.""" + serialized = json.dumps( + report.to_dict(), + ensure_ascii=False, + indent=2, + sort_keys=True, + allow_nan=False, + ) + return escape(serialized) + + +def _css() -> str: + """Return compact accessible styling without external resources.""" + return """ +:root { color-scheme: light dark; font-family: system-ui, sans-serif; } +* { box-sizing: border-box; } +body { margin: 0; background: Canvas; color: CanvasText; } +main { width: min(1120px, calc(100% - 32px)); margin: 0 auto 48px; } +.skip-link { position: absolute; left: 8px; top: -80px; padding: 10px; background: Canvas; color: CanvasText; z-index: 10; } +.skip-link:focus { top: 8px; } +.hero { padding: 48px 0 24px; } +h1 { margin: 0 0 8px; font-size: clamp(2rem, 5vw, 3.2rem); } +.subtitle { margin: 0; max-width: 78ch; } +section { margin-top: 20px; padding: 20px; border: 1px solid GrayText; border-radius: 10px; } +.review-required { border-inline-start: 8px solid #9c2f1f; } +.review-clear { border-inline-start: 8px solid #357a38; } +.notice { padding: 14px; border: 2px solid currentColor; font-weight: 650; } +.details-grid { display: grid; grid-template-columns: minmax(150px, 0.35fr) 1fr; gap: 8px 16px; } +.details-grid dt { font-weight: 700; } +.details-grid dd { margin: 0; overflow-wrap: anywhere; } +.table-scroll { overflow-x: auto; } +.table-scroll:focus-visible, pre:focus-visible { outline: 3px solid Highlight; outline-offset: 3px; } +table { width: 100%; border-collapse: collapse; } +caption { text-align: left; font-weight: 700; margin-bottom: 8px; } +th, td { padding: 10px; border: 1px solid GrayText; text-align: left; vertical-align: top; overflow-wrap: anywhere; } +code, pre { font-family: ui-monospace, monospace; } +pre { max-height: 32rem; overflow: auto; padding: 16px; border: 1px solid GrayText; white-space: pre-wrap; overflow-wrap: anywhere; } +.empty-state { font-style: italic; } +@media (max-width: 640px) { .details-grid { grid-template-columns: 1fr; } .details-grid dd { margin-bottom: 8px; } } +""".strip() + + +def _render_html(report: EssayScoreReport, title: str) -> str: + """Assemble one complete script-free report document.""" + engine = report.engine_descriptor + request = report.essay_request.scoring_request + review_class = "review-required" if report.human_review_required else "review-clear" + review_label = "Human review required" if report.human_review_required else "No structural trigger" + provenance = _definition_rows( + ( + ("Report ID", report.report_id), + ("Report handle", report.report_handle), + ("Report fingerprint", report.report_fingerprint), + ("Schema version", report.schema_version), + ("Request fingerprint", request.request_fingerprint), + ("Result fingerprint", report.scoring_result.result_fingerprint), + ("Assessment fingerprint", request.assessment_fingerprint), + ("Rubric fingerprint", request.rubric_fingerprint), + ("Task revision fingerprint", request.task_revision_fingerprint), + ("Engine ID", engine.engine_id), + ("Engine family", engine.engine_family_id), + ("Engine version", engine.engine_version), + ("Engine fingerprint", engine.engine_fingerprint), + ) + ) + criteria = _table( + caption="Criterion-level scoring outcomes", + headers=( + "Criterion", + "Status", + "Score category", + "Reason code", + "Evidence count", + "Observation fingerprint", + ), + rows=_criterion_rows(report), + empty_message="No criterion observations are available.", + ) + evidence = _table( + caption="Source-text-free evidence references", + headers=( + "Criterion", + "Source ID", + "Span ID", + "Role", + "Content fingerprint", + "Evidence fingerprint", + ), + rows=_evidence_rows(report), + empty_message="No evidence references are attached to this report.", + ) + return "\n".join( + ( + "", + '', + "", + '', + '', + '', + f"{escape(title)}", + f"", + "", + "", + '', + '
    ', + '
    ', + f"

    {escape(title)}

    ", + '

    Exact governed scoring provenance and transparent review routing without source text.

    ', + "
    ", + f'
    ', + '

    Review routing

    ', + f"

    {escape(review_label)}

    ", + f'

    {escape(_VALIDITY_NOTICE)}

    ', + _trigger_section(report), + "
    ", + '
    ', + '

    Exact provenance

    ', + provenance, + "
    ", + '
    ', + '

    Criterion outcomes

    ', + criteria, + "
    ", + '
    ', + '

    Evidence references

    ', + evidence, + "
    ", + '
    ', + '

    Canonical JSON

    ', + '

    The complete deterministic report payload is available below for audit reconstruction.

    ', + '
    ',
    +            _canonical_json(report),
    +            "
    ", + "
    ", + "
    ", + "", + "", + ) + ) + + +def render_essay_score_report_html( + report: EssayScoreReport, + output_path: str | Path, + *, + title: str | None = None, +) -> Path: + """Write one replay-verified, accessible standalone HTML audit report. + + The artifact contains exact criterion, evidence, engine, assessment, rubric, + request, result, and report provenance but no prompt, essay, or source text. + Its review state is not a validity or deployment decision. + """ + validated = _validated_report(report) + output = Path(output_path) + if output.suffix.lower() != ".html": + raise ValueError("essay score report output path must end with .html") + if title is not None and ( + not isinstance(title, str) or not title.strip() + ): + raise ValueError( + "essay score report title must be a non-empty string" + ) + resolved_title = _DEFAULT_TITLE if title is None else title + output.parent.mkdir(parents=True, exist_ok=True) + output.write_text(_render_html(validated, resolved_title), encoding="utf-8") + return output + + +__all__ = ["render_essay_score_report_html"] diff --git a/tests/test_scoring_essay_contracts.py b/tests/test_scoring_essay_contracts.py index dd4df322a..04948cdfd 100644 --- a/tests/test_scoring_essay_contracts.py +++ b/tests/test_scoring_essay_contracts.py @@ -190,6 +190,7 @@ def test_public_surface_is_explicit_and_documented() -> None: "build_essay_scoring_request", "build_essay_submission", "fit_essay_facets_calibration_report", + "render_essay_score_report_html", "score_essay_request", } assert set(essay.__all__) == expected diff --git a/tests/test_scoring_essay_report_html.py b/tests/test_scoring_essay_report_html.py new file mode 100644 index 000000000..568dbb5ae --- /dev/null +++ b/tests/test_scoring_essay_report_html.py @@ -0,0 +1,184 @@ +"""Tests for accessible standalone governed essay score reports.""" + +from __future__ import annotations + +import runpy +from pathlib import Path + +import pytest + +import fast_mlsirm.scoring.essay.report_html as report_html +from fast_mlsirm.scoring import AssessmentSpecError, ObservationStatus +from fast_mlsirm.scoring.essay import ( + EssayReviewFlag, + build_essay_score_report, + render_essay_score_report_html, +) + +_REPORT_FIXTURES = runpy.run_path( + str(Path(__file__).with_name("test_scoring_essay_reporting.py")) +) +essay_request = _REPORT_FIXTURES["essay_request"] +result_bundle = _REPORT_FIXTURES["result_bundle"] + + +def clean_report(): + """Return one deterministic report without structural review triggers.""" + request = essay_request() + _engine, descriptor, result = result_bundle(request) + return build_essay_score_report( + report_id="essay_score_report", + request=request, + result=result, + engine=descriptor, + metadata={"workflow_stage": "pilot_review"}, + ) + + +def review_report(): + """Return one deterministic report with mandatory and added triggers.""" + request = essay_request( + review_flags=(EssayReviewFlag.OFF_TOPIC_RESPONSE,) + ) + _engine, descriptor, result = result_bundle( + request, + claim_evidence=(), + alignment_status=ObservationStatus.ABSTAINED, + alignment_reason="insufficient_evidence", + ) + return build_essay_score_report( + report_id="review_required_report", + request=request, + result=result, + engine=descriptor, + additional_review_trigger_ids=("scorer_disagreement",), + ) + + +def test_html_renderer_surface_is_explicit_and_documented() -> None: + """The renderer module exposes one documented reporting operation.""" + assert report_html.__all__ == ["render_essay_score_report_html"] + assert render_essay_score_report_html.__doc__ + + +def test_clean_report_renders_deterministic_accessible_exact_values( + tmp_path: Path, +) -> None: + """A clean report remains exact, script-free, accessible, and deterministic.""" + report = clean_report() + first_path = tmp_path / "nested_report" / "first.html" + second_path = tmp_path / "second.html" + + returned = render_essay_score_report_html(report, first_path) + render_essay_score_report_html(report, second_path) + first = first_path.read_text(encoding="utf-8") + second = second_path.read_text(encoding="utf-8") + + assert returned == first_path + assert first == second + assert "" in first + assert '