Skip to content
Draft
Show file tree
Hide file tree
Changes from 15 commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
1e4bcc6
test(browser): observe pinned presentation cleanup
seonghobae Sep 9, 2026
7dd1318
test(browser): require hidden probe property reads
seonghobae Sep 9, 2026
7b407ca
fix(browser): read hidden probe text content
seonghobae Sep 9, 2026
ebce23e
fix(browser): read hidden presentation probe values
seonghobae Sep 9, 2026
c631b12
Merge remote-tracking branch 'origin/codex/presentation-pinned-chrome…
seonghobae Sep 9, 2026
ff723f1
test(browser): require causal presentation baseline
seonghobae Sep 9, 2026
a0a0848
fix(browser): require causal presentation transition
seonghobae Sep 9, 2026
431b32c
docs(browser): record presentation probe observation guard
seonghobae Sep 9, 2026
6019a9a
fix(browser): retain bounded session-start category
seonghobae Sep 9, 2026
bde3677
fix(browser): classify HTTP session startup failure
seonghobae Sep 9, 2026
1934a2e
fix(browser): classify closed startup diagnostics
seonghobae Sep 9, 2026
f8cb436
fix(browser): retain verbose startup categories
seonghobae Sep 9, 2026
da2e564
docs(browser): record sandbox owner dependency
seonghobae Sep 9, 2026
00724fe
fix(browser): restore startup diagnostic ownership
seonghobae Sep 9, 2026
d329f9e
docs(gap): track CodeQL owner successor
seonghobae Sep 9, 2026
70c710f
test(browser): require cleanup failure provenance in evidence
seonghobae Sep 9, 2026
3f8f1cc
fix(browser): preserve bounded cleanup failure provenance
seonghobae Sep 9, 2026
a297811
test(browser): bind profile cleanup provenance to primary failure
seonghobae Sep 9, 2026
0e77711
docs(evidence): trace cleanup failure provenance repair
seonghobae Sep 9, 2026
d31a8d5
test(browser): keep cleanup evidence class identity exact
seonghobae Sep 9, 2026
7ebbe58
docs(evidence): record cleanup contract harness correction
seonghobae Sep 9, 2026
2320dd7
test(browser): preserve nested cleanup provenance
seonghobae Sep 9, 2026
08133c3
fix(browser): preserve nested cleanup provenance
seonghobae Sep 9, 2026
f062cac
docs(browser): trace nested cleanup provenance
seonghobae Sep 9, 2026
8e68cf8
test(browser): preserve session-start cause through profile cleanup
seonghobae Sep 9, 2026
0050fff
fix(browser): retain bounded session cause across profile cleanup
seonghobae Sep 9, 2026
9488591
docs(browser): trace nested session-start cleanup provenance
seonghobae Sep 9, 2026
7e61280
test(browser): preserve nested session-start cause through cleanup
seonghobae Sep 9, 2026
960c1ac
fix(browser): retain bounded session-start cause through cleanup
seonghobae Sep 9, 2026
a88d2af
docs(traceability): record nested session-start cleanup repair
seonghobae Sep 9, 2026
4c9add7
docs(gap): restore canonical baseline ownership
seonghobae Sep 16, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,10 @@ The organization currently documents a **solo-maintainer** governance condition.
## Architecture constraints

- Keep Blink, V8, Skia, Viz, Dawn, Chromium sandboxing, Site Isolation, and Manifest V3 compatibility upstream-aligned.
- Pinned-Chromium presentation evidence may use only fixed CDP commands and declared static-fixture DOM outputs; page-provided scripts must never select commands or supply evaluation text.
- For hidden fixture outputs, read the bounded `textContent` property rather than rendered element text, and reject a probe baseline that already equals its target so an ACK cannot masquerade as a causal presentation transition.
- Browser-session failures may retain only a closed, credential-free WebDriver category; ChromeDriver startup/process diagnostics belong to their canonical owner lane and must not be captured here. Never emit remote or local diagnostic text into CI evidence.
- Hosted Ubuntu sandbox remediation is a canonical `.github` sandbox-helper workflow-contract dependency: do not add `--no-sandbox` or copy workflow setup into a product PR; adopt the reviewed immutable owner release and rerun the three browser trials.
- New product logic belongs in Rust control-plane modules behind narrow adapters.
- Rust crates must remain independently understandable and reusable.
- Keep logical origin, resolved destination, operating-system TCP peer, TLS service identity, proxy route, and HTTP semantics as separate authority boundaries.
Expand Down
6 changes: 5 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,10 @@ All notable changes to OriginWeave are documented in this file. The format follo

## [Unreleased]

- Keep Agent Task evidence to the standard bounded WebDriver session-creation category; ChromeDriver startup/process diagnostics remain with their canonical owner.
- Preserve the bounded standard WebDriver session-creation failure category in Agent Task evidence while continuing to redact remote driver diagnostics.
- Recognize the standard session-creation category from a ChromeDriver HTTP error response before applying the generic HTTP failure boundary.
- Added a pinned-Chromium presentation probe to the controlled Agent Task evidence lane. It records a baseline, applies fixed viewport/DPR/timezone overrides before the observed navigation, then resets and proves the baseline returns through declared fixture DOM outputs. Hosted browser evidence remains required.
- Keep WebDriver remote HTTP bodies, W3C error/message text, last-response startup detail, and mismatched remote `browserVersion` capability values out of CI exception strings while preserving fail-closed command/readiness/version decisions and the response-size bound.
- Publish success-shaped MV3/Agent Task compatibility JSON only after both owned loopback fixture servers complete their shutdown post-conditions; browser/trial gate failures still emit bounded diagnostic evidence before raising.
- Require loopback fixture-server cleanup to observe helper-thread termination after the bounded join, so a timed join cannot be treated as cleanup success while an owned server thread remains live.
Expand Down Expand Up @@ -121,4 +125,4 @@ All notable changes to OriginWeave are documented in this file. The format follo
- The hourly product agent has no Git metadata or repository authority. A separate post-verification publisher opens one PR and cannot approve or merge it.
- The unprivileged OpenCode user is restricted to loopback egress during model execution, preventing runner-wide allow-listed endpoints from becoming direct source-exfiltration channels.

[Unreleased]: https://github.com/ContextualWisdomLab/OriginWeave/compare/main...HEAD
[Unreleased]: https://github.com/ContextualWisdomLab/OriginWeave/compare/main...HEAD
4 changes: 4 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,4 +11,8 @@ Additional constraints:
- Do not merge logical origin, destination authorization, direct TCP peer proof, TLS service identity, proxy routing, or HTTP resource policy into one ambient authority.
- Do not add hostname reconnect, proxy-environment inheritance, dangerous certificate-verifier hooks, Common Name fallback, TLS 0-RTT, key logging, or secret extraction to a production TLS path.
- Keep changes bounded to one product gap and preserve modular crate boundaries.
- Browser presentation probes use fixed CDP commands and static-fixture DOM observations only; never turn page content into executable input.
- Read hidden fixture observations through the bounded `textContent` property, and require every pre-override value to differ from its target before accepting a presentation transition.
- Preserve only a closed WebDriver category in browser evidence; ChromeDriver startup/process diagnostics remain in their canonical owner lane, and diagnostic text is never serialized.
- Treat hosted Ubuntu sandbox remediation as a `.github` workflow-owner dependency; preserve Chromium sandboxing and require a released helper contract before rerunning browser evidence.
- Never claim a test, benchmark, browser integration, TLS identity, GPU execution, release, or merge succeeded without current exact-head evidence.
7 changes: 7 additions & 0 deletions docs/product-technical-gap-baseline.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,13 @@

This is a dated delivery baseline, not a substitute for the PRD, TRD, roadmap, architecture decisions, or live GitHub state. It keeps buyer-visible gaps, current issues, active pull-request evidence, and commercial completion tracks in one discoverable place. Protected `main` is the implementation boundary: code in an open pull request is not shipped behavior.

## Live continuity note: 2026-09-09

- The next #292 evidence slice is stacked on the existing pinned-Chrome Agent Task owner. It uses fixed Chromium CDP viewport/DPR/timezone commands, static fixture DOM observations, an observed pre-override baseline, and observed explicit-reset restoration. It is active-PR evidence only until the exact Chrome for Testing job succeeds.
- A failed Agent Task session start records only the bounded credential-free WebDriver failure category; this consumer neither captures nor classifies ChromeDriver startup/process diagnostics. Those diagnostics belong to their canonical owner lane, and no remote or local diagnostic text enters CI evidence.
- Historical #299 browser job `34316793780` on `f8cb436c` classified all three pre-navigation failures as `sandbox`. The remediation remains the canonical `.github#1792` sandbox-helper workflow contract, not a leaf workflow copy or `--no-sandbox`; owner PR `.github#1857` and its CodeQL wake-race successor `.github#2056` remain unreleased, so consumer evidence is blocked pending immutable owner adoption and a fresh three-trial replay.
- This runner proves a narrow browser-evidence contract, not a product Browser Session implementation. It does not transfer #293's standard-BiDi capability boundary into protected main or claim full-profile admission.

## Observed snapshot: 2026-08-26

### Protected-main truth
Expand Down
171 changes: 161 additions & 10 deletions scripts/ci/run_mv3_compatibility.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,10 @@
REPEATABILITY_TRIALS = 3
AGENT_TASK_REPEATABILITY_TRIALS = 3
AGENT_TASK_INPUT_VALUE = "originweave controlled input"
PRESENTATION_VIEWPORT_WIDTH = 1200
PRESENTATION_VIEWPORT_HEIGHT = 800
PRESENTATION_DEVICE_PIXEL_RATIO = 2
PRESENTATION_TIMEZONE = "Pacific/Kiritimati"
REQUEST_TIMEOUT_SECONDS = 5.0
STARTUP_TIMEOUT_SECONDS = 20.0
FIXTURE_TIMEOUT_SECONDS = 20.0
Expand Down Expand Up @@ -73,13 +77,20 @@ def __init__(self, cleanup_error: BaseException) -> None:


class AgentTaskSessionStartError(RuntimeError):
"""Classify a failed Agent Task browser session without exposing driver text."""
"""Record a failed Agent Task browser session without exposing driver text."""

def __init__(self, session_error: BaseException) -> None:
self.session_error_type = type(session_error).__name__
super().__init__("Agent Task browser session failed to start")


class WebDriverSessionNotCreatedError(RuntimeError):
"""Report the standard session-creation failure without remote diagnostics."""

def __init__(self) -> None:
super().__init__("WebDriver could not create a browser session")


def _free_loopback_port() -> int:
"""Reserve and release one loopback TCP port for a short-lived local service."""

Expand Down Expand Up @@ -143,6 +154,20 @@ def _json_request(
if len(raw) > MAX_WEBDRIVER_RESPONSE_BYTES:
raise RuntimeError("WebDriver response exceeded the bounded JSON limit")
if response.status >= 400:
try:
error_response = json.loads(raw.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError):
raise RuntimeError(
f"WebDriver HTTP request failed with status {response.status}"
) from None
error_value = (
error_response.get("value") if isinstance(error_response, dict) else None
)
if (
isinstance(error_value, dict)
and error_value.get("error") == "session not created"
):
raise WebDriverSessionNotCreatedError()
raise RuntimeError(
f"WebDriver HTTP request failed with status {response.status}"
)
Expand All @@ -153,6 +178,8 @@ def _json_request(
if not isinstance(decoded, dict):
raise RuntimeError("WebDriver returned a non-object JSON payload")
value = decoded.get("value")
if isinstance(value, dict) and value.get("error") == "session not created":
raise WebDriverSessionNotCreatedError()
if isinstance(value, dict) and value.get("error"):
raise RuntimeError("WebDriver command failed")
return decoded
Expand Down Expand Up @@ -230,6 +257,96 @@ def _get_element_semantics(
return role, label


def _presentation_cdp_path(session_id: str) -> str:
"""Return ChromeDriver's fixed vendor endpoint for this exact session."""

return _webdriver_path(session_id, "/goog/cdp/execute")


def _apply_presentation_probe(driver_port: int, session_id: str) -> None:
"""Apply only the pinned viewport, DPR, and timezone probe before navigation."""

_json_request(
driver_port,
"POST",
_presentation_cdp_path(session_id),
{
"cmd": "Emulation.setDeviceMetricsOverride",
"params": {
"width": PRESENTATION_VIEWPORT_WIDTH,
"height": PRESENTATION_VIEWPORT_HEIGHT,
"deviceScaleFactor": PRESENTATION_DEVICE_PIXEL_RATIO,
"mobile": False,
},
},
)
_json_request(
driver_port,
"POST",
_presentation_cdp_path(session_id),
{
"cmd": "Emulation.setTimezoneOverride",
"params": {"timezoneId": PRESENTATION_TIMEZONE},
},
)


def _reset_presentation_probe(driver_port: int, session_id: str) -> None:
"""Remove the exact probe overrides before reusing the browser session."""

_json_request(
driver_port,
"POST",
_presentation_cdp_path(session_id),
{"cmd": "Emulation.clearDeviceMetricsOverride", "params": {}},
)
_json_request(
driver_port,
"POST",
_presentation_cdp_path(session_id),
{"cmd": "Emulation.setTimezoneOverride", "params": {"timezoneId": ""}},
)


def _presentation_probe_target() -> dict[str, str]:
"""Return the exact fixed page-observed target for this evidence probe."""

return {
"viewport": f"{PRESENTATION_VIEWPORT_WIDTH}x{PRESENTATION_VIEWPORT_HEIGHT}",
"device_pixel_ratio": str(PRESENTATION_DEVICE_PIXEL_RATIO),
"timezone": PRESENTATION_TIMEZONE,
}


def _validate_presentation_probe_baseline(baseline: dict[str, str]) -> None:
"""Require an observable transition for every presentation surface under test."""

target = _presentation_probe_target()
if any(baseline.get(key) == value for key, value in target.items()):
raise RuntimeError("presentation probe baseline already matched target")


def _read_presentation_probe(driver_port: int, session_id: str) -> dict[str, str]:
"""Read only declared fixture observations through bounded element endpoints."""

observed: dict[str, str] = {}
for key, selector in {
"viewport": "#presentation-viewport",
"device_pixel_ratio": "#presentation-device-pixel-ratio",
"timezone": "#presentation-timezone",
}.items():
element_id = _find_element(driver_port, session_id, selector)
value = _json_request(
driver_port,
"GET",
_element_command_path(session_id, element_id, "/property/textContent"),
).get("value")
if not isinstance(value, str):
raise RuntimeError("presentation probe observation was malformed")
observed[key] = value
return observed


def _cleanup_browser_session(driver_port: int, session_id: str) -> None:
"""Delete one WebDriver session through the fixed loopback authority."""

Expand Down Expand Up @@ -567,7 +684,11 @@ def _run_agent_task_browser_pass(
driver_port = _free_loopback_port()
session_id: str | None = None
driver = subprocess.Popen(
[str(chromedriver_bin), f"--port={driver_port}", "--allowed-ips=127.0.0.1"],
[
str(chromedriver_bin),
f"--port={driver_port}",
"--allowed-ips=127.0.0.1",
],
stdout=subprocess.DEVNULL,
stderr=subprocess.STDOUT,
text=True,
Expand Down Expand Up @@ -634,6 +755,19 @@ def _run_agent_task_browser_pass(
).get("value")
if initial_url != fixture_url:
raise RuntimeError("Agent Task initial URL mismatch")
baseline_presentation = _read_presentation_probe(driver_port, session_id)
_validate_presentation_probe_baseline(baseline_presentation)
target_presentation = _presentation_probe_target()
_apply_presentation_probe(driver_port, session_id)
_json_request(
driver_port,
"POST",
_webdriver_path(session_id, "/url"),
{"url": fixture_url},
)
applied_presentation = _read_presentation_probe(driver_port, session_id)
if applied_presentation != target_presentation:
raise RuntimeError("presentation probe post-condition failed")
input_element = _find_element(driver_port, session_id, "#task-text")
input_role, input_name = _get_element_semantics(
driver_port,
Expand Down Expand Up @@ -745,6 +879,16 @@ def _run_agent_task_browser_pass(
url_unchanged = url_unchanged and accepted_outcome_url == initial_url
if not url_unchanged:
raise RuntimeError("Agent Task URL changed before accepted outcome")
_reset_presentation_probe(driver_port, session_id)
_json_request(
driver_port,
"POST",
_webdriver_path(session_id, "/url"),
{"url": fixture_url},
)
cleanup_presentation = _read_presentation_probe(driver_port, session_id)
if cleanup_presentation != baseline_presentation:
raise RuntimeError("presentation probe cleanup post-condition failed")
return {
"browser_version": browser_version,
"pre_action_baseline_verified": True,
Expand All @@ -757,6 +901,8 @@ def _run_agent_task_browser_pass(
"input_semantics_verified": True,
"submit_semantics_verified": True,
"extensions_disabled_requested": True,
"presentation_applied": True,
"presentation_cleanup_verified": True,
"duration_ms": round((time.monotonic() - started) * 1000),
}
finally:
Expand Down Expand Up @@ -829,6 +975,8 @@ def _run_agent_task_trial(
"input_semantics_verified": result["input_semantics_verified"],
"submit_semantics_verified": result["submit_semantics_verified"],
"extensions_disabled_requested": result["extensions_disabled_requested"],
"presentation_applied": result["presentation_applied"],
"presentation_cleanup_verified": result["presentation_cleanup_verified"],
"profile_cleaned": profile_cleaned,
"duration_ms": round((time.monotonic() - trial_started) * 1000),
}
Expand All @@ -850,6 +998,8 @@ def _agent_task_surfaces_complete(agent_task_trials: list[dict[str, Any]]) -> bo
and trial.get("url_unchanged") is True
and trial.get("input_semantics_verified") is True
and trial.get("submit_semantics_verified") is True
and trial.get("presentation_applied") is True
and trial.get("presentation_cleanup_verified") is True
and trial.get("profile_cleaned") is True
for trial in agent_task_trials
)
Expand Down Expand Up @@ -978,13 +1128,14 @@ def main() -> int:
http.client.HTTPException,
json.JSONDecodeError,
) as error:
agent_task_trials.append(
{
"trial_number": trial_number,
"passed": False,
"failure_type": type(error).__name__,
}
)
failed_trial: dict[str, Any] = {
"trial_number": trial_number,
"passed": False,
"failure_type": type(error).__name__,
}
if isinstance(error, AgentTaskSessionStartError):
failed_trial["failure_cause_type"] = error.session_error_type
agent_task_trials.append(failed_trial)

agent_task_successful_trials = sum(
1 for trial in agent_task_trials if trial.get("passed") is True
Expand Down Expand Up @@ -1048,4 +1199,4 @@ def main() -> int:


if __name__ == "__main__":
raise SystemExit(main())
raise SystemExit(main())
10 changes: 10 additions & 0 deletions tests/fixtures/agent_task_basic/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,9 @@ <h1>Controlled Agent Task</h1>
</form>

<output id="task-result" data-state="idle" aria-live="polite">idle</output>
<output id="presentation-viewport" hidden></output>
<output id="presentation-device-pixel-ratio" hidden></output>
<output id="presentation-timezone" hidden></output>

<p
data-originweave-untrusted="prompt-injection"
Expand All @@ -32,6 +35,13 @@ <h1>Controlled Agent Task</h1>
const taskText = document.getElementById("task-text");
const result = document.getElementById("task-result");

document.getElementById("presentation-viewport").textContent =
`${window.innerWidth}x${window.innerHeight}`;
document.getElementById("presentation-device-pixel-ratio").textContent =
String(window.devicePixelRatio);
document.getElementById("presentation-timezone").textContent =
Intl.DateTimeFormat().resolvedOptions().timeZone;

form.addEventListener("submit", (event) => {
event.preventDefault();
result.dataset.state = "submitted";
Expand Down
6 changes: 5 additions & 1 deletion tests/test_agent_task_action_transition_evidence_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -160,8 +160,12 @@ def test_surface_completeness_requires_transition_baseline_evidence(self) -> Non
trial["input_value_verified"] = True
self.assertFalse(complete([trial]))
trial["pre_click_baseline_verified"] = True
self.assertFalse(complete([trial]))
trial["presentation_applied"] = True
self.assertFalse(complete([trial]))
trial["presentation_cleanup_verified"] = True
self.assertTrue(complete([trial]))


if __name__ == "__main__":
unittest.main()
unittest.main()
10 changes: 10 additions & 0 deletions tests/test_agent_task_chromium_sandbox_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,5 +22,15 @@ def test_agent_task_browser_pass_does_not_disable_chromium_sandbox(self) -> None

self.assertNotIn('"--no-sandbox"', browser_pass_source)

def test_agent_task_browser_pass_does_not_own_chromedriver_diagnostics(self) -> None:
"""Keep ChromeDriver process diagnostics in their canonical owner lane."""

namespace = runpy.run_path(str(RUNNER), run_name="agent_task_sandbox_contract")
browser_pass_source = inspect.getsource(namespace["_run_agent_task_browser_pass"])

self.assertNotIn('"--verbose"', browser_pass_source)
self.assertNotIn('"--log-path=', browser_pass_source)
self.assertNotIn("_classify_chromedriver_startup_diagnostic", browser_pass_source)

if __name__ == "__main__":
unittest.main()
Loading
Loading