-
Notifications
You must be signed in to change notification settings - Fork 349
feat(resources_servers): add interactive browser RL environment (local Playwright + remote CDP backends) #1865
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
waple0820
wants to merge
12
commits into
NVIDIA-NeMo:main
Choose a base branch
from
waple0820:feat/lexmount-browser
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
12 commits
Select commit
Hold shift + click to select a range
9b2b263
feat(resources_servers): add lexmount_browser interactive browser env…
waple0820 55350f3
feat(lexmount_browser): add runnable example.sh, GRPO smoke config, a…
waple0820 0c716e2
fix(lexmount_browser): make example.sh work in a clean-room environment
waple0820 6f4d2f5
docs(lexmount_browser): document clean-room prerequisites for reviewers
waple0820 ff0533e
refactor(lexmount_browser): scope the PR to the environment contribution
waple0820 8200ff4
fix(lexmount_browser): unblock the event loop, bound observations, fa…
waple0820 821bc4c
refactor(interactive_browser): generic browser environment, vendor as…
waple0820 c13d6cf
refactor(interactive_browser): move the example provider's config nex…
waple0820 73fb0fc
fix(interactive_browser): report an unmeasurable rollout instead of s…
waple0820 d10dc9c
fix(interactive_browser): address review — open() leak, shared contex…
waple0820 cd402e0
feat(interactive_browser): ship the task_data schema every resources …
waple0820 6d3d40d
fix(interactive_browser): address review — page text, abandoned creat…
waple0820 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,204 @@ | ||
| # Interactive Browser — NeMo-Gym browser RL environment | ||
|
|
||
| A NeMo-Gym **stateful resources server** that turns a real browser into an RL | ||
| environment. Each rollout (`session_id`) owns one isolated live browser context; | ||
| the policy drives it with tool calls (navigate / click / type / observe / finish); | ||
| `verify()` returns the task reward from the live browser state. | ||
|
|
||
| This is the interactive complement to the existing read-only browsing envs | ||
| (`google_search`, `browsecomp_advanced_harness`): those search-and-extract text; | ||
| this one *operates* pages (stateful, multi-step web agency). | ||
|
|
||
| ## Backends | ||
|
|
||
| The server depends only on the small `BrowserBackend` contract | ||
| (`browser/base.py`): `open / goto / click / type / observe / current_url / text / | ||
| close`. Where the browser actually runs is one config block; nothing else in the | ||
| environment changes. | ||
|
|
||
| | `backend:` | Browser runs | Use it for | | ||
| | --- | --- | --- | | ||
| | `local_playwright` (**default**) | A Chromium launched on this machine, one per rollout | Local dev, CI, offline `site/` tasks. No account, no quota, no network. | | ||
| | `remote_cdp` | Anywhere reachable over CDP, supplied per rollout by a **session provider** | Keeping browsers off the training node: a browser container, another host, or a hosted browser service. | | ||
|
|
||
| ```yaml | ||
| backend: | ||
| local_playwright: | ||
| headless: true | ||
| ``` | ||
|
|
||
| ```yaml | ||
| backend: | ||
| remote_cdp: | ||
| connect_timeout_s: 60 | ||
| session_provider: | ||
| static_cdp: | ||
| cdp_url: "http://127.0.0.1:9222" # e.g. chromium --remote-debugging-port=9222 | ||
| ``` | ||
|
|
||
| ### Session providers | ||
|
|
||
| A provider answers one question for `remote_cdp`: *where is this rollout's | ||
| browser, and how do I give it back?* | ||
|
|
||
| ```python | ||
| class BrowserSessionProvider(Protocol): | ||
| name: str | ||
| async def acquire(self, spec: BrowserSessionSpec) -> BrowserSessionHandle: ... # returns a live cdp_url | ||
| async def release(self, handle: BrowserSessionHandle) -> None: ... # idempotent | ||
| ``` | ||
|
|
||
| | Provider | What it does | | ||
| | --- | --- | | ||
| | `static_cdp` (built in) | Uses an endpoint that already exists (`cdp_url`, or `$BROWSER_CDP_URL`). Release is a no-op. | | ||
| | `lexmount` (example) | Creates one isolated cloud browser session per rollout and deletes it afterwards. Code, config and docs live together in [`providers/lexmount/`](providers/lexmount/README.md). | | ||
|
|
||
| Providers are selected by name (`{provider_name: {kwargs}}`) and imported only | ||
| when selected, so an unused provider never pulls in its SDK. Publish one from | ||
| your own package with an entry point — no fork required: | ||
|
|
||
| ```toml | ||
| [project.entry-points."nemo_gym.browser_session_providers"] | ||
| my_browser_cloud = "my_pkg.provider:MyProvider" | ||
| ``` | ||
|
|
||
| `register_session_provider(name, cls)` does the same thing in-process, and | ||
| `register_backend(name, builder)` adds a backend that is not CDP-shaped at all | ||
| (a vendor SDK with its own DOM API implements `BrowserBackend` directly). | ||
|
|
||
| ## Tools & observation | ||
|
|
||
| Tools: `browser_navigate(url)`, `browser_click(element_id)`, | ||
| `browser_type(element_id, text)`, `browser_observe()`, `browser_finish(answer)`. | ||
| Observation is a **compact numbered list of interactive elements** (`[id] role: name`) | ||
| plus URL/title — deliberately token-cheap (raw HTML/pixels are far too expensive for | ||
| small policies and for training context length). `element_id`s come from the most | ||
| recent observation. | ||
|
|
||
| Element collection stops at `max_elements` (default 50) — probing each interactive | ||
| node costs several browser round-trips, so a large page is truncated rather than | ||
| fully scanned, and the observation says so. | ||
|
|
||
| ## Reward | ||
|
|
||
| `verify()` scores a per-task success spec in `verifier_metadata`: | ||
| `final_url` / `url_contains` / `dom_contains` / `answer_equals`. Sparse 0/1 outcome | ||
| reward by default (least reward-hackable); extend `_score()` with new keys as needed. | ||
|
|
||
| A spec carrying none of those keys **raises** rather than scoring 0 — a misspelled | ||
| key would otherwise give every rollout reward 0, which looks identical to a policy | ||
| that never solves the task. | ||
|
|
||
| ## Session lifetime | ||
|
|
||
| A browser is released when the rollout is scored (`verify`) or when the same | ||
| `session_id` is re-seeded. There is no independent episode TTL and no | ||
| client-side cap on concurrent sessions, so: | ||
|
|
||
| - a rollout abandoned without either (trainer crash, client disconnect) leaves | ||
| its browser open until the process exits (local) or the provider reclaims it; | ||
| - with a metered provider, size the account quota **above** the rollout | ||
| concurrency, with headroom for sessions still being torn down. | ||
|
|
||
| Every acquired provider session is released exactly once — including when the | ||
| CDP connect fails after the session was created, and however often `close()` is | ||
| called (`tests/test_registry_and_sessions.py`). | ||
|
|
||
| ## Run | ||
|
|
||
| ### 1. Backend contract tests (no GPU, no Gym serving stack) | ||
| ```bash | ||
| uv run --no-project --with playwright python -m playwright install chromium | ||
| uv run --no-project --with playwright --with pytest --with pytest-asyncio python -m pytest tests -q | ||
| ``` | ||
| Runs one contract against **both** shipped backends: `local_playwright`, and | ||
| `remote_cdp` against a Chromium the test suite starts with | ||
| `--remote-debugging-port` — so the remote path is covered in CI without any | ||
| third-party service. Tasks are the bundled offline `site/` (deterministic, ToS-safe). | ||
|
|
||
| > On a bare Linux container/VM, Chromium also needs system libraries (libnss3, | ||
| > libgbm1, ...): if the install warns `Host system is missing dependencies to | ||
| > run browsers`, run (root/sudo required) | ||
| > `uv run --no-project --with playwright python -m playwright install-deps chromium`. | ||
|
|
||
| ### 2. As a NeMo-Gym environment (rollouts / training) | ||
|
|
||
| Requirements for the policy endpoint — it must satisfy **both**: | ||
|
|
||
| 1. **Speak the Responses API** (`POST /v1/responses`). Gym's own model servers | ||
| already do — `vllm_model` exposes `/v1/responses` and forwards to vLLM's | ||
| chat-completions endpoint for you, so nothing extra is needed there. The | ||
| requirement applies when you point the agent straight at a gateway instead. | ||
| 2. **Parse tool calls into structured `function_call` items.** If the server | ||
| returns tool-call markup as plain text, the agent sees zero tool calls, the | ||
| browser is never driven, and every rollout "succeeds" with **reward 0.0** — | ||
| silently. For vLLM + Qwen-family models, launch with | ||
| `--enable-auto-tool-choice --tool-call-parser hermes`. | ||
|
|
||
| ```bash | ||
| # from the repo root: | ||
| gym env start --resources-server interactive_browser \ | ||
| --model-type openai_model --model <served-model-name> \ | ||
| --model-url https://your-endpoint/v1 --model-api-key <key> & | ||
| gym eval run --no-serve --agent interactive_browser_simple_agent \ | ||
| --input resources_servers/interactive_browser/data/example.jsonl \ | ||
| --output resources_servers/interactive_browser/data/example_rollouts.jsonl --limit 2 | ||
| ``` | ||
| Then plug into NeMo-RL GRPO via `examples/nemo_gym/run_grpo_nemo_gym.py`. | ||
|
|
||
| A provider's config ships with the provider, so a remote run loads it by path: | ||
|
|
||
| ```bash | ||
| gym env start --config resources_servers/interactive_browser/providers/lexmount/configs/lexmount.yaml \ | ||
| --model-type openai_model --model <served-model-name> \ | ||
| --model-url https://your-endpoint/v1 --model-api-key <key> & | ||
| ``` | ||
|
|
||
| Remote browsers cannot open this repo's offline `site/` pages (local `file://` | ||
| URIs), so remote runs use live-web tasks — `data/example_remote.jsonl`, which | ||
| that config points at. | ||
|
|
||
| ## Scope and maturity | ||
|
|
||
| This is a working environment meant to be extended, not a finished benchmark. | ||
| What that means concretely: | ||
|
|
||
| **Dataset.** `data/example.jsonl` is a small set of tasks over the bundled | ||
| offline `site/`, and `data/example_remote.jsonl` a handful of live-web ones. | ||
| They exist to exercise the loop end to end and keep CI deterministic. Point | ||
| `generate_data.py` at your own task source for anything beyond that. | ||
|
|
||
| **Verifier.** Outcome-only, checked once in `verify()`: final URL, URL substring, | ||
| DOM substring, or an exact answer match. That is enough for tasks with a single | ||
| unambiguous end state, and not enough for tasks where the path matters — a | ||
| policy that reaches the right URL by a shortcut scores the same as one that | ||
| completed the flow. Tasks whose correctness depends on what happened along the | ||
| way need a verifier of their own. | ||
|
|
||
| **Harness.** The observation is a numbered list of interactive elements plus | ||
| capped body text. No screenshots, no accessibility tree, no per-step reward. | ||
| A vision policy or a task needing layout will want more than this exposes; the | ||
| `BrowserBackend` contract is where to add it. | ||
|
|
||
| **Backends.** Local Chromium is the CI path. The hosted provider under | ||
| `providers/lexmount/` is one example of a remote backend and is optional — the | ||
| environment runs without its SDK installed. | ||
|
|
||
| ## Files (Gym `new-environment` spec) | ||
| - [x] `app.py` — resources server (seed_session + tools + verify) | ||
| - [x] `browser/` — `BrowserBackend` contract, `LocalPlaywrightBackend`, `RemoteCDPBackend` + `StaticCDPProvider`, registry | ||
| - [x] `providers/lexmount/` — example session provider: code, its own `configs/`, README (hosted browser service, optional SDK) | ||
| - [x] `configs/interactive_browser.yaml` | ||
| - [x] `site/` — bundled offline test site (deterministic tasks/CI) | ||
| - [x] `generate_data.py` + `data/example.jsonl` (offline) + `data/example_remote.jsonl` (live web) | ||
| - [x] `tests/` — backend contract (both backends) + selection/session bookkeeping | ||
| - [x] `requirements.txt`, `README.md` | ||
| - [x] `data/example_rollouts.jsonl` — 5 rollouts collected against a Responses-API endpoint (reward 1.0 on the offline site; reward wiring exercised end-to-end) | ||
| - [ ] GRPO training-signal run (train via NeMo-RL's `examples/nemo_gym/run_grpo_nemo_gym.py`) | ||
|
|
||
| ## Licensing | ||
| - Environment code: Apache 2.0 (matches NeMo-Gym). | ||
| - Reference backend: Playwright (Apache 2.0). | ||
| - Third-party browser-service SDKs (e.g. the `lexmount` provider): separate, | ||
| optional dependencies installed by the operator; not bundled, and not required | ||
| by the default backend. | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,2 @@ | ||
| # SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. | ||
| # SPDX-License-Identifier: Apache-2.0 |
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
could you add a short section on dataset availability and scope , such as if it is meant just as example with a few data points or to be extended on real datasets, if verifier and agent harness are robust or just example. for example the verifiers dont seem very robust and the harness doesnt seem to expose paragraph text or screenshot, and other useful feature, but is fine to be example
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Added a scope section in 86a9aa9 — it states plainly what the dataset, verifier and harness are and are not.
On why it's thin: this environment came out of our own browser RL runs, so the bundled tasks are what we needed to exercise the loop end to end and keep CI deterministic, not a curated benchmark. The verifier being outcome-only is the sharp edge — a policy that reaches the right URL by a shortcut scores the same as one that completed the flow. Real task sets want a verifier that checks the path, which is what I'd extend next.