Skip to content

Add Qwen2 config and sampling params - #6

Merged
apaz-cli merged 8 commits into
mainfrom
ap/sampling-params
Feb 28, 2025
Merged

Add Qwen2 config and sampling params#6
apaz-cli merged 8 commits into
mainfrom
ap/sampling-params

Conversation

@apaz-cli

Copy link
Copy Markdown
Contributor

No description provided.

@apaz-cli apaz-cli changed the title Add inference sampling params to config Add Qwen2 config and sampling params Feb 28, 2025
Comment thread src/zeroband/models.py Outdated
Comment thread src/zeroband/inference.py
Comment thread src/zeroband/inference.py
Comment on lines +149 to +156
if config.cpu_offload_gb != 0.0 and config.cpu_offload_percentage != 0.0:
raise ValueError("Cannot set both cpu_offload_gb and cpu_offload_percentage")
if config.cpu_offload_percentage != 0.0:
cpu_offload_gb = psutil.virtual_memory().available * config.cpu_offload_percentage
elif config.cpu_offload_gb != 0.0:
cpu_offload_gb = config.cpu_offload_gb
else:
cpu_offload_gb = 0.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can skip this for now as we don't need offloading

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we probably will in the future, and it was basically zero effort. The config I wrote isn't offloading.

Comment thread src/zeroband/models.py
Comment thread configs/inference/debug.toml Outdated
Comment thread configs/inference/Qwen32B/qwen32B.toml
@apaz-cli
apaz-cli merged commit 7dd4192 into main Feb 28, 2025
samsja pushed a commit that referenced this pull request Mar 30, 2026
…uv-source

update deepdive source to threading version (no longer multi-proc)
seanbell added a commit to clouddatalabs/scalerl-prime-rl that referenced this pull request Apr 26, 2026
…polish

- scripts/prebuild_tb_images.sh: drop `set -e`; report per-task
  succeeded/failed at end and exit non-zero if any failed (an upstream
  task with a broken Dockerfile, e.g. log-summary's missing logs/ dir,
  no longer aborts the remaining 27 builds).
- docs/SCALERL.md "4xB200" → "single 8-GPU B200 node (4 train + 4
  infer GPUs split per the [deployment] block)" — disambiguates the
  steady-state perf number for Snowflake reading capacity sizing
  (Snowflake POC critique PrimeIntellect-ai#5).
- configs/scalerl_math/rl.toml: add doc note that the implicit
  `use_token_client = true` default is correct ONLY for single-turn
  envs. Cross-references the TB config's full rationale so a future
  user copying scalerl_math to bootstrap a multi-turn env doesn't
  silently get the linear-history corruption mode (Snowflake POC
  critique PrimeIntellect-ai#6).
snimu added a commit that referenced this pull request Jun 4, 2026
- enabled_losses=None now validated as the full term list, so >1 echo term per
  env is caught at config time instead of at rollout time. [review #6]
- loss_overrides keys validated against `losses`; non-echo overrides rejected. [#7]
- warn (don't fail) when prompt-role echo is configured with renderer=None
  (MITO), where prompt_attribution is unavailable so it would silently no-op. [#8]
- token_export: add echo_mask/echo_weight columns + export sequences trained
  only via echo (gate on loss_mask OR echo_mask). [#9]
- doc notes: echo CE uses the rollout temperature (scale alpha to compensate,
  kept as-is); negative alpha is intentional (suppresses tokens). [#1, #10]
- tests for the new config validators.

Deferred to a follow-up pass (per the review): full per-sample primary routing /
rl-disable [#2b] + the <=1-primary validation it enables [#5], and the multi-run
losses fingerprint [#3]. Not run locally; ruff + py_compile clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
snimu added a commit that referenced this pull request Jun 5, 2026
…ht resolver

- #1 reserve loss-term names {sft, opd} (any term) and `rl` (non-primary), so an overlay can't
  silently overwrite a training_mode dispatch core or the rl primary in the trainer registry.
- #5 resolve the primary's advantage weight orchestrator-side: scale the per-token advantage by the
  advantage-weight's tau in process_group and drop adv_tau from the dppo_kl core / RLLossConfig. Now
  *any* primary core (dppo_kl or custom) gets the resolved advantage × tau — no per-core special-case.
  Bit-identical for the default tau=1.0.
- #4 overlay trainability = non-None AND non-zero, so a zero-weight overlay (e.g. advantage-weighted
  with zero advantage) no longer keeps an otherwise-empty batch alive past the empty-batch guard.
- #6 custom overlay weight resolver is group-aware: it now receives `WeightInputs{sample, rollouts}`
  (the full GRPO group) instead of a lone sample, so it can compute group-relative weights.
- #10 document the overlay_mask/overlay_weight token-export columns (schema v2) in the configs skill.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@hallerite hallerite mentioned this pull request Jul 1, 2026
6 tasks
@mikasenghaas
mikasenghaas deleted the ap/sampling-params branch August 5, 2026 04:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants