docs: align metadata tooling instructions - #18
Merged
Merged
Conversation
8x MI355X Qwen3-30B-A3B MoE pretraining MFU optimization. Built on ghcr.io/amdpilot-org/primus-mi355x-ready:v1. Two task.yaml variants for A/B testing the Phase 1 baseline agent: - task.yaml: phase1_baseline: true - task_nophase1.yaml: phase1_baseline: false Both pin the executor to Kimi-K2.6 at 10.235.24.154:30000.
The previous base (ghcr.io/amdpilot-org/primus-mi355x-ready:v1) was a slimmed copy that did NOT include /workspace/primus_train/Primus. Phase 1 spent its full max_turns budget trying to bootstrap from scratch. Switch to primus-mi355x-flat:v1 (locally available, has Primus + Primus-Turbo pre-installed and patched for triton 3.4.0). Also install uv at /root/.local/bin in our Dockerfile so the kimi-cli runtime's source $HOME/.local/bin/env succeeds — without this, executor trials exit 137 immediately after image switch. Tag the new image primus-qwen3-30b-mfu-base:v1 so amdpilot triggers a build the first time it runs.
Two cumulative fixes for n08-09 (8x MI355X): 1. Without PYTORCH_ROCM_ARCH=gfx950 (and AITER_ROCM_ARCH, HSA_NO_SCRATCH_RECLAIM, HIP_FORCE_DEV_KERNARG, etc) the torch HIP runtime can't dispatch kernels and benchmarks die immediately with hipErrorInvalidDeviceFunction. These were set in xiao/baizhou's working containers but missing from amdpilot's docker run line. Add to both Dockerfile ENV and task.yaml container.env so they apply via either path. 2. Replace /workspace/detect_interface.sh with a /proc-based detector (the original needed `ip` from iproute2, unavailable in the slim base). bench_mfu.sh now auto-detects GLOO/NCCL socket IFNAME at runtime if bench_config.env doesn't pin one — without this, Megatron's distributed init fails fast (3-5s) before training starts.
Enables tag + push of the Phase 1 baseline image to
docker.io/jhinpan/primus-qwen3-30b-mfu-phase1 after a successful
phase1 commit. Template: {date}-{metric} (e.g. 20260422-278p80) +
:latest. Other nodes can then `docker pull` that tag and skip Phase 1
entirely.
Switch repository from docker.io/jhinpan to ghcr.io/amdpilot-org so all nodes in the org can pull the verified phase1-baseline image directly. Bump push timeout_s to 9000 (2.5h) to absorb the one-time 42 GB base-layer seed; subsequent pushes only upload the phase1-commit delta (~500 MB - 1 GB) via GHCR cross-repo mount.
Remove references to missing metadata schema files and point reviewers at the current registry and validation helper tests.
Remove the inherited Primus eval bundle from the docs-only branch so the PR only updates metadata tooling guidance.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Test plan
PYTHONPATH=. python3 -m pytest tests -q