prepare: serialise concurrent model preparation with flock - #134
Conversation
# Conflicts: # docs/docker.md # docs/gotchas.md
|
Merging — and you are right about the premise, which I got wrong. I said a stale lock would refuse to prepare. The rest, exercised with the lock lines as they appear in
The bounded wait is what makes this a different change from the one I pushed back on. Refusing on sight turns the normal case — the entrypoint's prepare racing Gotcha 59 is correct now: #124 merged a minute ago with 58. |
Why: the entrypoint runs prepare before every start, so a booting container races `docker compose run --rm prepare`, and two containers starting together after a crash race each other. Every step is idempotent but the script is not concurrency-safe: two runs against one model dir can interleave a shard rewrite with an index write, leaving shards missing from model.safetensors.index.json, a config that disagrees with the tensors, or a half-fetched fast variant -- all of which surface far from the cause. What: docker/prepare.sh takes an exclusive flock on <models dir>/.prepare.lock, where the directory comes from dirname "$BASE", so BASE_MODEL_DIR cannot move the lock inside (or out of) the model dir. The script waits up to PREPARE_LOCK_WAIT seconds (default 600) for a holder and only then refuses, so a legitimate concurrent prepare is waited for rather than rejected on sight. The lock is advisory and released when the holder exits, so a leftover .prepare.lock file is inert -- it is a lock, not a marker, and nothing has to clean it up. Docs: gotcha 59 plus a bullet in docs/docker.md where the entrypoint's prepare step is described. Split out of syv-ai#124 at review request, so that PR is only the chat-template effort translation.
264875e to
828c957
Compare
|
Merging — and you are right about the premise, which I got wrong. I said a stale lock would refuse to prepare. The rest, exercised with the lock lines as they appear in
The bounded wait is what makes this a different change from the one I pushed back on. Refusing on sight turns the normal case — the entrypoint's prepare racing Gotcha 59 is correct now: #124 merged a minute ago with 58. Rebased your branch onto main myself rather than bouncing it back — #124 merged a few minutes before this and touched the same two files, so the conflict was my sequencing, not yours. Checked after rebasing: the lock is taken at line 24, the step loop starts at line 59, so it still precedes every mutating step — including #124's |
…i) into local main Upstream's six commits since the last sync: syv-ai#124, syv-ai#134, syv-ai#137, syv-ai#138 are the reviewed squashes of branches local main already carries (resolve_config.sh is byte-identical), so they land as the reviewed variants of the same features -- the launchers' fail-closed source guard, resolve_api_key.sh, and warmup.sh using that resolver instead of its own chain. syv-ai#140 and syv-ai#141 are new content: the README setup decision tree, .env.example, single-user/README.md and .github/FUNDING.yml. Conflicts, one hunk each: - batch/start_qwen.sh, single-user/start_qwen.sh: took upstream's guarded source (the launchers do not run under `set -e`); kept local's select_model.sh call in the single-user launcher. - docs/docker.md: kept local's paragraph (upstream never had it). - docs/gotchas.md: took upstream's indentation on gotcha 58's continuation. - docker/prepare.sh auto-merged into a duplicated translate block (local hoists DIRS and passes them to harden_chat_template.py; upstream nests them); kept local's version, a superset of upstream's. Verified: bash -n on every touched shell file, test_resolution.sh passes, patches/ untouched by the merge.
# Conflicts: # docs/docker.md # docs/gotchas.md
…i) into local main Upstream's six commits since the last sync: syv-ai#124, syv-ai#134, syv-ai#137, syv-ai#138 are the reviewed squashes of branches local main already carries (resolve_config.sh is byte-identical), so they land as the reviewed variants of the same features -- the launchers' fail-closed source guard, resolve_api_key.sh, and warmup.sh using that resolver instead of its own chain. syv-ai#140 and syv-ai#141 are new content: the README setup decision tree, .env.example, single-user/README.md and .github/FUNDING.yml. Conflicts, one hunk each: - batch/start_qwen.sh, single-user/start_qwen.sh: took upstream's guarded source (the launchers do not run under `set -e`); kept local's select_model.sh call in the single-user launcher. - docs/docker.md: kept local's paragraph (upstream never had it). - docs/gotchas.md: took upstream's indentation on gotcha 58's continuation. - docker/prepare.sh auto-merged into a duplicated translate block (local hoists DIRS and passes them to harden_chat_template.py; upstream nests them); kept local's version, a superset of upstream's. Verified: bash -n on every touched shell file, test_resolution.sh passes, patches/ untouched by the merge.
Split out of #124 at @mhenrichsen's request, so that PR stays a chat-template fix and this can be reviewed on its own merits.
Why
docker/prepare.shis idempotent per step but not concurrency-safe, and concurrency is the normal case rather than user error: the entrypoint runspreparebefore every start, so a booting container racesdocker compose run --rm prepare, and two containers starting together after a crash race each other. Two runs against one model dir can interleave a shard rewrite with an index write, and the damage surfaces far from its cause — a shard missing frommodel.safetensors.index.json, a config that disagrees with the tensors on disk, or a half-fetched fast variant thatverify.shthen reports somewhere else entirely.What
docker/prepare.shtakes an exclusiveflockon<models dir>/.prepare.lock:dirname "$BASE", so the lock sits beside the model dir andBASE_MODEL_DIRcannot move it inside it. The hunk originally attached to Translate chat-template effort vocabulary so OpenAI clients do not get HTTP 400 #124 used${BASE_MODEL_DIR:-/app/models}, which put the lock inside the model directory whenever that variable was set;PREPARE_LOCK_WAIT(default 600 s) bounds the wait, and only a timeout refuses, so a legitimate concurrent prepare is waited for rather than rejected on sight;.prepare.lockfile is inert. Nothing has to clean it up, and a stale file cannot block a prepare.One correction to the underlying premise
The original hunk was rejected partly because "a stale lock now refuses to prepare".
flockis released when the holding process exits, so that failure needs a live holder to occur at all; the real cost was refusing a legitimate concurrent prepare, which the bounded wait removes. Said plainly because the fix here targets the real failure mode rather than the stated one.Docs
Gotcha 59 (main's last entry is 57, #124 claims 58, so this takes 59 — it needs renumbering if #124 has not merged first) and a bullet in
docs/docker.mdwhere the entrypoint'spreparestep is described.Verification
flockis present in the built image (util-linux 2.39.3, checked insideghcr.io/syv-ai/qwen38-27b-rtx3090:latest) and in the shell used to test. The lock lines were extracted verbatim fromdocker/prepare.shand exercised:BASE_MODEL_DIRset<models>/.prepare.lock; nothing written inside the model dirPREPARE_LOCK_WAIT=1PREPARE_LOCK_WAIT=15).prepare.lock, no holderbash -n docker/prepare.shLimits
No full
preparerun (re-downloads ~19.5 GB) and no image rebuild. The exercised lines are the script's own lock lines, run under the same shell andflockversion the image ships.