Skip to content

fix(agent): one engine-presence rule, and complete the engine step when the engine is installed - #205

Merged
gen16k merged 2 commits into
mainfrom
fix/179-engine-detect
Jul 27, 2026
Merged

gen16k merged 2 commits into
mainfrom
fix/179-engine-detect

Conversation

@gen16k

@gen16k gen16k commented Jul 26, 2026 •

Copy link
Copy Markdown
Contributor

The engine-detect session of waired-ai/waired#932's Wave 1 — the
highest fan-out blocker in that plan. Both issues land in
cmd/waired-agent/setup_desired.go and #187 depends on #179, so they
ship as one branch, one commit each.

Fixes #179
Fixes #187

Why

#179. The wizard's engine_installed came from the hardware
profiler's exec.LookPath probe; the daemon resolves its engine binary
by stat under the state dir. That is where waired's own install lands —
the normal case, and the only one on Linux — so /setup/state and
/inference/status contradicted each other on the same host at the same
instant. Everything downstream follows: the model pull is never
admitted, so the step contents never change, so the setup-progress dedup
suppresses every push, so last_check freezes, so the wizard renders
"no updates from this device", so the executor lease expires and the run
ends as executor_gone. A browser-driven setup showed no progress while
the terminal was visibly downloading. Third instance of this bug class
in this repo.

#187. The "Install the AI software" row stayed at "working on it"
for the whole model download: the engine step only reached done in the
ready case, and readiness there means the active model is ready. The
executor's own Done phase had no arm at all, and installed shadowed
the failed-phase arm, so a genuinely failed install on a
half-configured engine spun forever.

What changed

cmd/waired-agent/engine_resolve.go (new)

The rule now exists once, in a file named after it:

  • resolveOllamaBinary(goos, stateDir, borrowed) — extracted verbatim
    from the bootstrap's ollamaResolver closure, which now calls it.
    Takes goos rather than reading runtime.GOOS so the Linux-strict
    branch is table-testable from any runner (CLAUDE.md §Cross-OS parity).
  • vllmVenvActive(stateDir) — the <state-dir>/runtimes/vllm venv
    check, previously open-coded in two places.
  • engineInstalledOnHost(goos, stateDir, cfg, engine) — the single
    predicate, with the three-strikes history recorded at the top of the
    file so the next person reaching for exec.LookPath reads why not.

Call sites

  • setupEngineState — the wizard's answer. Also drops the profiler's
    30 s cache from this path, so an engine the executor just installed is
    visible on the next frame instead of up to half a minute later.
  • engineViable — the boot engine decision, which read the same PATH
    probe and so chose no-engine on a host whose engine waired had
    installed itself, then resolved a binary anyway. Included on the
    user's call: fixing only the wizard would have left the class alive in
    the function right below it.

snapshot's engine arm

Reordered strongest-evidence-first, and documented as the contract it
is:

ready -> phase==failed -> installed -> phase==done
      -> leaseLive && !elevated -> leaseLive -> everSeen -> default

The completion arm moves from case ready: to case installed:, and
phase == done gets the arm it never had. An engine that is serving
beats a stale failed phase from an earlier attempt; a failure the
executor reported beats mere on-disk presence, because a half-configured
install leaves a binary behind and still failed.

No new step ids — that set is a spec-level decision tracked separately
(waired-ai/waired#934).

Behaviour change worth naming

In bundled mode on Linux, a system ollama on $PATH no longer makes
the host viable at boot. It never actually was: the bootstrap refuses to
spawn it by design, because it is not the pinned engine. This only makes
the boot decision agree with what follows it — hasUsableEngine already
answered via the resolver, so subsystem_state was already correct.
reuse mode is the supported way to borrow a user's engine and is
unaffected; TestChooseEngine_BundledLinuxIgnoresSystemOllamaOnPATH
pins both directions.

Decision: pull admission stays gated on engine presence

#179's proposal item 4 asks to remove that gate. It was written against
the broken probe. The gate is the waired-ai/waired#904 guarantee that
lets the wizard write engine and model in one gesture — firing
PullModel at a host with no engine can only fail, and the operator
would watch a red "Download the AI model" for the whole install that
later turns green on its own. TestSetupPullWaitsForEngineInstall locks
it. With a correct probe plus the existing false->true re-admission
hook, item 4's intent is met without removing the belt.

Tests

Test-after rather than test-first for the extraction itself (it is a
verbatim move with the existing suite as its bar); the two regression
bars below were written against the failing behaviour.

  • engine_resolve_test.go (new, untagged so it runs on all three OSes)
    • TestEngineInstalledOnHost_StateDirEngineWithEmptyPATH — the fix(agent): setup state derives engine_installed from a PATH-only probe that cannot see the bundled engine #179
      bar: empty PATH, engine under the state dir, across all three
      goos values and all three ollama_source settings.
    • TestSetupEngineState_SeesStateDirEngineImmediately — the same bar
      at the wizard's own endpoint, and it pins the freshness half: the
      two calls straddle the install microseconds apart, which the old
      30 s-cached implementation could not have answered. Built with a
      nil profiler on purpose — this path must not touch it.
    • TestResolveOllamaBinary_LinuxBundledIsStrict /
      _NonLinuxKeepsFallback — the GOOS table. The non-Linux case
      asserts on error identity (download.ErrNotInstalled) rather than
      on finding a stub, because ResolveBinary's binary name is fixed by
      the running OS's build tag, not by the goos argument.
  • inference_choose_engine_test.go — chooseEngineProfiler's
    ollamaInstalled bool is gone (it no longer decides ollama viability;
    keeping it would have been a trap). Tests express "ollama is
    installed" on disk via fakeBundledOllama, mirroring the existing
    fakeVLLMVenv. Two new tests cover the state-dir-with-empty-PATH host
    and the bundled-vs-reuse PATH contrast.
  • setup_desired_test.go — TestSetupEngineStepMatrix walks the whole
    (installed, ready, phase) space. Every row holds a live elevated
    lease, so each arm has to beat the running fallback the bug landed
    on. Includes (true, false, failed) -> Failed, the case fix(agent): engine_install is gated on model readiness and ignores the executor's own done phase #187 names.
    TestSetupEngineStepDoneDoesNotOutrunTheModelStep pins that
    completing the engine row does not make the model row look complete.
    TestSetupSnapshotStatuses's installed -> running expectation
    updated to done — that assertion was the bug.

Verification

All of CONTRIBUTING.md §"Building and testing", green:

gofmt -l .                                  # clean
go vet ./... && (cd proto && go vet ./...)
golangci-lint run                           # 0 issues
go test ./... -timeout 10m
(cd proto && go test ./...)
go build -tags prod ./... && go vet -tags prod ./...
go test -tags prod ./internal/buildflag/...
make verify-cross                           # linux/windows/darwin amd64+arm64

Scope left out


Note on a CI failure seen before the rebase

The first run (and its rerun, which re-materialises the same merge
commit) failed TestGrantRoleGate_NoOwnerLatchFromForeignPeer in
internal/inference at waitForInFlight — a 2 s wait on goroutine
scheduling. It is not reachable from this diff: dependencies flow
cmd → internal, so internal/inference's test binary cannot compile
anything changed here, and every file in this PR is under
cmd/waired-agent/. It passes 300× locally on 4 pinned CPUs, and the
full go test ./... passes repeatedly under the same constraint;
main and the three sibling PRs open at the time were green on it.
The branch has since been rebased onto current main, which produced a
fresh run.

gen16k added 2 commits July 26, 2026 18:33
…PATH

The setup wizard's `engine_installed` came from the hardware profiler's
`exec.LookPath` probe, while the daemon resolves its engine binary by
stat under the state dir. That is where waired's own install lands — the
normal case, and the only one on Linux — so `/setup/state` and
`/inference/status` contradicted each other on the same host at the same
instant.

The consequences cascade: the model pull is never admitted, so the step
contents never change, so the setup-progress dedup suppresses every
push, so `last_check` freezes, so the wizard renders "no updates from
this device", so the executor lease expires and the run ends as
`executor_gone`. A browser-driven setup showed no progress at all while
the terminal was visibly downloading.

Third instance of this bug class in this repo, so the rule now exists
once, as `engineInstalledOnHost` in a file named after it:

  * `resolveOllamaBinary(goos, stateDir, borrowed)` — extracted verbatim
    from the bootstrap's `ollamaResolver` closure, which now calls it.
    Takes `goos` rather than reading `runtime.GOOS` so the Linux-strict
    branch is table-testable from any runner.
  * `vllmVenvActive(stateDir)` — the `<state-dir>/runtimes/vllm` venv
    check, previously open-coded.
  * `engineInstalledOnHost(...)` — the single predicate, now used by
    both `setupEngineState` (the wizard's answer) and `engineViable`
    (the boot engine decision, which read the same PATH probe and so
    chose "no-engine" on a host whose engine waired had installed
    itself, then resolved a binary anyway).

Behaviour change worth naming: in bundled mode on Linux a system ollama
on $PATH no longer makes the host viable. It never was — the bootstrap
refuses to spawn it, by design, because it is not the pinned engine —
so this only makes the boot decision agree with what follows it.

Dropping the profiler from this path also drops its 30 s cache, so an
engine the executor just installed is visible on the next frame instead
of up to half a minute later.

Pull admission stays gated on engine presence. #179 proposes removing
that gate, but it was written against the broken probe: the gate is the
waired#904 guarantee that lets the wizard write engine and model in one
gesture (`TestSetupPullWaitsForEngineInstall` locks it), and with a
correct probe plus the existing false->true re-admission hook the intent
is met without it.

Fixes #179

Signed-off-by: gen16k <gen16k@gmail.com>
…t when the model is

The wizard's "Install the AI software" row stayed at "working on it" for
the entire model download. Two independent defects in the engine_install
arm of the setup snapshot, plus an ordering bug riding along:

  * It only reached `done` in the `ready` case, and readiness there means
    the active MODEL is ready — not that the engine is installed. During
    a download the state is `installed && !ready`, which fell into the
    `installed -> running` arm indefinitely. The wizard renders a
    `running` step with no byte counts as "working on it", while the
    progress bar tracks the step that does carry bytes, so two rows
    appeared active at once.
  * There was no arm for the executor's own `Done` phase, so the explicit
    completion it reports — which exists to advance the wizard — was
    discarded.
  * `installed` shadowed the failed-phase arm. A half-configured engine
    leaves a binary on disk and still failed, so a genuinely failed
    install rendered as "working on it" forever.

The arm order is now strongest-evidence-first, and is documented as the
contract it is:

  ready -> phase==failed -> installed -> phase==done
        -> leaseLive && !elevated -> leaseLive -> everSeen -> default

An engine that is serving beats a stale failed phase from an earlier
attempt; a failure the executor reported beats mere on-disk presence.

No new step ids: the step-id set is a spec-level decision tracked
separately.

TestSetupEngineStepMatrix walks the whole (installed, ready, phase)
space with a live elevated lease on every row, so each arm has to beat
the `running` fallback the bug used to land on.

Fixes #187

Signed-off-by: gen16k <gen16k@gmail.com>
@gen16k
gen16k force-pushed the fix/179-engine-detect branch from 05dccc7 to f83d7f0 Compare July 26, 2026 18:34
@gen16k
gen16k merged commit 75646f5 into main Jul 27, 2026
14 of 15 checks passed
gen16k added a commit that referenced this pull request Aug 1, 2026
…that decides it (#138)

Linux's install.sh ran `waired runtimes install ollama` from inside
linux_apt_install — before the daemon was up, and long before anyone was
asked whether this computer should run models. In the 0.0.2-rc7 install
review the ~1.4 GB engine landed under "AI engine (Ollama)" ahead of
"Sign in and set up", and the browser wizard opened with "Install the AI
software" already done.

macOS and Windows dropped their pre-install in #55/#73; Linux kept its
state-dir one on the assumption that init took the standalone path. #119
made the daemon path the default on all three OSes, #205 fixed the
PATH-only engine_installed probe that was the recorded blocker, and the
standalone path itself has since been deleted — so the carve-out has had
no basis for a while. packaging/install/README.md and docs-site already
described the target behaviour everywhere; only the code disagreed.

`waired init` now owns both the decision and the install on Linux too:
the wizard's executor step (runSetupEngineInstall), or
ensureDaemonPathEngine when no browser is driving. Both install into the
daemon-declared state dir, which is the same
/var/lib/waired/runtimes/ollama/bin/ollama the strict bundled resolver
requires. --skip-ollama / WAIRED_NO_OLLAMA keep working; they now only
tell init to stay out of it.

Also in this change:

* the done banner MEASURES the engine line instead of being told it by an
  earlier step — same three arms and strings as darwin_next_steps
* the pre-install summary lists sign-in first and says the download
  happens only if the operator opts into running models here (mirrored
  into install.ps1's Show-InstallSummary)
* a headless Linux install stops promising "opens your web browser": init
  prints a link there (login_gate.go resolveBrowserGate), so the summary
  says so
* zstd leaves the apt prerequisites. It was installed for the retired
  upstream ollama.com/install.sh; the engine tarball's zstd layer is
  decompressed in-process now (internal/runtime/ollama_install.go)

Hosts that never reach init (--no-init, no terminal, non-systemd) finish
with no engine until the first `sudo waired init` — the same contract as
Windows -SkipInit, and the point of the gate.

Regression cover: two installtest-dash cases assert the "AI engine
(Ollama)" section and the bundled-Ollama install log are absent from a
fresh install, next to positive asserts on the new summary and banner
lines so the negatives cannot pass vacuously.

Fixes #138

Signed-off-by: gen16k <gen16k@gmail.com>
gen16k added a commit that referenced this pull request Aug 1, 2026
…that decides it (#138) (#372)

Linux's install.sh ran `waired runtimes install ollama` from inside
linux_apt_install — before the daemon was up, and long before anyone was
asked whether this computer should run models. In the 0.0.2-rc7 install
review the ~1.4 GB engine landed under "AI engine (Ollama)" ahead of
"Sign in and set up", and the browser wizard opened with "Install the AI
software" already done.

macOS and Windows dropped their pre-install in #55/#73; Linux kept its
state-dir one on the assumption that init took the standalone path. #119
made the daemon path the default on all three OSes, #205 fixed the
PATH-only engine_installed probe that was the recorded blocker, and the
standalone path itself has since been deleted — so the carve-out has had
no basis for a while. packaging/install/README.md and docs-site already
described the target behaviour everywhere; only the code disagreed.

`waired init` now owns both the decision and the install on Linux too:
the wizard's executor step (runSetupEngineInstall), or
ensureDaemonPathEngine when no browser is driving. Both install into the
daemon-declared state dir, which is the same
/var/lib/waired/runtimes/ollama/bin/ollama the strict bundled resolver
requires. --skip-ollama / WAIRED_NO_OLLAMA keep working; they now only
tell init to stay out of it.

Also in this change:

* the done banner MEASURES the engine line instead of being told it by an
  earlier step — same three arms and strings as darwin_next_steps
* the pre-install summary lists sign-in first and says the download
  happens only if the operator opts into running models here (mirrored
  into install.ps1's Show-InstallSummary)
* a headless Linux install stops promising "opens your web browser": init
  prints a link there (login_gate.go resolveBrowserGate), so the summary
  says so
* zstd leaves the apt prerequisites. It was installed for the retired
  upstream ollama.com/install.sh; the engine tarball's zstd layer is
  decompressed in-process now (internal/runtime/ollama_install.go)

Hosts that never reach init (--no-init, no terminal, non-systemd) finish
with no engine until the first `sudo waired init` — the same contract as
Windows -SkipInit, and the point of the gate.

Regression cover: two installtest-dash cases assert the "AI engine
(Ollama)" section and the bundled-Ollama install log are absent from a
fresh install, next to positive asserts on the new summary and banner
lines so the negatives cannot pass vacuously.

Fixes #138

Signed-off-by: gen16k <gen16k@gmail.com>


Verified on a real host by dispatching installtest-inference.yml (os=linux)
against the branch: the done banner reports "installed by sign-in", init then
prints "Installing the AI engine (one-time download)", and the leg still finds
the bundled binary at /var/lib/waired/runtimes/ollama/bin/ollama with the model
ready in the :9475 store. Its one failing assert (benchmark throughput) fails
the same way on main's nightly (#300).
@gen16k
gen16k deleted the fix/179-engine-detect branch August 3, 2026 19:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant