Skip to content

feat(agent): let the daemon re-authenticate an enrolled device (#175) - #296

Merged
gen16k merged 1 commit into
mainfrom
feat/175-daemon-reauth
Jul 28, 2026
Merged

gen16k merged 1 commit into
mainfrom
feat/175-daemon-reauth

Conversation

@gen16k

@gen16k gen16k commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Why

waired init on a host that already has an identity was the last input that
selected the standalone enrollment implementation. The daemon could not renew
credentials for a device it had already enrolled — Start returned its
idempotent no-op — so re-auth had nowhere else to go.

This is the behaviour half of closing #175. The deletion follows in its own PR.

The control plane needed nothing

EnrollDevice already matches an existing Device on the machine key and
renews the row (renewDeviceTx, "#115 Phase C"), and the renewal path is
exempt from the per-network device cap. setup.Enroll loads the existing
machine key rather than generating one, so it falls straight into that path.

So: agent wiring only. No proto change, no waired-side PR, no CP deploy to
wait on.

What changed

File Change
internal/management/login.go LoginStartRequest.Reauth
cmd/waired-agent/login.go Start skips the already-active short-circuit when asked; run rebuilds the session instead of publishing a second one
cmd/waired-agent/main.go the node-key rotator's teardown-then-activate becomes one rebuildSession function, shared with the login controller under one mutex
cmd/waired/init_route_daemon.go routeLocal and every fact that used to steer to it
cmd/waired/login_client.go passes reauth; names the version-skew failure
cmd/waired/init_benchmark.go disabled / stopped end the benchmark wait

chooseEnrollRoute now decides one thing: is the agent there. Its table
test is the full input space rather than a sample of it — if a later change
reintroduces an input that steers elsewhere, there is nowhere for it to steer.

Two defects fixed on the way

--auth-key on an already-enrolled host printed a successful sign-in for a
run that renewed nothing.
It took routeDaemon and hit the daemon's no-op.
Now it re-authenticates. Pinned by a test.

waitForBenchmark waited ten minutes for a model that was never coming.
It read disabled and stopped as "engine is up, a download must be in
flight" and sat on "Waiting for the model to finish downloading…" until the
deadline before reporting it had given up — on a host with local inference
switched off. waitForBundledModel has always treated those two states as
terminal (init_pull.go:97); now both do.

Measured, not estimated: install test (linux) on this branch's parent took
11m55s against 1m45s on main before the harness migration, and the
whole delta is one ten-minute wait per leg, three legs, every PR. It became
CI's problem when #290 migrated the harness to the daemon path; it has been a
real user's problem on any gateway-only host for longer than that.

waitForBenchmark is shared with waired runtimes benchmark, so that command
stops hanging on an inference-off host too.

Version skew

An agent predating reauth ignores the field (handleLoginStart decodes
leniently) and answers with the no-op status: active, no session id. The CLI
names that — "the background service is too old to renew that sign-in" —
rather than reporting the protocol symptom, or success.

What is left behind on purpose

The code behind the removed route is unreachable from this commit but still
present, marked as such. Deleting it here would have buried a ~250-line
behaviour change under ~3,300 lines of deletion.

Verification

  • gofmt -l . / go vet ./... / golangci-lint run (0 issues) /
    go test ./... -timeout 12m
  • go build -tags prod ./..., go vet -tags prod ./...,
    go test -tags prod ./internal/buildflag/...
  • make verify-cross
  • bash scripts/ci/harness-failure-strings-guard.sh,
    python3 scripts/ci/decision-log-guard.py
  • node scripts/i18n-sync.mjs --check — 32 pairs in sync

Not verified locally: a real re-auth against a live control plane. The
3-OS installtest legs exercise the enrollment path itself; the re-auth branch
is covered by the daemon-side and CLI-side tests, and by the CP-side evidence
above that renewDeviceTx is the path an existing machine key takes.

Refs #175

@github-actions

Copy link
Copy Markdown

📘 Docs preview for this PR: https://waired-docs--pr-296-a3zow9ff.web.app

Rebuilt on each push; the preview channel auto-expires in 7 days.

`waired init` on a host that already has an identity was the last input
that selected the standalone enrollment implementation. The daemon could
not renew credentials for a device it had already enrolled -- Start
returned its idempotent no-op -- so re-auth had nowhere else to go.

The control plane needed nothing: EnrollDevice already matches an
existing Device on the machine key and renews the row (renewDeviceTx,
"#115 Phase C"), and the renewal path is exempt from the device cap. So
this is agent wiring only, no proto change and no CP deploy.

LoginStartRequest gains `reauth`. The daemon honours it by skipping the
already-active short-circuit, and by rebuilding the live session
afterwards rather than publishing a second one: the running session is
using the credentials that were just replaced, and activate refuses to
publish over a current session anyway. That rebuild is the same
teardown-then-activate the node-key rotator has always done, so it is now
one function sharing one mutex with the rotator instead of two copies
that could race.

chooseEnrollRoute therefore loses routeLocal and, with it, every fact
that used to steer: no input picks an implementation any more, because
there is only one. What remains is "is the agent there", and the table
test is now the full input space rather than a sample of it.

The code behind the removed route is unreachable from this commit but
still present; #175's next PR deletes it together with the `waired init`
flags that exist only to steer it. Splitting it that way keeps this
change reviewable as a behaviour change.

Two defects fixed on the way:

* `--auth-key` on an already-enrolled host used to print a successful
  sign-in for a run that renewed nothing (routeDaemon -> the daemon's
  no-op). It now re-authenticates.
* waitForBenchmark read `disabled` and `stopped` as "engine is up, a
  download must be in flight" and waited out the full ten-minute
  deadline before giving up -- on a host with no model and no intention
  of getting one. waitForBundledModel has always treated those two as
  terminal (init_pull.go:97); now both do. Measured cost: ten minutes
  per installtest leg, three legs, every PR. Found because #290's
  harness migration made the daemon path the one CI takes.

An agent predating `reauth` ignores the field and answers with the
no-op status. The CLI names that skew ("the background service is too
old to renew that sign-in") instead of reporting the protocol symptom or,
worse, success.

Refs #175

Signed-off-by: gen16k <gen16k@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant