Repository navigation
fix: authenticate demo box git pull with the job's own GITHUB_TOKEN - #898
Conversation
…OKEN Every deploy-demo-box run since 2026-08-13 has failed at the "Pull latest main" step with `fatal: could not read Username for 'https://github.com': No such device or address` (exit 128), so main was never actually reaching the box: the stack kept serving whatever was already deployed. The box's persistent clone at /home/sakib/hive authenticated its HTTPS remote from some credential external to this workflow. That credential worked through the 2026-08-12 20:04 UTC run (its own log shows a real fetch, not an error) and was gone by the next deploy 17 hours later, with the workflow file byte-identical across that gap. That is an expired or rotated credential, not a missing config line. Fix: stop depending on any credential stored on the box. The Pull latest main step now authenticates with this job's own ephemeral GITHUB_TOKEN via `git -c http.extraheader`, scoped to the single fetch/pull invocation and never written to disk, the same mechanism actions/checkout uses internally. Nothing here needs future rotation. GIT_TERMINAL_PROMPT=0 plus an explicit ::error:: on a failed fetch makes a bad credential fail loud and specific instead of the opaque prompt error this outage produced. actions/checkout replacing the box's persistent clone entirely was considered and rejected: every later step in this job assumes /home/sakib/hive holds the box's own untracked .env and Docker build cache, and actions/checkout's default `clean: true` would delete that .env on first use. Grepped the repo for other automated pulls on the box: none. scripts/ install.sh has a similar git fetch/checkout pair, but it targets /opt/hive for a human-run, interactive Enterprise install, not an unattended CI loop, so it is out of scope here. Buglog entry (for the follow-up buglog-only PR against main): {"error_message": "deploy-demo-box \"Pull latest main\" step: fatal: could not read Username for 'https://github.com': No such device or address (exit 128)", "root_cause": "the demo box's persistent git clone at /home/sakib/hive authenticated its HTTPS remote from a credential external to the workflow (stored PAT or credential helper on the box); that credential worked through the 2026-08-12 20:04 UTC deploy and was gone by the next deploy 17 hours later with no workflow change in between, so it expired or was rotated out from under an unattended git pull with no fallback", "fix": "authenticate the box's own fetch/pull with the deploy job's ephemeral GITHUB_TOKEN via git -c http.extraheader, scoped per-invocation and never persisted to disk, plus GIT_TERMINAL_PROMPT=0 and an explicit ::error:: so a bad credential fails loud instead of an opaque prompt error", "tags": ["ci", "deploy-demo-box", "git", "credentials", "self-hosted-runner"]}
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
Warning Review limit reached
Next review available in: 47 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Free Run ID: 📒 Files selected for processing (1)
ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Free Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
📝 WalkthroughWalkthroughThe deploy job grants read access to repository contents. Its persistent checkout validates ChangesDemo deployment authentication
Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to The change makes the demo deployment pull use the job's temporary repository token and adds clearer failure reporting; no actionable merge-blocking risk remains beyond normal checks and review. Note 🎁 Summarized by CodeRabbit FreeYour organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login. Comment |
CodeRabbit review on PR #898: the previous ::error:: text asserted the fetch failure is a credential/permission problem specifically, which would mislead debugging if a future failure is actually a missing ref, DNS, TLS, or remote-URL issue instead. Point at the preceding git error output as the real diagnostic source and keep the same two most likely places to check, without asserting the cause.
|
Adversarial review status for this PR (pipeline mode, dynamic selector). Ran: CodeRabbit CLI ( SKIPPED, structural capability gap: |
plain adversarial streamRead-only pass on Checked:
Findings: 2 MEDIUM (post-pull commit auditability, remote-identity verification), 1 LOW (inconsistent error annotation on the pull-failure path). No CRITICAL, no HIGH. Nothing here blocks merge on its own. |
Adversarial review on PR #898 found the extraheader token authenticates any public github.com repo, not just this one, so a repointed origin on the box would pull the wrong source while still reporting success. Check the remote URL before using the token, and fail loud if it does not match this repository. Also add an explicit ::error:: on the pull-failure path (only the fetch had one) and log the resulting commit SHA after a successful pull, so a future silent-stale case is visible in the run log rather than requiring a read of raw git output.
plain adversarial streamRead-only pass on Checked:
Findings: 2 MEDIUM (post-pull commit auditability, remote-identity verification), 1 LOW (inconsistent error annotation on the pull-failure path). No CRITICAL, no HIGH. Nothing here blocks merge on its own. |
…he remote URL Security review on PR #898 found three HIGH severity issues in the previous commit's fix and two MEDIUM ones worth acting on rather than answering with a rebuttal. HIGH, fixed: - The Basic-auth header was passed via `git -c ...` on the command line, world-readable through /proc/<pid>/cmdline on a stock box for the life of the fetch and pull. Switched to GIT_CONFIG_COUNT/KEY_n/VALUE_n environment variables, readable only by the same uid or root, same effect, no argv exposure. - The base64 Basic-auth blob derived from the token was never masked. GitHub only auto-masks the literal token, not a value derived from it. Added an explicit ::add-mask:: on the blob before it is used, the same step actions/checkout itself takes for the identical value. - The remote-URL guard added in the previous commit printed the box's origin URL unredacted in its error message. If that URL ever embeds a credential, the guard would leak it into the run log while checking for exactly this class of problem. Redacted once, reused everywhere the value is printed. MEDIUM, fixed: - The previous comment claimed the token is never written to disk. False: a `${{ }}` expression inside a run: block materializes into a script under the runner's temp directory, and this runner is not ephemeral. The token now reaches the script only via the GH_TOKEN env var, never interpolated into the script text. - Nothing reset credential.helper, so a green run did not actually prove the new mechanism authenticated rather than a leftover on-box credential. GIT_CONFIG_KEY_1/VALUE_1 empties it for this invocation. Also added three secret-free diagnostic lines (credential.helper state, redacted remote URL, presence of ~/.git-credentials) so the next outage does not start from the same unknowns this PR's own Verification section flagged as unprovable without box access. - pull --ff-only moves the ref and writes the worktree as two separate steps. A run killed mid-pull (this job's own timeout, a cancel, or a box reboot) can leave HEAD advanced over a half-written tree that the next run reports as current. Added a tracked-file dirty check after the pull, plus a leading `rm -f .git/index.lock` for the matching stale-lock case (safe: the workflow-level concurrency group serializes this job).
Security review on PR #898 pointed out that contents:read, added for the Pull latest main fix, also widens the job-scoped GH_TOKEN this later step already used at only actions:read. Correct and unavoidable at job-level permissions granularity, just undocumented until now. Also links issue #902, filed for the separate, pre-existing question of who can trigger this job via workflow_dispatch.
security-reviewer stream, second pass on
|
Security review, second pass on PR #898, confirmed the three HIGH argv/ mask/redaction fixes hold in the actual code, then found three more: MEDIUM: `git config --show-origin --get-all credential.helper` prints the helper's value, and a helper can be an inline shell snippet with a PAT embedded, exactly the on-box credential this PR suspects. Switched to `--name-only --get-regexp`, which reports that a helper is configured without printing what it contains. MEDIUM: the unconditional `rm -f .git/index.lock` could stomp a lock held by a genuinely running process on this shared checkout (git gc --auto, a human on the box), since the workflow-level concurrency group only serializes this workflow's own runs against each other, not every writer of this clone. Age-gated to `find -mmin +15 -delete`, matching this job's own timeout-minutes. LOW: the redaction regex stopped at the first @, which would leave part of a credential exposed if the secret itself contained an unencoded @. Made greedy to match the last @ instead. Also added a `.netrc` presence check alongside the existing `.git-credentials` one, and an explicit ::error:: on an empty GH_TOKEN instead of silently building a useless auth header from it.
Security review pointed out --name-only alone drops the config file origin that made this diagnostic useful in the first place. --show-origin combines safely with --name-only (origin path, still no value), so add it back: reports whether a helper is configured AND which file configures it, with nothing printable left.
Summary
Every
deploy-demo-boxrun since 2026-08-13 has failed at the "Pull latest main" step:Exit 128 at the pull means the recreate never ran, so main never actually reached the box: the stack keeps serving whatever was already deployed, silently, behind a red check.
Root cause
The
deployjob never checks out the repo itself; itcds into the box's own long-lived clone at/home/sakib/hiveand runsgit checkout main; git fetch origin; git pull --ff-only origin main. That clone's HTTPS remote authenticates from some credential external to this workflow (a stored PAT or a git credential helper on the box; this job has no shell access to the box outside its own steps to inspect which).Evidence this is an expired or rotated credential, not a missing config line:
31635783176(2026-08-12 20:04 UTC): thedeployjob's own log showsgit fetchsucceeding with a real ref listing (From https://github.com/sakibsadmanshajib/hive, multiple branch updates), i.e. the credential worked.31704949535(2026-08-13 13:27 UTC, ~17 hours later): the identicalgit fetch origincall fails with the Username error above.git log -- .github/workflows/deploy-demo-box.ymlbetween the commits at those two runs (3549971e..7faae2d3) is empty: the workflow file is byte-identical across the gap. Nothing in this repo changed to cause the failure.That combination (same command, same file, works then stops working with no repo change in between) is a credential that expired or was rotated out from under an unattended, non-interactive
git pull, which then hits git's default terminal-prompt path and fails fast since the runner has no tty.Fix
Stop depending on any credential stored on the box at all. The "Pull latest main" step now authenticates the fetch/pull with this job's own ephemeral
GITHUB_TOKEN, viagit -c http.extraheader, scoped to the single invocation and never written to disk. This is the same mechanismactions/checkoutuses internally. Addedcontents: readto the job'spermissionsblock (it previously only hadactions: read) so the token can read repo contents. Nothing here needs future rotation: the token is minted fresh every run.Also added:
GIT_TERMINAL_PROMPT=0so any future bad credential fails in one git-native line instead of the opaque "could not read Username / No such device or address" this outage actually produced.::error::on a failed fetch naming the real failure class (credential/permission, not a missing branch) and where to look on the box.Alternative considered and rejected
Replacing the box's persistent clone with
actions/checkoutentirely (having the runner's own already-authenticated checkout do the fetch) was considered first, since it's the more architecturally clean fix. Rejected for this PR: every later step in thedeployjob assumes/home/sakib/hiveis a long-lived directory carrying the box's own untracked.env(every secret this stack runs on, never committed) and Docker build cache.actions/checkout's defaultclean: truerunsgit clean -ffdx, which would delete that untracked.envon first use. That is a materially bigger and riskier change than fixing the credential the existing clone already uses, and cannot be safely verified before merge (see Verification below).Same blast radius, checked
Grepped the repo for every other place a workflow or script pulls the demo box's own clone the same way. Only
deploy-demo-box.ymldoes this in an unattended CI context.scripts/install.shhas a similargit fetch/git checkout main --quietpair, but it targets/opt/hivefor a human-run, interactive Enterprise self-host install (usesprompt_valuehelpers elsewhere in the same script), not an automated redeploy loop, so it is out of scope for this fix.Buglog entry
{"error_message": "deploy-demo-box \"Pull latest main\" step: fatal: could not read Username for 'https://github.com': No such device or address (exit 128)", "root_cause": "the demo box's persistent git clone at /home/sakib/hive authenticated its HTTPS remote from a credential external to the workflow (stored PAT or credential helper on the box); that credential worked through the 2026-08-12 20:04 UTC deploy and was gone by the next deploy 17 hours later with no workflow change in between, so it expired or was rotated out from under an unattended git pull with no fallback", "fix": "authenticate the box's own fetch/pull with the deploy job's ephemeral GITHUB_TOKEN via git -c http.extraheader, scoped per-invocation and never persisted to disk, plus GIT_TERMINAL_PROMPT=0 and an explicit ::error:: so a bad credential fails loud instead of an opaque prompt error", "tags": ["ci", "deploy-demo-box", "git", "credentials", "self-hosted-runner"]}Verification
Verified:
python3 -c "import yaml; yaml.safe_load(...)").git -c http.https://github.com/.extraheader="AUTHORIZATION: basic ...") matchesactions/checkout's own internal implementation for HTTPS token auth, a well established pattern.git logon the workflow file across that gap; not guessed.Could not verify (self-hosted runner on a physical box, no local reproduction possible):
contents: read-scopedGITHUB_TOKENwith the permissions this job declares, end to end.https://github.com/sakibsadmanshajib/hive.git(assumed from the existing error message and job comments; the new::error::step surfacesgit remote -vguidance if this assumption is wrong)..netrc, a globalcredential.helperthat might intercept before the per-invocation-cflag takes effect) interferes.git -cconfig takes precedence over global config for the same key, but this is asserted from git's documented config precedence, not observed on the box.The only real proof is the next
deploy-demo-boxrun on main after this merges. Watch the "Pull latest main" step specifically: it should show a realgit fetch/pullref update (like the 2026-08-12 20:04 run) instead of the Username error, and the deploy should proceed to recreate the stack.Test plan
deploy-demo-boxrun (gh run watch)CREATEDtimestamps in the "Dump container status" step or a follow-updocker compose ps)Summary by CodeRabbit