Skip to content
Merged
Show file tree
Hide file tree
Changes from 4 commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
a3630a2
Gather the audit's files under audit/, as layout 4
SiddarthAA Aug 14, 2026
99a7c7c
Point the audit-lane e2e tests at the layout-4 schedule path
SiddarthAA Aug 14, 2026
42a78d9
Resume the CTA that opened the sign-in dialog, not always the reminder
SiddarthAA Aug 14, 2026
dc5581e
Report a scheduled audit's harmful findings, so the machine can tell you
SiddarthAA Aug 14, 2026
71af070
Merge the scheduled-audit controls into the audit page, delete /settings
SiddarthAA Aug 14, 2026
069c19b
Bound a first digest to one interval, and mask secrets that arrive cut
SiddarthAA Aug 14, 2026
9b35c7a
Give scheduled audits their own page, and let the audit be a report
SiddarthAA Aug 14, 2026
345a3d4
Stop a redaction test asserting a path shape only true on my machine
SiddarthAA Aug 14, 2026
5ff4434
feat(audit): schedule audits from the CLI, with terminal sign-in
SiddarthAA Aug 14, 2026
a906ba4
feat(settings): rebuild /settings as a console for the service
SiddarthAA Aug 14, 2026
3c8ced1
fix(ui): one colour per section eyebrow; settings masthead
SiddarthAA Aug 14, 2026
7a99771
fix(settings): offer the next step when a session is already dead
SiddarthAA Aug 14, 2026
7fb1488
Stop the redactor printing the username, and three more from review
SiddarthAA Aug 14, 2026
5d9259c
Five defects from an adversarial review pass
SiddarthAA Aug 14, 2026
b47fcac
Give the schedule CLI's sign-in the frame the wizard uses
SiddarthAA Aug 15, 2026
280e835
Answer the email question up front with --schedule --email
SiddarthAA Aug 15, 2026
ce3852c
Stop the layout-4 changelog contradicting itself
SiddarthAA Aug 15, 2026
93a66ee
Stop the digest shipping assigned secrets, and two redaction misfires
NiveditJain Aug 15, 2026
087144d
Do not read an old `audit.auto` as consent to send findings off the m…
NiveditJain Aug 15, 2026
c951d10
Warn about the daemon on the command that strands it
NiveditJain Aug 15, 2026
61cc76e
Four more from the review: a stranded token, a dead session, two claims
NiveditJain Aug 15, 2026
2695960
Bring the docs to what shipped, and re-gate /settings
NiveditJain Aug 15, 2026
2c7bcf5
Seven follow-ups from the review
NiveditJain Aug 15, 2026
22b4bd9
Close the test gaps, and stop the changelog contradicting itself
NiveditJain Aug 15, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 11 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,19 @@
# Changelog

## 1.0.1-beta.0 — 2026-08-12
## 1.0.1-beta.0 — 2026-08-14

### Features

- Report harmful findings from a scheduled audit, so the machine can tell you what its agent did instead of asking you to go and look. A new `[audit] email_enabled` — a SEPARATE switch from `auto`, because `audit --help` promises the scan "runs fully offline — no account or network required" and that must stay true for anyone who wants scheduled scanning and nothing else. Off by default, like `auto`, and for a stronger version of the same reason: the failure direction is a machine mailing an account nobody pointed it at. **The window is applied per event, not through `--since`.** `--since` filters on transcript MTIME, which is right for deciding which files to open and wrong as a window: a session left open for a month has a fresh mtime, so `--since 7d` hands back that whole transcript including month-old events, and the first digest anyone received would describe everything their agent had ever done as though it happened that week. The scan stays unfiltered and the window is applied here, against the timestamps `AuditCount` already carries. Where activity straddles the boundary the report counts the EXAMPLES inside it rather than the policy's total — the cache stores counts, not event lists, so there is nothing to subtract; undercounting is the safe direction because the server's threshold reads these, and it can delay a digest but never invent one. **Harm is `deny` + `sanitize`**, plus `protect-env-vars` by hand: `severityForBuiltin` derives severity from the NAME PREFIX, so a policy that blocks `env`/`printenv` outright reads as hygiene, and its whole subject is an agent reaching for the environment — inheriting a scoring heuristic's blind spot into a security digest would be the wrong kind of consistency. Examples are redacted before they leave, against `SECRET_PATTERNS` — now exported from `builtin-policies.ts`, so blocking and redacting share one definition of "secret" rather than growing a second list beside it that eventually disagrees. Masking runs BEFORE path-shortening, because shortening can cut a path mid-token and a credential sliced in half stops matching its own pattern and ships as a fragment. `~/.failproofai/audit/machine.json` holds the machine id and the digest watermark, both `identity` class: regenerate the id and the server sees a new machine on every logout, reset the watermark and the next report re-covers months. The id is minted fresh rather than reusing `state/telemetry-id`, so opting into a digest never links the anonymous telemetry person to a verified address. The whole path runs in the audit CHILD, never the daemon — refresh rotation is theft-detecting, and keeping the token inside the audit lock is what stops a cross-process race from revoking every session a user has. Scheduled runs only, and nothing in it can fail a scan: every error is an outcome, so a dead network or an expired session leaves the local audit working and its dashboard correct. (#698)

- Gather everything the audit owns under `audit/`, as layout 4. `auth.json` becomes `audit/session.json`, `next-audit.json` becomes `audit/reminder.json`, and `state/audit-schedule.json` becomes `audit/schedule.json`, so one directory answers "what does the audit know about this machine" the way `policies/` answers it for enforcement. Two new paths join them, and the split between them is the design rather than tidiness: `session.json` holds the tokens and is `user-typed`, while `machine.json` holds this machine's report id and its digest watermark and is `identity`. Both fields have to outlive a sign-out — regenerate the id and the server sees a brand-new machine on every logout, reset the watermark and the next digest re-reports months of history as though it just happened — so they cannot live in the file a sign-out deletes. **`auditDir` is now deliberately absent from `HOME_CLASSES`.** It was classified `derived` wholesale, which was correct for a directory holding two caches and became a trap the moment a credential moved in: `resettablePaths()` is a filter over that table, so a reset and every future migration would have deleted the user's tokens. It is MIXED now and classified per-file, exactly like `state/` already is, and the `COVERED_BY_PARENT` guard records it as the second entry mapping to itself. The migration is three moves and no deletions, each a rename with a copy fallback because `audit/` and the home root land on different filesystems once `$HOME` is a network mount and `rename(2)` returns `EXDEV` there. A missing source is success — most homes never signed in, so two of the three files are absent on the majority of machines — and an existing destination wins, because re-running the step is exactly what happens when a later step in the same chain throws and the user retries; the stale original is dropped rather than left at the root, since a second copy of a bearer credential is a liability. `session.json`'s mode is reasserted to `0600` afterwards rather than assumed, because a rename preserves it and the copy fallback inherits the umask. All three are backed up first: `auth.json` is a live credential that, unlike every other file in that list, was never on a delete list and so has never had a copy taken before a migration touched it. `next-audit.json` is MOVED rather than retired even though the scheduled-audit work replaces reminders, because a migration that deleted it before that work landed would drop a cadence a person chose with no way back if the follow-up slipped. (#695)
Comment thread
SiddarthAA marked this conversation as resolved.
Outdated

### Fixes

- Resume the CTA that opened the sign-in dialog, instead of assuming it was the reminder. The reminder and "invite a friend" buttons share one `AuthDialog`, and which one opened it was tracked only as `authCopy` — the headline and subhead to show — while `handleAuthed` unconditionally called `persistReminder`. So the dialog knew which button had been pressed for the purpose of its own COPY and not for the purpose of its own EFFECT, and the invite path did the reminder path's work: a user who clicked *invite a friend*, read "Oops! Login required", and signed in got a 7-day reminder they never asked for, and no invite dialog — their actual intent dropped on the floor. An explicit `pendingAction` now carries the intent (and, for a reminder, the cadence whose button was actually pressed, so a re-render between click and verify cannot change which one lands); the copy is DERIVED from it, so the two can no longer disagree, and a third CTA means adding a case rather than remembering to branch inside a handler that has no idea it is shared. Dismissing the dialog clears the intent, because leaving it set would make the next sign-in — from any other CTA — resume something the user had walked away from; and "no pending action" is now expressible at all, which it was not before. The component's tests were the other half of the story: they covered which COPY each CTA shows and nothing else, so they were exactly as green on the broken version as on the fixed one. Three tests now pin the effect — invite resumes the invite dialog and writes no reminder, a cadence button still writes its reminder, and a dismissed dialog abandons the intent. (#698)

- Stop `detectLayout()` deriving a landmark's layout from whatever this build speaks. `config.toml` with no `config.json` returned `LAYOUT_VERSION - 1`, which read correctly while current was 3 and became silent data loss at 4: a genuine layout-2 home was reported as layout 3, so `planMigration` ran only the 3 → 4 step — which finds none of layout 3's files, moves nothing, and stamps the home as current. `config.toml` and `credentials.toml` would never be carried into JSON, orphaning the cloud token and `daemon.configured` on a machine that then reads as fully migrated. A landmark identifies ONE layout and is never relative. The `config.json` branch above it had the same shape with a different ending: that file proves "layout 3 or later" and cannot separate the two, so a layout-3 home that lost its `VERSION` was called current, the 3 → 4 move never ran, and the user was silently signed out with `auth.json` still sitting on disk. What actually separates 3 from 4 is where the audit's files sit, so it now asks that directly — any of the three still at the root means stale — and when none are present the two layouts are identical on disk, the step would move nothing, and current is the correct non-destructive answer. Found by the layout-4 bump: the assertion that caught it was pinned to `2` and started failing the moment the constant moved, which is the whole reason it was written that way. (#695)

- Move the nightly doc translation onto the canary box too, so one machine and one installer carry both scheduled jobs. Runner minutes were the entire cost of both crons; the LLM spend is identical wherever they run. The runner image already knew how to lock, check out a ref and hand off to a script from that checkout, so `$CANARY_JOB` now selects WHICH script — `jobs/canary.sh` (the integration suite, 11:00 local) or `jobs/translate.sh` (the translation, 02:00 local) — resolved to a path rather than through a case statement, so a third job is a new file in the repo and never an image rebuild. Everything per-run is keyed by job: the **lock** above all, because one shared lock lets a canary wedged on a vendor CLI swallow the night's translation and the swallow is a clean `exit 0` that reports nowhere; also the clone, since translate commits and switches branches inside its checkout, and the log. `install.sh` grew `--jobs`, per-job `--at-*` flags and one cron line per job, each behind its own marker so installing one never strips the other's; it validates credentials **per job**, so installing only the canary never demands a translation PAT, and it prints the timezone cron resolved, because "02:00" read as UTC on an IST box is 07:30 and the person reading the output is the one who would be surprised. Three things collapse in the move and are why the job is shorter than the workflow it replaces: the 14-way matrix was runner parallelism, not translation structure (cli.ts already fans out over pages x languages under one limit, so one process at `TRANSLATE_MAX_CONCURRENT=16` reproduces CI's exact peak of `max-parallel: 4` x 4 — which deletes the artifact round-trip, the per-language cache fragments and the ~35-line script that merged them); the Actions cache layer becomes a 13 KB file symlinked into the checkout from the work dir; and `consolidate`'s re-checkout-and-overlay existed only because its siblings ran on other machines. The one genuinely new credential is a push token — Actions minted a repo-scoped `GITHUB_TOKEN` that died with the job, and a box needs a long-lived fine-grained PAT, which is why it goes in a git credential helper rather than the remote URL: git echoes the remote back on a push error and the Slack crash-note carries the log tail. The translate job posts **nothing** to Slack — its output is the pull request it opens, which the PR list already says; its failures land in the run log and the exit code. The canary keeps reporting on every run including the quiet ones, so silence from it means the box did not run rather than that all was well. (#694)

- Audit the documentation weekly, on the same box. `mintlify validate` and `validate:mdx` answer "does this build", per PR, on the pages a PR touches — and pass happily on a corpus that builds perfectly and is quietly wrong: a page nobody has edited since the CLI it documents was rewritten, a page in the nav that is gone, a page in **no** nav and so unreachable by any reader, an in-body link to something renamed, a translation still describing last quarter's behaviour. None of that fails a build, which is precisely the shape a periodic sweep catches and a per-PR gate structurally cannot. `docs-audit` runs Mondays at 04:00 and posts what it found. It is the cheapest job on the box — **no gateway key, no push token, no sibling containers**, so it installs on a machine holding no credentials at all beyond the webhook — and that is deliberate: an audit that could also FIX what it finds would need write access and a much longer argument about what it may change unattended. It **reports and exits 0 by design**; `--fail-on-findings` exists for a future caller that wants a gate and is off by default, because a docs audit that turns the build red the day a page crosses an age threshold gets switched off within a week, and then there is neither a gate nor a report. It reports two ways: the weekly Slack post, and one `[auto] docs audit` tracking ISSUE kept current on GitHub — opened when there is something to do, its body refreshed each week, and closed when a week comes back clean, so an open issue always means "there is something to do" rather than "this ran once, months ago". An issue and not a PR, deliberately: a report is not a change, so a weekly PR would either sit open forever or auto-merge a file nobody reads, and an audit opening a FIXING PR would have almost nothing safe to put in it — a dangling nav entry might mean "delete the entry" or "restore the page", an orphan page might be deliberately unlisted, a broken link has no inferable target, and each is a judgement this job cannot make. Its token is correspondingly weak, `Issues: read+write` and nothing else, since it never changes a file; leave it empty and the job degrades to Slack alone. `countActionable` decides open-vs-closed and deliberately EXCLUDES stale and never-translated pages, because the nightly translation closes both by itself and counting them would hold the issue open forever — the only way a tracking issue can actually fail. The judgement lives in `scripts/docs-audit.ts` — pure functions taking the git log, the file list and the cache as arguments, so every detector is unit-tested in **both** directions (it fires on the bad case, and stays silent on the good one) without a repo, a docs tree or a clock; the shell job is only box wiring around `bun run docs:audit`, which anyone can run by hand. Two details worth knowing: it reads the same translation cache the nightly job writes, or every page would report as never-translated every week — a 672-line finding that is an artefact of where a file lives rather than a fact about the docs; and it skips link forms it cannot resolve (external, anchors, relative) rather than guessing, because the first finding nobody can reproduce is what gets the whole weekly post ignored. It also hardened the ref check. Matching the NAME against one known-stale branch (`origin/failproofaid`) only ever caught that one branch — a merged-and-deleted feature branch sailed straight through, which is exactly what was sitting in a real `secrets.env`: `CANARY_REF=origin/feat/canary-local-runner`, so the box would have tested a frozen tree forever and never said so. The installer now asks the REMOTE whether the branch still exists, which catches every deleted branch without naming any, and warns (without refusing) on anything that is not `origin/main` — legitimate for a one-off, rarely right for a cron line. Scheduling it also taught the installer to say weekly at all: a spec is now `"M H"` or a full five-field cron expression, and a job name may carry a dash (`docs-audit` is a valid path component and an invalid shell variable name), so every per-job lookup goes through one conversion rather than each site remembering. (#694)
Expand Down
2 changes: 1 addition & 1 deletion __tests__/actions/update-scheduled-audit.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,7 @@ describe("scheduled-audit write actions", () => {
expect(readConfig().telemetry.enabled).toBe(false);
expect(JSON.parse(readFileSync(configFile(), "utf8")).telemetry).toEqual({ enabled: false });
// And the audit write actually landed alongside it.
expect(readConfig().audit).toEqual({ auto: true, intervalDays: 14 });
expect(readConfig().audit).toEqual({ auto: true, intervalDays: 14, emailEnabled: false });
});

it("preserves an unrelated cloud/collector setting across a scan write", async () => {
Expand Down
Loading
Loading