feat(juicefs): Step 4 — plumb the tailnet-bound DB port the cross-node lane needs - #2728
Conversation
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Docker Hardening ValidationHardening Validation ReportValidated: Mon Aug 24 21:03:06 UTC 2026Services CheckedPMOVES.AI Docker Hardening Validation[INFO] Checking: pmoves/docker-compose.hardened.yml [INFO] Validating: hi-rag-gateway-v2 [INFO] Validating: extract-worker [INFO] Validating: langextract [INFO] Validating: presign [INFO] Validating: render-webhook [INFO] Validating: retrieval-eval [INFO] Validating: pdf-ingest [INFO] Validating: jellyfin-bridge [INFO] Validating: invidious-companion-proxy [INFO] Validating: ffmpeg-whisper [INFO] Validating: media-video [INFO] Validating: media-audio [INFO] Validating: hi-rag-gateway-v2-gpu [INFO] Validating: hi-rag-gateway-gpu [INFO] Validating: deepresearch [INFO] Validating: supaserch [INFO] Validating: publisher-discord [INFO] Validating: mesh-agent [INFO] Validating: nats-echo-req [INFO] Validating: nats-echo-res [INFO] Validating: comfy-watcher [INFO] Validating: grayjay-plugin-host [INFO] Validating: agent-zero [INFO] Validating: p7-room-orchestrator [INFO] Validating: archon [INFO] Validating: channel-monitor [INFO] Validating: pmoves-yt [INFO] Validating: notebook-sync [INFO] Validating: supabase_service_role_key [INFO] Validating: supabase_jwt_secret [INFO] Validating: p7_control_token ====================================== |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f75ce5cf83
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Hold this — pending the outcome of #2730. This PR multi-homes If metadata follows, this is pointed the wrong way. It isn't wrong — it's scoped to a topology now under review in #2730, which reopens the two Separately, and independent of the placement decision: the volume's recorded bucket is |
…sting Both §4.1 choices were correct for the requirement as written in June. The requirement changed, and two facts landed since, so they are reopened rather than quietly outgrown. NEW REQUIREMENT (operator, 2026-08-24): Jellyfin, PMOVES.YT and media hosted on the KVMs, which is also where egress lives to bypass a slow local uplink; local nodes keep working during a VPS outage; no single node offline stops operators viewing. That is a placement-and-availability requirement. §4.1 chose for a single-node stack. WHAT IS ACTUALLY DEPLOYED. The spec has JuiceFS replacing MinIO via `juicefs gateway`. pmoves-media instead runs ON TOP of MinIO -- Storage "minio", Bucket "http://minio:9000/juicefs" -- with Postgres metadata on b850 and pmoves-minio-1 also on b850. Both halves of the storage layer on one workstation; no VPS anywhere. This is the Status line playing out: the replacement gateway was deferred, the content mount shipped meanwhile and picked the deprecated store as its backend, and because both are "JuiceFS" the stack reads as though MinIO had already been replaced. RECORDED BUCKET CANNOT RESOLVE OFF-HOST. `minio:9000` is a Docker-internal hostname baked in at format time. docker-compose.juicefs.yml:39 anticipated exactly this and the single-node fallback was used anyway. The existing guard does not catch it: juicefs-cross-node-setup.sh:73-82 refuses Storage "file", and ours is "minio", so preflight PASSES and the mount fails on open -- the precise failure the guard exists to prevent. A remote node running its own `minio` container would silently bind a different object store. MINIO IS ARCHIVED, NOT MERELY EOL. Repository archived read-only April 2026: no releases, no reviewed patches, no community binaries, console reduced to a read-only browser. The #1862 last-real-tag pin is now a permanently frozen base. DATA BACKEND reopened: `file://` is explicitly single-node, caused the 2026-08-04 blocker, and is now refused by the preflight. Candidate Garage, whose stated design goal is "a lightweight geo-distributed data store ... made for multi-sites (datacenters, offices, households) interconnected through regular Internet connections" -- this fleet's topology stated literally. Its non-goals cost nothing here: it does not emulate POSIX, and JuiceFS supplies that above it. R2 noted as the managed alternative given CF is already wired. METADATA ENGINE reopened: the June rationale included "supports the multi-node vision". It does not. JuiceFS's docs discuss no replication/failover/HA for PostgreSQL, and best-practices says outright not to use a multi-server distributed architecture for PG metadata. Postgres is a correct single-server choice and an incorrect HA one. TiKV is the only engine the docs push toward a dedicated cluster. Migration is bounded: `juicefs dump | juicefs load`, object data untouched, writes suspended for the cutover. THE CONSTRAINT: one strongly-consistent metadata store cannot be writable on both sides of a partition. What the docs do offer is asymmetric and worth stating precisely rather than promising symmetry -- cached reads survive an outage, and `--writeback` buffers writes locally at a documented data-loss risk. True both-directions availability needs two filesystems with object-level replication, which should be chosen deliberately, not discovered. Consequence for work in flight: PR #2728 assumes b850 stays the metadata host. It is not wrong, it is scoped to a topology now under review, and should not merge on momentum. Deliberately decides nothing. §4.1 stays unedited as the record of the June decision; this replaces "confirmed" with the evidence needed to choose again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FkwiW3VY1xWmahTAtVioxz
…sting (#2730) * docs(arch): reopen the two "confirmed" storage choices against KVM hosting Both §4.1 choices were correct for the requirement as written in June. The requirement changed, and two facts landed since, so they are reopened rather than quietly outgrown. NEW REQUIREMENT (operator, 2026-08-24): Jellyfin, PMOVES.YT and media hosted on the KVMs, which is also where egress lives to bypass a slow local uplink; local nodes keep working during a VPS outage; no single node offline stops operators viewing. That is a placement-and-availability requirement. §4.1 chose for a single-node stack. WHAT IS ACTUALLY DEPLOYED. The spec has JuiceFS replacing MinIO via `juicefs gateway`. pmoves-media instead runs ON TOP of MinIO -- Storage "minio", Bucket "http://minio:9000/juicefs" -- with Postgres metadata on b850 and pmoves-minio-1 also on b850. Both halves of the storage layer on one workstation; no VPS anywhere. This is the Status line playing out: the replacement gateway was deferred, the content mount shipped meanwhile and picked the deprecated store as its backend, and because both are "JuiceFS" the stack reads as though MinIO had already been replaced. RECORDED BUCKET CANNOT RESOLVE OFF-HOST. `minio:9000` is a Docker-internal hostname baked in at format time. docker-compose.juicefs.yml:39 anticipated exactly this and the single-node fallback was used anyway. The existing guard does not catch it: juicefs-cross-node-setup.sh:73-82 refuses Storage "file", and ours is "minio", so preflight PASSES and the mount fails on open -- the precise failure the guard exists to prevent. A remote node running its own `minio` container would silently bind a different object store. MINIO IS ARCHIVED, NOT MERELY EOL. Repository archived read-only April 2026: no releases, no reviewed patches, no community binaries, console reduced to a read-only browser. The #1862 last-real-tag pin is now a permanently frozen base. DATA BACKEND reopened: `file://` is explicitly single-node, caused the 2026-08-04 blocker, and is now refused by the preflight. Candidate Garage, whose stated design goal is "a lightweight geo-distributed data store ... made for multi-sites (datacenters, offices, households) interconnected through regular Internet connections" -- this fleet's topology stated literally. Its non-goals cost nothing here: it does not emulate POSIX, and JuiceFS supplies that above it. R2 noted as the managed alternative given CF is already wired. METADATA ENGINE reopened: the June rationale included "supports the multi-node vision". It does not. JuiceFS's docs discuss no replication/failover/HA for PostgreSQL, and best-practices says outright not to use a multi-server distributed architecture for PG metadata. Postgres is a correct single-server choice and an incorrect HA one. TiKV is the only engine the docs push toward a dedicated cluster. Migration is bounded: `juicefs dump | juicefs load`, object data untouched, writes suspended for the cutover. THE CONSTRAINT: one strongly-consistent metadata store cannot be writable on both sides of a partition. What the docs do offer is asymmetric and worth stating precisely rather than promising symmetry -- cached reads survive an outage, and `--writeback` buffers writes locally at a documented data-loss risk. True both-directions availability needs two filesystems with object-level replication, which should be chosen deliberately, not discovered. Consequence for work in flight: PR #2728 assumes b850 stays the metadata host. It is not wrong, it is scoped to a topology now under review, and should not merge on momentum. Deliberately decides nothing. §4.1 stays unedited as the record of the June decision; this replaces "confirmed" with the evidence needed to choose again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FkwiW3VY1xWmahTAtVioxz * docs(arch): correct two claims this revision got wrong, and record the decisions Both Codex findings are right, and both overturn things the first draft asserted confidently. §0.6 SAID LAB NODES WOULD "DEGRADE TO CACHED READS + BUFFERED WRITES" DURING A VPS OUTAGE. They will not. `--writeback` defers object-BLOCK uploads; it does not queue metadata TRANSACTIONS, and a block cache is not a metadata service. Open, create, rename and stat are all metadata operations, so an unreachable metadata engine does not degrade the mount, it fails it. The repo already said so and the draft contradicted a checked-in runbook: "The JuiceFS metadata + storage home is B850 ... the mounts below FAIL if B850 is not up and reachable over the tailnet." -- JUICEFS_CROSS_NODE_MOUNT_RUNBOOK.md Replaced with a blunt availability table: during a VPS outage lab nodes lose pmoves-media entirely, local work is unaffected, operators keep viewing. The caching options are still worth setting over a slow uplink -- for BANDWIDTH, not availability. Saying otherwise would have sold an outage mode that does not exist. §0.5 CALLED POSTGRES "AN INCORRECT HA ENGINE". Over-read of the vendor warning. "Do not use a multi-server distributed architecture" is about multi-writer and sharded layouts; a replicated primary with standby failover behind ONE endpoint still presents JuiceFS a single writable Postgres, so the warning does not apply. The true statement is narrower: the current single supabase-db is not HA, which is a deployment gap rather than an engine verdict. Replicated Postgres is now listed FIRST among candidates -- it is the only one needing no metadata migration, so if it satisfies the requirement the whole engine change is avoidable. That correction may save the migration this revision was heading toward. OPERATOR DECISIONS recorded in a new §0.8: asymmetric availability ACCEPTED with the corrected cost in view, and object substrate = Garage self-hosted on the KVMs (over managed R2), keeping the egress economics the KVMs exist for. Metadata engine stays open and is now the first question rather than a foregone one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FkwiW3VY1xWmahTAtVioxz --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Docker Hardening ValidationHardening Validation ReportValidated: Tue Aug 25 06:39:06 UTC 2026Services CheckedPMOVES.AI Docker Hardening Validation[INFO] Checking: pmoves/docker-compose.hardened.yml [INFO] Validating: hi-rag-gateway-v2 [INFO] Validating: extract-worker [INFO] Validating: langextract [INFO] Validating: presign [INFO] Validating: render-webhook [INFO] Validating: retrieval-eval [INFO] Validating: pdf-ingest [INFO] Validating: jellyfin-bridge [INFO] Validating: invidious-companion-proxy [INFO] Validating: ffmpeg-whisper [INFO] Validating: media-video [INFO] Validating: media-audio [INFO] Validating: hi-rag-gateway-v2-gpu [INFO] Validating: hi-rag-gateway-gpu [INFO] Validating: deepresearch [INFO] Validating: supaserch [INFO] Validating: publisher-discord [INFO] Validating: mesh-agent [INFO] Validating: nats-echo-req [INFO] Validating: nats-echo-res [INFO] Validating: comfy-watcher [INFO] Validating: grayjay-plugin-host [INFO] Validating: agent-zero [INFO] Validating: p7-room-orchestrator [INFO] Validating: archon [INFO] Validating: channel-monitor [INFO] Validating: pmoves-yt [INFO] Validating: notebook-sync [INFO] Validating: supabase_service_role_key [INFO] Validating: supabase_jwt_secret [INFO] Validating: p7_control_token ====================================== |
Register entries for the claude-review pin bump (this PR) and the PR #2728 compose regen, with CI evidence for both.
Register entries for the claude-review pin bump (this PR) and the PR #2728 compose regen, with CI evidence for both.
Register entries for the claude-review pin bump (this PR) and the PR #2728 compose regen, with CI evidence for both.
…it crash (#2746) * fix(ci): bump claude-code-action v1.0.193 -> v1.0.205 to clear the init crash claude-review failed on every open PR since ~2026-08-23 with is_error:true / num_turns:1 / cost:0 at SDK init — upstream regression anthropics/claude-code-action#1720 (reproduces v1.0.160–v1.0.201). The "directory mismatch for tsconfig.json" line is Bun oven-sh/bun#25730 surfacing through the same crash (#1568). Upstream shipped CLI/SDK bumps v1.0.202–205 (Claude Code 2.1.241–2.1.245) in the two days after the reports; this takes the newest. All four workflows that pin the action move together so the other review lanes cannot hit the same wall. Our workflow inputs were verified against upstream docs/usage.md first — plugin_marketplaces/plugins/prompt/oauth-token are all documented; the config was fine, the pin simply predates the regression. * docs(agnote): CLAIM + RELEASE the CI unblock pair lane Register entries for the claude-review pin bump (this PR) and the PR #2728 compose regen, with CI evidence for both. * docs(agnote): amendment — reruns + CodeQL suppression greened the three PRs without waiting for the pin merge
…e lane needs
The compose has published this port since March:
ports: - ${SUPABASE_DB_BIND:-127.0.0.1}:${SUPABASE_DB_PORT:-54322}:5432
but it has never been reachable, and the reason is not the bind. supabase-db sits
on pmoves_data and pmoves_api, both `internal: true`. A published port on a
container attached only to internal networks maps and then answers nothing.
Multi-homing onto pmoves_external is what actually plumbs it — the pattern the
NATS bus already uses, and the one the lane handoff specified.
This does not put a database on the open internet:
- SUPABASE_DB_BIND is set per node to that node's TAILNET address, never
0.0.0.0, so the listener exists only on the tailnet interface. The default
stays 127.0.0.1, so any node that does not set it publishes nothing new and
this change is inert there.
- pg_hba (supabase/config/pg_hba.conf, #2702) admits ONLY juicefs_meta from
100.64.0.0/10 and REJECTS every other role there. That is the control that
keeps this from being a superuser surface. The scoped role alone never was —
it changes which credential JuiceFS uses, not which one is accepted.
sslmode is a recorded decision rather than an inherited default, which is what
JUICEFS_META_CREDENTIAL_RUNBOOK.md §1.3 asked for at this step: the cross-node
DSN carries sslmode=disable, JuiceFS's own PostgreSQL best-practices advise
against it, and the justification for keeping it is that WireGuard encrypts the
tailnet transport and the metadata role is non-superuser and pg_hba-scoped.
Revisit if the DB is ever reachable off-tailnet.
Operators should publish on 5432, not the 54322 default: PORT_REGISTRY already
assigns 5432 to this service, and juicefs-cross-node-setup.sh:28 defaults
DB_PORT=5432, so the canonical port means the remote node needs no override.
Registry row updated from "internal only" to reflect that.
Verified: compose parses; supabase-db resolves to
[pmoves_data, pmoves_api, pmoves_external]; split overlays regenerated with the
pinned toolchain; both 5432 and 54322 are free on this host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FkwiW3VY1xWmahTAtVioxz
Step 4 hand-edited the networks: lists but not the generated PMOVES_NETWORKS env that topology.TopologyContext.from_env() reads — the injector drift gate failed. Re-ran the injector and compose-split; one mirrored env line per file, no other changes.
Review P1: joining supabase-db to pmoves_external handed every internet-facing container on that bridge a direct route to supabase-db:5432, and pg_hba's 172.16.0.0/12 catch-all accepts every role from any docker bridge -- the tailnet scoping never sees those sources. Replaced with pmoves_db_egress (172.30.8.0/24, internal: false, external network carrying ONLY the database): - pg_hba scopes the new subnet above the catch-all: juicefs_meta only, everything else rejected -- identical treatment to the tailnet rules. Legitimate inbound still DNATs in with its real 100.x source and hits the 100.64.0.0/10 block; only a container that joined this bridge sources from 172.30.8.x, and that bridge is supposed to hold the database alone. - Subnet is .8, not .7: pmoves_public is compose-declared at 172.30.7.0/24 (docker-compose.yml) -- .7 would collide. - Network ensured alongside the others in ensure-networks and ensure-overlay-networks; overlays regenerated via injector + split. The 5090's reachability Monitor is unaffected: SUPABASE_DB_BIND still pins the listener to the node's tailnet address, and the dedicated bridge is what makes the published port answer at all.
cc3c2c7 to
d287702
Compare
|
P1 addressed in |
Docker Hardening ValidationHardening Validation ReportValidated: Tue Aug 25 15:04:47 UTC 2026Services CheckedPMOVES.AI Docker Hardening Validation[INFO] Checking: pmoves/docker-compose.hardened.yml [INFO] Validating: hi-rag-gateway-v2 [INFO] Validating: extract-worker [INFO] Validating: langextract [INFO] Validating: presign [INFO] Validating: render-webhook [INFO] Validating: retrieval-eval [INFO] Validating: pdf-ingest [INFO] Validating: jellyfin-bridge [INFO] Validating: invidious-companion-proxy [INFO] Validating: ffmpeg-whisper [INFO] Validating: media-video [INFO] Validating: media-audio [INFO] Validating: hi-rag-gateway-v2-gpu [INFO] Validating: hi-rag-gateway-gpu [INFO] Validating: deepresearch [INFO] Validating: supaserch [INFO] Validating: publisher-discord [INFO] Validating: mesh-agent [INFO] Validating: nats-echo-req [INFO] Validating: nats-echo-res [INFO] Validating: comfy-watcher [INFO] Validating: grayjay-plugin-host [INFO] Validating: agent-zero [INFO] Validating: p7-room-orchestrator [INFO] Validating: archon [INFO] Validating: channel-monitor [INFO] Validating: pmoves-yt [INFO] Validating: notebook-sync [INFO] Validating: supabase_service_role_key [INFO] Validating: supabase_jwt_secret [INFO] Validating: p7_control_token ====================================== |
…850 fallback producer - §2: sequencing (data move first, two freezes, never combined), secret key re-injection after load, all-mounts freeze, multi-host DSN passes through but failover COULD-NOT-MEASURE (max_life_time=0; sandbox test), Patroni per rto-rpo-targets.md:86, follow-on plan scope. Review P2-4: #2728 is MERGED; interim path, retire the tailnet exposure after move. - §1.4: D2 options under the fleet-wide layout, not decided. - §1.6: delivery to every storage node is COULD-NOT-MEASURE (G2); Spark is never the sole bundle producer, b850 is the named fallback. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…de); CNM markers Review P2-4 in the parent doc; key import syntax and MagicDNS-in-bridge kept visible as COULD-NOT-MEASURE. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…MinIO bridge) (#3200) * docs(juicefs): scaffold Garage migration plan (WIP) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): Garage migration plan + runbook (replaces the interim MinIO bridge) Plan only; nothing executed. Target: Garage v2.4.1 on the 3 KVMs, RF=3 (writes survive one node down; RF=2 goes read-only), s3_region us-east-1 to match JuiceFS's default, bucket-scoped key, 5 funnel labels with shape checks and the key-print hazards. Metadata engine is gate D1 (data move alone gives the inverse of the accepted asymmetry). Runbook steps 0 and (a)-(e) each carry a verification gate; rollback targets the untouched MinIO bucket. KVM capacity and tailnet throughput are COULD-NOT-MEASURE (Tailscale SSH refused) with operator commands; tailscale ping RTTs recorded. Adds a §0.9 pointer in JUICEFS_OBJECT_STORE_MIGRATION.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): fix #3200 review P1-1..3 (sync flags, URL scheme, freeze length) - P1-1: drop --enable-checkpoint (absent in ce-v1.3.0; v1.4.x only). Re-runs are incremental by size, not resumable. - P1-2: Garage side of every sync uses minio:// + --no-https; s3:// with a MagicDNS host is parsed virtual-host style. Gate A gains a --dry sync that proves the parse before pass 1. - P1-3: --check-all stays in Gate B, outside the freeze; c2 is a --check-new delta. Time/bandwidth table carries the check-all egress. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): measured KVM capacity (D3 open), tailnet binds, external port-probe gate - §1.1: Hostinger REST, read-only, 2026-09-27: kvm2 KVM 2 / 100 GB, kvm4-1 and kvm4-2 KVM 4 / 200 GB, all data_center_id 17, no extra volumes. RF=3 with a full copy per node is INFEASIBLE as specified; options (a) KVM 8 upgrade, (b) RF=2 on kvm4s, (c) add a disk node are listed, not chosen. Free-after-OS pending the peer survey. - §1.3 (review P3-1, P3-2, P3-3 names): snapshots dir is a sibling of data_dir; S3/admin bind to the tailnet address; pmoves_ secret names. No Hostinger rule mentions 3900/3901/3903 and default policy is not exposed, so an external probe of all three ports is a required gate. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): §1.6 secrets: real exposure window, KVM delivery unconfirmed, pmoves_ names Review P2-2 (exposure window, history), P2-5 (KVM delivery vehicle is COULD-NOT-MEASURE, G2 item), P3-3 (pmoves_ docker secret names; manifest enforces only min_length/prefix), and the D1 consequence that JUICEFS_GARAGE_SECRET_KEY must stay deliverable for re-injection after load. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): fleet-wide Garage mesh, decided tiers, measured KVM survey (D3 open) Operator direction 2026-09-27: every capable node is a Garage storage node and JuiceFS client over the mesh. Tiers DECIDED: tier 1 always-on = KVMs + Spark (Spark DOWN on 2026-09-27, recorded as a live example of the risk); tier 2 = 5090, Z890, Knuckles, 4090. - §1.0 prior art: UNIVERSAL_MEDIA_ARCHITECTURE, MEDIA_DATA_ARCHITECTURE_PLAN, FLEET_ACCESS_NATS_HUB §4 sidecar, #2288, egress.mk:318, parent §0.4. - §1.1 KVM shell survey (root, read-only): free now 78/15/31G; unused volumes are NOT garbage -> HARD caution + G0 identification item. - §1.2 capacity semantics corrected from Garage v2.4.1 layout docs (smallest-node rule only for exactly 3 zones); layout has no priority attribute, so tiers are expressed via capacity weighting (applied) and optional zone grouping (OPEN); arithmetic of today's tension shown. - §1.3a Windows nodes (Docker Desktop, sidecar, funnel): COULD-NOT-MEASURE. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): §2 D1 DECIDED replicated Postgres; §1.4 D2 reopened; b850 fallback producer - §2: sequencing (data move first, two freezes, never combined), secret key re-injection after load, all-mounts freeze, multi-host DSN passes through but failover COULD-NOT-MEASURE (max_life_time=0; sandbox test), Patroni per rto-rpo-targets.md:86, follow-on plan scope. Review P2-4: #2728 is MERGED; interim path, retire the tailnet exposure after move. - §1.4: D2 options under the fleet-wide layout, not decided. - §1.6: delivery to every storage node is COULD-NOT-MEASURE (G2); Spark is never the sole bundle producer, b850 is the named fallback. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): §3 per-set env files, live-mount-shape remount, stale sessions, rollback invariant Review P2-2 (jfs() takes credential sets via --env-file from a tmpfs umask-077 dir; history off), P2-1 (Step 0 records the live mount shape; c5 remounts through the same target, not juicefs-mount-local; gate on status sessions + recorded role/network; #3150 precondition), P3-5 (c1 stale sessions: wait for expiry, never force), P2-3 (invariant is "no writes after c1 except the carry-back"; Gate D gc scan > 0). Step (a)/Gate A generalised to the fleet-wide first cut. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): soak/failover split, measured KVM RTT, gates and decisions updated Review P2-6 (node-down test on a non-endpoint node; endpoint failover is its own gated drill G6a), P3-4 (drop dd-over-ssh; per-pair throughput plan; KVM<->KVM RTT measured 1-5 ms direct; iperf3 not installed, method OPEN D8). Risks, gates (G0: #3150 + KVM volume identification; G2: per- node delivery, b850 fallback) and the decisions table reflect D1 decided, tiers decided, D3 per-node declarations OPEN, D7 zone grouping OPEN. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): parent §0.7-0.9 (PR #2728 merged, D1 decided, fleet-wide); CNM markers Review P2-4 in the parent doc; key import syntax and MagicDNS-in-bridge kept visible as COULD-NOT-MEASURE. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): fold re-review N1-N4 + Hostinger 7-day KVM metrics - N1: Garage key reaches sync via the minio:// env fallback (MINIO_ACCESS_KEY/SECRET_KEY in the dst env file, no userinfo on the Garage URL); only the MinIO credential stays in argv. Gate A --dry uses the exact form. - N2: zone redundancy must print "maximum" (Gate A, G1); never run `layout config -r`; §1.2 notes the one cluster-level parameter. - N3: partition share is bounded by, not proportional to, capacity. - N4: c5 remount shape is the #3150 path; cross-node-setup.sh on main also uses --network host. - Minor: tier-2 replication lands on residential downlinks; check-all reads can come from any replica holder. - D3/D8: Hostinger 7-day metrics (kvm4s +35/+64 GiB/7d, unknown writer owned outside this plan; peak <=39 Mbit/s, not the binding constraint). B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(juicefs): measured iperf3, MagicDNS-in-bridge, kvm4-1 inbound blocker, BuildKit cache - Time table re-derived at the measured ~18 Mbit/s Knuckles uplink (~10.5 h per 85 GB pass; 100 Mbit/s kept as hypothetical). KVM<->KVM 345-383 Mbit/s. D8 MEASURED. - MagicDNS does not resolve inside a Docker bridge (measured): runbook containers use the host-resolved tailnet IPv4 or host networking; the choice is recorded in Gate A. - kvm4-1 accepted no inbound tailnet connection (cause COULD-NOT-MEASURE): G0 / Gate A BLOCKER for kvm4-1 as a Garage node. - KVM growth is BuildKit state (buildx_buildkit_pmoves-shared0_state, ~140/146 GB); other unused volumes are empty. Reclaim through the build-cache road, never volume deletion; owned outside this plan. B850-CLAUDE-FUNNEL (Knuckles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Step 4 of the juicefs cross-node lane. Unblocks the 5090, which is holding a Monitor on
:5432reachability.The port was never the problem
Compose has published it since March:
But
supabase-dbsits onpmoves_data+pmoves_api, bothinternal: true. A published port on a container attached only to internal networks maps and then answers nothing. Multi-homing ontopmoves_externalis what actually plumbs it — the pattern the NATS bus already uses, and what the lane handoff specified.This does not put a database on the internet
SUPABASE_DB_BIND0.0.0.0— the listener exists only on the tailnet interface. Default stays127.0.0.1, so this change is inert on any node that doesn't set it.pg_hba(#2702)juicefs_metafrom100.64.0.0/10, rejects every other role thereThat second row is the one that matters. The scoped role alone was never the control — it changes which credential JuiceFS uses, not which one is accepted.
sslmode — decided, not inherited
JUICEFS_META_CREDENTIAL_RUNBOOK.md§1.3 asked for this call to be made at Step 4. The cross-node DSN carriessslmode=disable; JuiceFS's own PostgreSQL best practices advise against it.Keeping it, on the record: WireGuard encrypts the tailnet transport, and the metadata role is non-superuser and pg_hba-scoped. Revisit if the DB ever becomes reachable off-tailnet.
Publish on 5432, not the 54322 default
PORT_REGISTRYalready assigns 5432 to this service, andjuicefs-cross-node-setup.sh:28defaultsDB_PORT=5432— so the canonical port means the remote node needs no override. Registry row updated from "internal only".Verified
supabase-db→[pmoves_data, pmoves_api, pmoves_external]Activation
Needs a
supabase-dbrecreate withSUPABASE_DB_BIND=<tailnet-ip>andSUPABASE_DB_PORT=5432. That same recreate activates #2702's pg_hba mount, which has been merged but never applied — the container has been up 8 days with zero tailnet rules. Acceptance test is the asymmetry:juicefs_metaconnects from the 5090,supabase_adminis rejected.🤖 Generated with Claude Code
https://claude.ai/code/session_01FkwiW3VY1xWmahTAtVioxz