Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,7 @@ The graph has sixteen lanes: SCM compatibility · namespace · P-derive · obser
- [ ] **Bring the two new machines up as model-serving members, through the same path as everything else** — Both new machines join through the same standing-up sequence every other machine uses, reach a stated model and runtime, and serve requests on an endpoint. Which model is actually being served, and whether the endpoint is healthy, are read back independently, and each machine stays separately observable. They serve models — they are not job runners and not build caches. Why: Standing these up by hand would create two more machines nobody can answer questions about, which is the state this whole lane exists to end. Going through the convergence path also proves that path on a machine class it has never seen. [authority](dag/gunbc/spark/dgx_procurement.dag) — requires all: fleet-main-revision-authority, fleet-atomic-convergence-verdict, compute-exact-work-contract, fleet-spark-host-enrollment
- [ ] **A serving route that has stopped serving stops being routable, and the seats held against it are freed** — Whether a route can take new work is one reading, bound to the endpoint process that answered it, and every consumer projects from that reading rather than deciding separately. The reading covers three ways a route stops serving: the front door does not answer inside a bounded deadline; the endpoint process was replaced between observations; and the engine core is no longer advancing while requests are admitted. Seats and grants held against a launch that is gone become unusable rather than waiting out their term. Why: Group A was dark for 43 hours in September 2026 and nothing in the repository could say so: the health probes had no deadline, so a wedged engine held the roster gate, dispatch selection and placement open with no refusal, and one harness POST was held open for over thirty minutes against a backend that would never answer. A stalled engine still answers its front door, so the route stayed offerable throughout. [authority](dag/gunbc/serving/serving_availability.dag)
- [ ] **Serve DeepSeek V4.1 on four Sparks by building the runtime it needs, or say honestly that none exists** — One exact runtime candidate whose identity is fixed before anything is built: the pinned public vLLM source and build recipe, the kernel wheels re-admitted against their acquisition, the ordered engine patches by digest, the checkpoint and tokenizer manifests, which Engram tensors are read from storage and which stay resident, the materialization route, the topology and the device. Identity is not permission: a keyed candidate still owes an exact, time-bounded operator authorization before any host runs it. Why: The published V4.1 runtime cannot serve this checkpoint on four GB10s: gunbc.spark.serving_deployment_selection candidate_component_budget, reading gunbc.spark.pair_serving_observed group_rank_free_memory_observed, excludes resident Engram at TP4 by a measured shortfall on the Engram term after the artifact's other components — not a near miss against device total — and every Engram placement that runtime offers draws on that same pool. Without a candidate identity, a build receipt, a probe and a launch can each be attributed to bytes that produced none of them, which is how a deployment comes to be described by evidence it never generated. [authority](dag/gunbc/spark/v41_runtime_candidate.dag) — requires all: fleet-spark-inference-serving, serving-liveness-route-withdrawal
- [ ] **Cut D D0: suspend Group A under an exact operator consent, read the fleet inside it, settle to one typed terminal, and never free a host early** — gunbc.spark.pair_serving_d0 over gunbc.spark.pair_serving_authority_log on the fabric DB: one durable genesis per group partition, the host-placement partition as the linearization point, prepare -> append-claim -> authority -> finalize as one saga bound by a typed AuthorityWriteIntent, consumed/cancelled exclusivity, claim-bound quiescence evidence, total crash recovery by lifecycle identity, the operator consent slot as a fabric-DB head (gunbc.durable_cas_fabric_storage, generation-bearing objects), and a wet door that refuses while the store's write walls are missing. STATE (2026-09-20, wound down under operator direction to re-prioritize v1 performance and v2 migration): the whole stack is one branch, plan/dsv41-cut-d-2b, re-rooted onto the rewritten main and source-approved through twenty-one review rounds; the operator ruled it may land under the widened fabric-DB principal drop. Its stacked follow-up branch, plan/dsv41-cut-d-2b-admission-cleanup, deletes the twenty transition-admission rows the stack consumed and carries the detached-process wet-lane admission row. Why: Without D0 the V4.1 cut over Group A has no transaction: no consent that is spent exactly once across executors, no state in which the group is neither serving nor being mutated by two writers, and no receipt of the fleet as it was read before the authority moved -- the prior state, in which membership and placement were roster words and a crashed lane left no recoverable trace. [authority](docs/plans/dsv41-cut-d-redesign.md)
- [ ] **One answer for the whole fleet, or exactly which machines are not there yet** — Each machine produces a receipt — what was wanted, what was there, what was planned, what was applied, and what was read back afterwards. Those fold into a single answer for the fleet, joined with the lifecycle cells enrollment constructs: every required host-and-phase cell is observed, refused, unreachable, drifted, or converged, and a required cell nobody has observed prevents convergence rather than defaulting green. Anything short of all of them is not converged, with the machines that refused, could not be reached, or have drifted each named separately. Why: One machine can hold the new value while the others hold the old one and nothing anywhere says so, which is exactly why the honest answer to whether something is applied is currently to go and check by hand. And the fleet phase matrix currently has exactly one real observation producer cell — one host, one phase, hard-coded — so every other lifecycle claim in the fleet is unobserved by construction; enrollment makes obligations countable, not satisfied. [authority](docs/plans/fleet-acceptance-criteria.md)
- [ ] **Deploy every copy of the dashboard through the one deployment mechanism** — Deploying a dashboard uses the same mechanism as everything else we deploy: work out what should be there, compare it to what is, and change only the difference. Which machine it goes to becomes a parameter, not a second copy of the logic. Why: There are currently two deployment engines for the same three files. The second one does not know what it owns, and reaches the machine outside the checked path — so it can tear down something that was never ours. [authority](docs/plans/gunbc-served-dashboard-design.md)
- [ ] **Serve the roadmap page from emitted native code, never the tree-walking interpreter** — The interpreted accept loop is a declared scaffold whose dissolution trigger has now fired in production: a two-hour page wedge, one core pinned in memcpy, thirteen connections queued behind a single thread. The fix is ordered. Wire roadmap_serve_handle through the existing emit-on-demand path FIRST — interpreted concat copies its accumulator on every call, quadratic; emitted concat moves, linear: a complexity-class change, not tuning. Then memoize page bodies by content hash of their inputs so a request is lookup-and-write. Concurrency only after bodies are immutable artifacts. A request deadline returns a typed refusal instead of eating the process. Why: The 2026-08-08 srv1 outage: gunbc serve wedged for two hours rendering the daily page, every route behind it dead, and it re-wedges on the next load of the same shape. The scaffold row already names emit-on-demand as its dissolution and the module carries zero references to it — declared and left standing. Separately, the belt reconcile has been dead as long as GcpProjectId has failed to resolve, burning about two minutes of CPU on every timer tick before exiting. [authority](docs/plans/gunbc-served-dashboard-design.md)
Expand Down
24 changes: 24 additions & 0 deletions dag/extdeps/docker/cli.dag
Original file line number Diff line number Diff line change
Expand Up @@ -375,6 +375,30 @@ fn docker_image_inspect_command(image: NonEmptyStr) -> ArgvCommand {
// naming "No such container" when the reference names nothing. No --format, for the same
// portable-word reason as the image form. The argv and the sudoers Cmnd_Spec that authorizes it are
// the same words, so a grantee rendering both reads this one function.
// THE RUNNING CONTAINERS SELECTED BY ONE `docker ps` FILTER, one container id per line (`docker ps
// --filter <key>=<value> --quiet`): empty output is "no such running container". `docker ps` lists
// RUNNING containers only, which is the question a quiescence reading asks; a stopped container is
// not a residual realization of an effect. Two filters this repository asks by: `name=<name>`
// (docker's name filter is an unanchored regular expression, so it matches any running container
// whose name contains the text -- wider than exact, and the wider reading is the safe one for a
// residue question) and `ancestor=<image>` (containers created from that image reference, for an
// effect that starts an unnamed container). `--quiet` rather than a `--format` template because
// the argv crosses an ssh leg whose portable word alphabet has no braces.
type DockerPsFilter
= DockerPsByName { name: NonEmptyStr }
| DockerPsByAncestor { image: NonEmptyStr }

fn docker_ps_filter_wire(f: DockerPsFilter) -> String {
match f {
DockerPsByName { name: n } => join(["name=", n as String], "")
DockerPsByAncestor { image: i } => join(["ancestor=", i as String], "")
}
}

fn docker_ps_running_command(filter: DockerPsFilter) -> ArgvCommand {
argv_command(program: docker_binary_path, arguments: ["ps", "--filter", docker_ps_filter_wire(f: filter), "--quiet"])
}

fn docker_container_inspect_command(name: NonEmptyStr) -> ArgvCommand {
argv_command(program: docker_binary_path, arguments: ["container", "inspect", name as String])
}
Expand Down
1 change: 1 addition & 0 deletions dag/extdeps/exec/command.dag
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,7 @@ fn argv_command(program: NonEmptyStr, arguments: List<String>) -> ArgvCommand
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_exec_command"),
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_image_pull_command"),
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_container_inspect_command"),
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_ps_running_command"),
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_build_command"),
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_build_spec_command"),
decl_ref(module_path: "extdeps.docker.cli", decl_name: "docker_image_digest_command"),
Expand Down
26 changes: 26 additions & 0 deletions dag/extdeps/http/client.dag
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,16 @@ fn http_client_get_argv(url: NonEmptyStr, max_time: Second) -> List<String> {
// 5 rules that a bounded "forever" is not an "unknown" error. Measured on this deployment, a wedged
// backend held one harness POST open for over thirty minutes while the router answered an unrelated
// probe in 0.15ms -- so the turn was neither progressing nor failing, and no arm existed to say so.
// StatusWithin -- A STATUS IS AN ANSWER; ONLY NO STATUS IS SILENCE. Every GET carries -f, which
// turns a 4xx or 5xx into a failed operation -- right for a caller that wants a body, wrong for a
// caller asking whether anything is THERE: a front door returning 401 or 500 answered. That
// operation asks only for the status line, within the caller's deadline, and fails only when no
// HTTP response was received at all -- and WHY is carried as curl's exit code, because "no
// response" is not one fact: a failed connect (curl 7) and a deadline (curl 28) are different
// readings for a caller asking whether anything is THERE, and neither proves absence on its own.
// The consumer maps the code; the transport reports it. The codes are curl's own facts and live
// with the tool authority: extdeps.tools.curl curl_exit_outcome.
//
// PostStdinWithin, A BOUNDED POST WHOSE TRANSPORT OUTCOME IS KEPT, NOT FOLDED INTO SUCCESS. The request body goes
// on stdin (curl --data-binary @-), so no body byte passes through argv or a shared file. Without
// -f, an HTTP error status is an ANSWER (exit 0), and the status is appended as the final line by
Expand Down Expand Up @@ -142,6 +152,22 @@ service http.Client {
}
}

operation StatusWithin {
input { url: NonEmptyStr, max_seconds: NonEmptyStr }
output { status: String from "stdout", success: Bool from "exit_success", exit_code: Int from "exit_code" }
readonly
transport shell {
argv: ["curl", "-sS", "-o", "/dev/null", "-w", "%\{http_code}", "--max-time", "{max_seconds}", "{url}"]
}
exit {
0 => Unit
nonzero => String "http client status GET within a caller deadline received no response"
}
mock_response {
0 => { status: "", success: false, exit_code: 28 } "hermetic: no live HTTP endpoint; a bounded presence reading refuses rather than fabricate a status"
}
}

operation PostStdinWithin {
input { url: NonEmptyStr, request_body: String, connect_seconds: NonEmptyStr, max_seconds: NonEmptyStr }
output {
Expand Down
103 changes: 103 additions & 0 deletions dag/extdeps/linux/proc_net_tcp.dag
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
module extdeps.linux.proc_net_tcp

import std.types { NonEmptyStr, List, Bool, Int, Port }
import std.string_type { String }
import v2.std.optional { Present, Absent }
import std.algebra { trim }
import extdeps.numeric.base16 { base16_word_value }
import extdeps.external_authority { ExternalAuthority }
import extdeps.uri { Uri, Https }

// THE SHAPE OF /proc/net/tcp AND /proc/net/tcp6 (proc(5), "/proc/net/tcp"): one header line, then
// one row per socket, whitespace-separated. The fields this module reads are the second
// (local_address: hex IP, a colon, a 4-hex-digit port) and the fourth (st: the socket state as two
// hex digits, where 0A is TCP_LISTEN -- include/net/tcp_states.h). Everything else on the row is
// carried past. The port is the same 16-bit hex field in both files; only the address width
// differs (8 hex digits for v4, 32 for v6), and this parser reads the port from the right of the
// colon so the width is not its business.
//
// WHAT THIS ESTABLISHES: which local ports have a socket in LISTEN state on the host whose procfs
// was read, whatever bound them. It does not say which process owns the socket (that is the inode
// joined through /proc/<pid>/fd, a separate read) and it says nothing about a listener in another
// network namespace -- a container with its own netns has its own /proc/net/tcp. A consumer that
// needs "nothing listens on this port on this host" reads both files and, when containers may
// hold the port, the host's namespace is the one the enrolled address reaches.
data extdeps_external_authority_anchor: ExternalAuthority = ExternalAuthority {
uri: Uri {
scheme: Https
locator: "man7.org/linux/man-pages/man5/proc.5.html#/proc/net/tcp"
}
}

data tcp_listen_state_hex: String = "0A"

// THE LOCAL PORT IS std.types Port WHERE ONE IS BOUND. The field is a 4-hex-digit 16-bit word;
// 0000 is a real value the kernel prints for a socket with no bound port, and Port (1..65535) has
// no constructor for it, so the row says which it read rather than carrying 0 as a port.
type TcpLocalPort
= TcpPortBound { port: Port }
| TcpPortUnbound

type TcpSocketRow {
local_address: String
local_port: TcpLocalPort
state: String
}

type ProcNetTcpTable
= ProcNetTcpParsed { rows: List<TcpSocketRow> }
| ProcNetTcpUnparseable { line: String }

fn tcp_row_fields(line: String) -> List<String> {
filter(split(s: trim(s: line), delimiter: " "), f => f != "")
}

fn tcp_socket_row(line: String) -> TcpSocketRow? {
let fields = tcp_row_fields(line: line)
match get(xs: fields, index: 1) {
Absent => none
Present { value: local } =>
match get(xs: fields, index: 3) {
Absent => none
Present { value: st } => {
let parts = split(s: local, delimiter: ":")
if length(parts) != 2 { none } else {
match get(xs: parts, index: 1) {
Absent => none
Present { value: port_hex } =>
match base16_word_value(word: port_hex, max_digits: 4) {
Absent => none
Present { value: port } => Present { value: TcpSocketRow { local_address: local, local_port: if port == 0 { TcpPortUnbound } else { TcpPortBound { port: port } }, state: st } }
}
}
}
}
}
}
}

// The table: the header line is skipped by its first field ("sl"), blank lines are skipped, and
// a row that does not carry the two fields this module reads refuses the whole table naming the
// line -- a socket table with an unreadable row is not a table with fewer sockets.
fn proc_net_tcp_table(text: String) -> ProcNetTcpTable {
let lines = filter(split(s: text, delimiter: "\n"), l => trim(s: l) != "")
fold(lines, init: ProcNetTcpParsed { rows: [] as List<TcpSocketRow> }, f: (acc, line) =>
match acc {
ProcNetTcpUnparseable { line: _ } => acc
ProcNetTcpParsed { rows: rs } =>
match get(xs: tcp_row_fields(line: line), index: 0) {
Absent => acc
Present { value: first } =>
if first == "sl" { acc } else {
match tcp_socket_row(line: line) {
Absent => ProcNetTcpUnparseable { line: line }
Present { value: r } => ProcNetTcpParsed { rows: concat(rs, [r]) }
}
}
}
})
}

fn tcp_rows_listening_on(rows: List<TcpSocketRow>, port: Port) -> List<TcpSocketRow> {
filter(rows, r => r.state == tcp_listen_state_hex && (match r.local_port { TcpPortBound { port: p } => p == port TcpPortUnbound => false }))
}
40 changes: 40 additions & 0 deletions dag/extdeps/linux/procfs.dag
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,8 @@ type ProcfsPath
| ProcSelfStatus
| ProcSelfCgroup
| ProcNetUnix
| ProcNetTcp
| ProcNetTcp6
| ProcSwaps
| ProcSysKernelHostname
| ProcNetPnp
Expand All @@ -36,6 +38,8 @@ fn procfs_path_literal(p: ProcfsPath) -> NonEmptyStr {
ProcSelfStatus => "/proc/self/status"
ProcSelfCgroup => "/proc/self/cgroup"
ProcNetUnix => "/proc/net/unix"
ProcNetTcp => "/proc/net/tcp"
ProcNetTcp6 => "/proc/net/tcp6"
ProcSwaps => "/proc/swaps"
ProcSysKernelHostname => "/proc/sys/kernel/hostname"
ProcNetPnp => "/proc/net/pnp"
Expand Down Expand Up @@ -121,6 +125,40 @@ service linux.Procfs {
}
}

operation ReadNetTcp {
input {}
output {
value: String from "stdout"
success: Bool from "exit_success"
}
readonly
transport shell { argv: ["cat", "/proc/net/tcp"] }
exit {
0 => Unit
nonzero => String "/proc/net/tcp read failed"
}
mock_response {
0 => { value: " sl local_address rem_address st tx_queue rx_queue tr tm->when retrnsmt uid timeout inode\n", success: true } "hermetic linux.Procfs.ReadNetTcp: a socket table with no rows"
}
}

operation ReadNetTcp6 {
input {}
output {
value: String from "stdout"
success: Bool from "exit_success"
}
readonly
transport shell { argv: ["cat", "/proc/net/tcp6"] }
exit {
0 => Unit
nonzero => String "/proc/net/tcp6 read failed"
}
mock_response {
0 => { value: " sl local_address remote_address st tx_queue rx_queue tr tm->when retrnsmt uid timeout inode\n", success: true } "hermetic linux.Procfs.ReadNetTcp6: a socket table with no rows"
}
}

operation ReadPidCgroup {
input { pid: NonEmptyStr }
output {
Expand Down Expand Up @@ -185,6 +223,8 @@ fn procfs_path_provision(p: ProcfsPath) -> ProcfsPathProvision {
ProcSwaps => ProvisionNotYetQualified
ProcSysKernelHostname => ProvisionNotYetQualified
ProcNetPnp => ProvidedWhenConfigured { option: "CONFIG_IP_PNP" as NonEmptyStr }
ProcNetTcp => ProvisionNotYetQualified
ProcNetTcp6 => ProvisionNotYetQualified
}
}

Expand Down
Loading