Repository navigation
Model the GitHub observation record and the bounded scan producer - #11564
Conversation
…y URL First piece of the ingest layer. An observation is addressed by WHAT WAS ASKED -- a closed request set whose identity is derived from the fields that change the answer -- so a later model change is a pure fold over stored records rather than another traversal of GitHub. Every figure this lane has produced so far came from a fetch into a temp directory that no longer exists. Parts fold through the hash rather than a separator: a delimiter-joined key collides the moment a field contains the delimiter, and a code-search query is the field most likely to contain anything. An empty part folds as a named placeholder so ["", "a"] and ["a"] stay distinct. Completeness is stored beside the payload because a reader cannot recover it from the bytes: a page of 30 hits looks identical whether it is the whole answer or what arrived before the search timed out. That conflation is not hypothetical -- this lane made it twice in the prototype. The rate-limit receipt is DECLARED UNOBSERVABLE rather than modelled. GitHub reports remaining budget in response headers; this repository's REST transport models request headers only and no operation reads an output from a response header, so X-RateLimit-Remaining cannot be seen. BudgetUnobservable carries that at every record and names the consequence: a scan paces from the DECLARED limits in extdeps.api_rate_limit, which is a published policy rather than a measurement, so it assumes nothing else is spending the same credential. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ation is The obligation this closes has been declared in gunbc.ci_workflow_source_read since before this lane started: "a scan producer that pages perform_code_search_page under a declared bound, reads each survivor at an exact ref, and canonicalizes them against a run receipt ... the obligation is the page walk and its receipt, which is where a bounded scan becomes a declared population rather than a sample nobody sized." Every decision it needs was already modelled and had no caller. THE WALK IS A FOLD OVER A BOUNDED PAGE LIST, NOT A LOOP UNTIL EMPTY. "Keep fetching until a page comes back short" has no bound -- a query whose result set grows under it, or an endpoint answering full pages, walks until a rate limit or an operator stops it. deepest_retrievable_page derives the list from the endpoint's own ceiling, so the most a scan can spend is known before it spends anything, and there is no caller-supplied limit that could be set past what the endpoint will serve. Early exhaustion still stops it, which needs care because a fold cannot break: the walk state carries a stop reason and every later step becomes a pass-through that performs NO effect. On an instrument rate-limited to ten requests a minute, fetching pages past a known answer is the difference between a scan that finishes and one that does not. THE 1000-RESULT CEILING IS NOW A DECLARATION. extdeps.github.code_search has always said in prose that the endpoint caps a query at 1000 results however many exist, and a number living only in prose cannot be compared or refused against. total_count is GitHub's claim about what EXISTS; retrievable_result_ceiling is what it will surrender. When the first exceeds the second the query is SATURATED and no walk can enumerate it -- the repair is a narrower query, not more pages. scan_is_a_declared_population requires all three: the walk exhausted rather than hit its bound, upstream never declared it gave up, and the query was enumerable. Each guards a distinct way this lane's own prototype reported a sample as a census. Thirteen witnesses drive the classification without a network, which is possible because ci_workflow_source_read split the decision from the fetch -- its annotation records why: the mock publishes one result, so the truncated arm is unreachable through the effectful path and an earlier revision had to carry a hand-copied twin that could not go red when the real decision was wrong. Each wall carries its control: a full page must NOT exhaust the walk, a query at the ceiling must stay enumerable, and zero pages read leaves the size unobserved rather than reporting an enumerable partition of size zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
briansrls
left a comment
There was a problem hiding this comment.
REQUEST CHANGES — source-reviewed exact head 5e5f060b1d9496b233d009c0fe90c6d89d997143.
The bounded page-walk decisions are useful, but the artifact does not yet have the provenance or persistence its title and body claim.
-
body_digestis not a digest of the body.scan_observationsets it tocontent_hash_atom(query). Every page of one query therefore records the same digest regardless of the response or returned hits. The observation cannot verify, recover, or identify what GitHub actually answered. -
The receipt is not bound to its population.
ScanReceiptstores onlyhits_read; the hits are not inside it and no hits identity is derived.scan_hits_if_declared_population(receipt, hits)accepts an arbitrary caller-supplied hit list alongside a green receipt. A receipt from population A can authorize population B. Carry the rows in the sole-constructed receipt or bind them through a derived content identity and expose them only from that receipt. -
Nothing durable is materialized. There are no raw response bytes, stored rows, artifact/store reference, or publish/readback route. After the process exits, later model changes cannot fold the original response without calling GitHub again. Either narrow this PR honestly to an in-memory scan-classification primitive, or land the immutable payload/materialization binding that makes “durable” and “zero-refetch re-derivation” true.
-
The derived page bound has an invalid input domain and under-fetches valid page sizes.
1000 / per_pagedivides by zero for zero, accepts negative or >100 values, and floors for non-divisors (per_page=30yields 33 pages/990 retrievable rows). Use the endpoint's typed 1..100 page-size contract and ceiling division, or refuse invalid sizes before constructing the walk.
Additional residual: an exactly-1000 result set ending on a full page is always classified StoppedAtBound, so the PartitionFullyEnumerable { total_count: 1000 } control can never yield a declared population. That is fail-closed, but it means the current “at the ceiling is enumerable” claim does not hold end to end.
The next cut named in the PR—exact-ref source canonicalization—should consume the same receipt-bound payload, not a second independently supplied hit list.
Side-chat REWORK. Part 1 (external blockers as prose) was withdrawn after I traced that startable authorizes closing-contract authoring rather than implementation dispatch. Parts 2 and 3 stood, and both were my errors. THE CENSUS ROWS ARE NOT AN AUTHORITY. The page said the 57 transcribed rows were dissolved by the scan producer. That is materially wrong in two ways. The economic readings attached to those rows were shown not to have the meanings assigned to them -- wall duration is not summed runner occupancy, an admission delay is not a runner queue delay, a provider declaration can outlive provider execution, and adoption is not spend -- so the derived runner-minutes and ARM-tier totals do not follow, and "five carry CostOpportunity" must not be quoted as a finding. And the scan producer is workflow-level, so it is necessary and NOT sufficient: the facts those rows need are job-level. The job-level arc is now a roadmap node instead of a sentence. ACQUISITION IS MODELED, NOT ACHIEVED. The page called gunbc#11552 an end-to-end App manifest flow. Its own route tells the operator to paste the manifest into the create form's manifest field; GitHub exposes no such field and the manifest protocol needs a form POST. I watched that step fail live in this session and wrote "end to end" anyway. Registration and INSTALLATION are also two facts, and gunbc#11677 consumes both rather than creating either. That is now a node with the remaining work named. EXTERNAL BLOCKERS ARE GONE. gunbc#11552, #11564, #11656, #11669, #11671, #11677 and #11679 are all merged. The nodes and the page said otherwise. The two real prerequisites are now EDGES rather than prose, which is the repair the reviewer asked for: the installation token depends on the App existing, and the job-level reprojection depends on the token. Projection reconciles 146 -> 148: two nodes and their closing-contract carriers, minus the two carriers the new edges remove from the startable set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two modules that turn a discovery scan from a directory of fetched JSON into a durable, self-describing population. Both close obligations the corpus already declared.
The observation record
An observation is addressed by WHAT WAS ASKED -- a closed request set whose identity is derived from the fields that change the answer -- so a later model change is a pure fold over stored records rather than another traversal of GitHub. Every figure this lane has produced so far came from a fetch into a temp directory that no longer exists, which is why none of them is reproducible.
Parts fold through the hash rather than a separator: a delimiter-joined key collides the moment a field contains the delimiter, and a code-search query is the field most likely to contain anything.
Completeness is stored beside the payload because a reader cannot recover it from the bytes -- a page of 30 hits looks identical whether it is the whole answer or what arrived before the search timed out.
The rate-limit receipt is DECLARED UNOBSERVABLE rather than modelled. GitHub reports remaining budget in response headers; this repository's REST transport models request headers only and no operation reads an output from a response header, so X-RateLimit-Remaining cannot be seen. BudgetUnobservable carries that at every record and names the consequence: a scan paces from the DECLARED limits in extdeps.api_rate_limit, which is published policy rather than measurement, so it assumes nothing else is spending the same credential.
The scan producer
gunbc.ci_workflow_source_read has carried this obligation since before this lane started: "a scan producer that pages perform_code_search_page under a declared bound ... the obligation is the page walk and its receipt, which is where a bounded scan becomes a declared population rather than a sample nobody sized." Every decision it needs was already modelled and had no caller.
The walk is a fold over a bounded page list, not a loop until empty. "Keep fetching until a page comes back short" has no bound. deepest_retrievable_page derives the list from the endpoint's own ceiling, so the most a scan can spend is known before it spends anything, and no caller-supplied limit can be set past what the endpoint will serve.
Early exhaustion still stops it, which needs care because a fold cannot break: the walk state carries a stop reason and every later step becomes a pass-through performing NO effect. On an instrument rate-limited to ten requests a minute, fetching past a known answer is the difference between a scan that finishes and one that does not.
The 1000-result ceiling is now a declaration. extdeps.github.code_search has always said in prose that the endpoint caps a query however many results exist, and a number living only in prose cannot be compared or refused against. When total_count exceeds it the query is SATURATED and no walk can enumerate it -- the repair is a narrower query, not more pages.
scan_is_a_declared_population requires all three: the walk exhausted rather than hit its bound, upstream never declared it gave up, and the query was enumerable. Each guards a distinct way this lane's own prototype reported a sample as a census.
Evidence
Thirteen witnesses drive the classification without a network, which is possible because ci_workflow_source_read split the decision from the fetch. Each wall carries its control: a full page must NOT exhaust the walk, a query at the ceiling must stay enumerable, and zero pages read leaves the size unobserved rather than reporting an enumerable partition of size zero.
Seven core claims verified PASS by execution under claim_batch. required-ci witnesses lane on srv1: parse, declarations and the touched-subject compile green.
Not in this change
The producer canonicalizes nothing yet -- wiring its hits through perform_workflow_source_read at an exact ref is the next cut, and it is what turns an index hint into an observation. No job-level ingest: that needs the API budget the GitHub App work (gunbc#11552) is building.
This branch supersedes gunbc#11541, which was mis-cut from a session branch and carried an unrelated PR's twelve commits.
🤖 Generated with Claude Code