docs: mark the metadata tree generated, and say not to read it - #477
Merged
Merged
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #477 +/- ##
=======================================
Coverage 87.69% 87.69%
=======================================
Files 43 43
Lines 3893 3893
Branches 993 993
=======================================
Hits 3414 3414
Misses 277 277
Partials 202 202 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Every fortnight the metadata sync opens a PR whose diff is thousands of lines of upstream data with one CHANGELOG.md entry somewhere in it, and the review that matters - did anything unexpected ride along? - is the hardest thing to see in it. Marking resources/** generated collapses the data on GitHub, leaves it out of the repository's language statistics and code search, and leaves the part worth reading in plain view. The same goes for the two C# tables derived from that metadata: CountryCodeToRegionCodeMap.cs and ShortNumbersRegionCodeSet.cs, which upstream generates with BuildMetadataProtoFromXml and this port mirrors rather than regenerating. Neither is content a reviewer reads line by line, which is what the marking is about; it says nothing about who may edit them, and CountryCodeToRegionCodeMap.cs stays hand-maintained here as AGENTS.md describes. This is read by GitHub's linguist, not by git: no text, eol, diff or merge attribute is set, so nothing about how git stores, diffs or merges these files changes - including the `git diff` that lib/update-changelog.sh uses to decide whether a release is metadata-only. AGENTS.md and the syncing-upstream-metadata skill say so where it matters, with the caveat that a collapsed diff is not an unchanged file: the "expect only resources/** and one CHANGELOG entry" check on a sync PR is a check of the file list, not of what GitHub renders.
… hides it The linguist marking only changes what GitHub renders; an agent reading PhoneNumberMetadata.xml or a 3.8 MB geocoding table pays the full cost either way. So the hard rule now covers reading as well as editing, with the sizes and the extraction one-liner spelled out, the diagnosis skill gives the extraction command for each file instead of naming the file to open, and the sync skill says to review a metadata PR by its file list rather than pulling thousands of lines of upstream data into context.
twcclegg
force-pushed
the
claude/linguist-generated-metadata
branch
from
September 24, 2026 00:20
d10e0fc to
341b5b5
Compare
…eport The diagnosis skill told you to extract the pattern that rejected a number and quote it, on the grounds that this "turns 'it's a metadata issue' into something the reporter can act on upstream". It does not. Ten years of Google's list traffic has never turned on someone pointing at a metadata file: a report states a number or range that does not match expectations, and is then asked for evidence from the national regulator or an equivalent authority. The pattern is not evidence for that claim - it is a restatement of the behaviour being disputed. It is not even the authority on this port's own behaviour. resources/ is a build input that PhoneNumbers.MetadataBuilder compiles into the embedded binary metadata the library actually reads, so a scratch test answers "what does it do" in seconds and the demo comparison in step 2 already settles metadata-vs-port-bug. Reading the XML adds nothing to either and costs a large part of a context window. So 3a now says what upstream actually asks for, and says not to go digging in resources/ to explain an answer. AGENTS.md keeps the extraction one-liner but scopes it to the one case that needs it - working on the parser and needing to see the schema - rather than pointing at a skill section that no longer recommends reading the tree at all.
… marking CountryCodeToRegionCodeMap.cs and ShortNumbersRegionCodeSet.cs come out of the linguist list. The first is the reason: nothing regenerates it, AGENTS.md says to edit it by hand, and marking a hand-maintained file generated collapses its diff and drops it out of GitHub code search for no gain. ShortNumbersRegionCodeSet.cs goes with it rather than leaving half a pair marked on a question still open - it can come back once CountryCodeToRegionCodeMap.cs is settled one way or the other. resources/phonemetadata.proto and resources/phonenumber.proto are exempted from the resources/** marking. They are the schema, not the data, and a .proto change is the one thing in this tree that is never routine: it trips the sync's gate and stops the run outright because it may need porting by hand. That makes it exactly the diff a reviewer needs expanded. They are also small - 22 KB for both - and are the one part of resources/ the hard rule allows reading, so keeping them in code search is worth something too. Listed by name rather than by glob, so a new .proto upstream is marked as data by default and has to be looked at and exempted deliberately. 461 files marked now, down from 465.
…rd rule resources/carrier/README and resources/timezones/README.md are upstream prose describing the data, not the data - the same category as the .proto schema, and 2.7 KB between them. An upstream change to either is a readable signal worth seeing expanded, so they join the exemption list. 459 files marked now. The resources/ hard rule had grown to twenty lines carrying six separate ideas. Split: the tree itself in one bullet, the two derived C# tables in another, both tightened.
This was referenced Sep 24, 2026
Open
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changes
.gitattributesmarking the data underresources/**aslinguist-generated=true— upstream's XML metadata, the geocoding/carrier/timezone tables, and the generatedlocale/country_names.txt. 459 files. The two.protofiles and the two upstream READMEs are listed by name and exempted; nothing outsideresources/is touched.AGENTS.md: the don't-hand-edit rule becomes don't hand-edit it, and don't read it. Nobody, human or agent, needs to read the data — the one exception is working on the parser and needing the schema, and the schema is the two.protofiles (22 KB together), so those are fine to open. Sizes and an extraction one-liner are in the rule for the case where you genuinely need one rule out of the XML.diagnosing-number-behaviourskill: section 3a no longer tells you to quote the metadata pattern at the reporter. It says what upstream actually asks for — evidence from the numbering authority — and says not to go digging inresources/to explain an answer.syncing-upstream-metadataskill: review a metadata PR by its file list (git diff --name-only origin/main...,--statfor shape), not by pulling the data diff into context; a collapsed file is still a changed file, which is why the check is the file list.Why
Every fortnight the sync opens a PR whose diff is thousands of lines of upstream data with one
CHANGELOG.mdentry buried in it, and the thing a reviewer actually needs to check — did anything unexpected ride along? — is the hardest part to see.linguist-generatedcollapses the data, leaves it out of the language statistics and code search, and leaves the part worth reading in plain view.That does nothing for the second reader. An agent asked why some number validates the way it does will still read a megabyte of XML to find one 500-line
<territory>block, becauselinguist-generatedonly changes what GitHub renders. So the agent-facing half has to be said in words, which is what theAGENTS.mdand skill changes do. Marking the tree without saying that would have looked like a fix and not been one.Why the
.protofiles and the READMEs are exemptThey are the schema, not the data, and a
.protochange is the one thing in this tree that is never routine:lib/github-actions-metadata-update.shexits atEXIT_NEEDS_ATTENTIONwhen the upstream diff contains one, because it may need porting by hand. That makes it precisely the diff a reviewer needs expanded rather than collapsed. They are also the one part ofresources/the hard rule allows reading, so keeping them in code search is worth something on its own.resources/carrier/READMEandresources/timezones/README.mdare exempt for the same reason on the prose side: 2.7 KB between them, describing what the data is and where it came from, and an upstream change to either is a readable signal rather than noise.All four are listed by name rather than by glob, so anything new upstream is marked as data by default and has to be looked at and exempted deliberately.
Why the derived C# tables are not marked
An earlier revision of this PR also marked
CountryCodeToRegionCodeMap.csandShortNumbersRegionCodeSet.cs, on the grounds that upstream generates both withBuildMetadataProtoFromXmland this port mirrors them. That was wrong for the first one: nothing regenerates it here, its own header still says "todo make this file automatically generated",lib/github-actions-metadata-update.shdeliberately treats a change to it as hand-written content, andAGENTS.mdsays to edit it by hand when you need to. Marking a hand-maintained file generated collapses its diff and drops it out of code search for no gain.ShortNumbersRegionCodeSet.cscomes out alongside it rather than leaving half a pair marked while that question is open. Both can be revisited onceCountryCodeToRegionCodeMap.csis settled one way or the other.Why 3a stopped recommending the metadata
The skill used to say that quoting the exact pattern which rejected a number "turns 'it's a metadata issue' into an answer the reporter can act on upstream". It does not. Upstream reports are not settled by pointing at a metadata file — they state a number or range that does not match expectations and are then asked for evidence from the national regulator or an equivalent authority. The pattern is a restatement of the disputed behaviour, not evidence for the claim.
It is not the authority on this port's behaviour either:
resources/is a build input thatPhoneNumbers.MetadataBuildercompiles into the embedded binary metadata the library actually reads (seedocs/api-differences-from-java.md). A scratch test answers "what does it do" in seconds, and the comparison against Google's demo in step 2 already settles metadata-vs-port-bug. Reading the XML adds nothing to either.Effects, and what is not affected
linguist-generatedis read by GitHub, not by git. Notext,eol,difformergeattribute is set, so how git stores, diffs and merges these files is unchanged — including thegit diffthatlib/update-changelog.shuses to decide whether a release is metadata-only.git grep, thediagnosing-number-behaviourworkflow and the triage automation all read the checkout, so none of them are affected. The.protofiles and both.cstables stay searchable.resources/is Protocol Buffer, and it is now exempt; the.txtand.xmlare data languages and were never counted.linguist-vendoredwould have the same effect and is arguably the better word for a verbatim upstream copy, but it does not fitlocale/country_names.txt, which is generated here bylib/DumpLocale.javarather than vendored — so one attribute across the set reads truer than splitting it.Verification
git check-attrover every tracked file: 459 marked. Both.protofiles and both READMEs come backfalse;PhoneNumberMetadata.xml,ShortNumberMetadata.xml,geocoding/,carrier/,timezones/andlocale/country_names.txtcome backtrue;CountryCodeToRegionCodeMap.cs,ShortNumbersRegionCodeSet.cs,PhoneNumberUtil.cs,CHANGELOG.mdand the csproj files all come backunspecified. A hypotheticalresources/nested/new.protocomes backtrue, confirming the exemptions are per-file and not globs.git check-attr -asets nothing butlinguist-generatedon any of them.git statusis clean after adding the file — no renormalization, since notextattribute is involved. The one command still quoted inAGENTS.mdwas run against the current tree: the GB<territory>range returns 518 lines ofPhoneNumberMetadata.xml's 32,000.