Skip to content

docs: mark the metadata tree generated, and say not to read it - #477

Merged
twcclegg merged 5 commits into
mainfrom
claude/linguist-generated-metadata
Sep 24, 2026
Merged

twcclegg merged 5 commits into
mainfrom
claude/linguist-generated-metadata

Conversation

@twcclegg

@twcclegg twcclegg commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

Changes

  • New .gitattributes marking the data under resources/** as linguist-generated=true — upstream's XML metadata, the geocoding/carrier/timezone tables, and the generated locale/country_names.txt. 459 files. The two .proto files and the two upstream READMEs are listed by name and exempted; nothing outside resources/ is touched.
  • AGENTS.md: the don't-hand-edit rule becomes don't hand-edit it, and don't read it. Nobody, human or agent, needs to read the data — the one exception is working on the parser and needing the schema, and the schema is the two .proto files (22 KB together), so those are fine to open. Sizes and an extraction one-liner are in the rule for the case where you genuinely need one rule out of the XML.
  • diagnosing-number-behaviour skill: section 3a no longer tells you to quote the metadata pattern at the reporter. It says what upstream actually asks for — evidence from the numbering authority — and says not to go digging in resources/ to explain an answer.
  • syncing-upstream-metadata skill: review a metadata PR by its file list (git diff --name-only origin/main..., --stat for shape), not by pulling the data diff into context; a collapsed file is still a changed file, which is why the check is the file list.

Why

Every fortnight the sync opens a PR whose diff is thousands of lines of upstream data with one CHANGELOG.md entry buried in it, and the thing a reviewer actually needs to check — did anything unexpected ride along? — is the hardest part to see. linguist-generated collapses the data, leaves it out of the language statistics and code search, and leaves the part worth reading in plain view.

That does nothing for the second reader. An agent asked why some number validates the way it does will still read a megabyte of XML to find one 500-line <territory> block, because linguist-generated only changes what GitHub renders. So the agent-facing half has to be said in words, which is what the AGENTS.md and skill changes do. Marking the tree without saying that would have looked like a fix and not been one.

Why the .proto files and the READMEs are exempt

They are the schema, not the data, and a .proto change is the one thing in this tree that is never routine: lib/github-actions-metadata-update.sh exits at EXIT_NEEDS_ATTENTION when the upstream diff contains one, because it may need porting by hand. That makes it precisely the diff a reviewer needs expanded rather than collapsed. They are also the one part of resources/ the hard rule allows reading, so keeping them in code search is worth something on its own.

resources/carrier/README and resources/timezones/README.md are exempt for the same reason on the prose side: 2.7 KB between them, describing what the data is and where it came from, and an upstream change to either is a readable signal rather than noise.

All four are listed by name rather than by glob, so anything new upstream is marked as data by default and has to be looked at and exempted deliberately.

Why the derived C# tables are not marked

An earlier revision of this PR also marked CountryCodeToRegionCodeMap.cs and ShortNumbersRegionCodeSet.cs, on the grounds that upstream generates both with BuildMetadataProtoFromXml and this port mirrors them. That was wrong for the first one: nothing regenerates it here, its own header still says "todo make this file automatically generated", lib/github-actions-metadata-update.sh deliberately treats a change to it as hand-written content, and AGENTS.md says to edit it by hand when you need to. Marking a hand-maintained file generated collapses its diff and drops it out of code search for no gain.

ShortNumbersRegionCodeSet.cs comes out alongside it rather than leaving half a pair marked while that question is open. Both can be revisited once CountryCodeToRegionCodeMap.cs is settled one way or the other.

Why 3a stopped recommending the metadata

The skill used to say that quoting the exact pattern which rejected a number "turns 'it's a metadata issue' into an answer the reporter can act on upstream". It does not. Upstream reports are not settled by pointing at a metadata file — they state a number or range that does not match expectations and are then asked for evidence from the national regulator or an equivalent authority. The pattern is a restatement of the disputed behaviour, not evidence for the claim.

It is not the authority on this port's behaviour either: resources/ is a build input that PhoneNumbers.MetadataBuilder compiles into the embedded binary metadata the library actually reads (see docs/api-differences-from-java.md). A scratch test answers "what does it do" in seconds, and the comparison against Google's demo in step 2 already settles metadata-vs-port-bug. Reading the XML adds nothing to either.

Effects, and what is not affected

linguist-generated is read by GitHub, not by git. No text, eol, diff or merge attribute is set, so how git stores, diffs and merges these files is unchanged — including the git diff that lib/update-changelog.sh uses to decide whether a release is metadata-only.

  • Marked files drop out of GitHub code search for this repository: the XML and the prefix tables. A local git grep, the diagnosing-number-behaviour workflow and the triage automation all read the checkout, so none of them are affected. The .proto files and both .cs tables stay searchable.
  • The language bar is unaffected. The only programming language in resources/ is Protocol Buffer, and it is now exempt; the .txt and .xml are data languages and were never counted.

linguist-vendored would have the same effect and is arguably the better word for a verbatim upstream copy, but it does not fit locale/country_names.txt, which is generated here by lib/DumpLocale.java rather than vendored — so one attribute across the set reads truer than splitting it.

Verification

git check-attr over every tracked file: 459 marked. Both .proto files and both READMEs come back false; PhoneNumberMetadata.xml, ShortNumberMetadata.xml, geocoding/, carrier/, timezones/ and locale/country_names.txt come back true; CountryCodeToRegionCodeMap.cs, ShortNumbersRegionCodeSet.cs, PhoneNumberUtil.cs, CHANGELOG.md and the csproj files all come back unspecified. A hypothetical resources/nested/new.proto comes back true, confirming the exemptions are per-file and not globs. git check-attr -a sets nothing but linguist-generated on any of them.

git status is clean after adding the file — no renormalization, since no text attribute is involved. The one command still quoted in AGENTS.md was run against the current tree: the GB <territory> range returns 518 lines of PhoneNumberMetadata.xml's 32,000.

@codecov

codecov Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 87.69%. Comparing base (d6a67ec) to head (e534fe7).

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #477   +/-   ##
=======================================
  Coverage   87.69%   87.69%           
=======================================
  Files          43       43           
  Lines        3893     3893           
  Branches      993      993           
=======================================
  Hits         3414     3414           
  Misses        277      277           
  Partials      202      202           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@twcclegg twcclegg changed the title docs: mark the metadata tree and its derived tables linguist-generated docs: mark the metadata tree generated, and say not to read it Sep 20, 2026
Every fortnight the metadata sync opens a PR whose diff is thousands of lines of
upstream data with one CHANGELOG.md entry somewhere in it, and the review that
matters - did anything unexpected ride along? - is the hardest thing to see in
it. Marking resources/** generated collapses the data on GitHub, leaves it out of
the repository's language statistics and code search, and leaves the part worth
reading in plain view.

The same goes for the two C# tables derived from that metadata:
CountryCodeToRegionCodeMap.cs and ShortNumbersRegionCodeSet.cs, which upstream
generates with BuildMetadataProtoFromXml and this port mirrors rather than
regenerating. Neither is content a reviewer reads line by line, which is what the
marking is about; it says nothing about who may edit them, and
CountryCodeToRegionCodeMap.cs stays hand-maintained here as AGENTS.md describes.

This is read by GitHub's linguist, not by git: no text, eol, diff or merge
attribute is set, so nothing about how git stores, diffs or merges these files
changes - including the `git diff` that lib/update-changelog.sh uses to decide
whether a release is metadata-only.

AGENTS.md and the syncing-upstream-metadata skill say so where it matters, with
the caveat that a collapsed diff is not an unchanged file: the "expect only
resources/** and one CHANGELOG entry" check on a sync PR is a check of the file
list, not of what GitHub renders.
… hides it

The linguist marking only changes what GitHub renders; an agent reading
PhoneNumberMetadata.xml or a 3.8 MB geocoding table pays the full cost either
way. So the hard rule now covers reading as well as editing, with the sizes and
the extraction one-liner spelled out, the diagnosis skill gives the extraction
command for each file instead of naming the file to open, and the sync skill says
to review a metadata PR by its file list rather than pulling thousands of lines of
upstream data into context.
@twcclegg
twcclegg force-pushed the claude/linguist-generated-metadata branch from d10e0fc to 341b5b5 Compare September 24, 2026 00:20
…eport

The diagnosis skill told you to extract the pattern that rejected a number and
quote it, on the grounds that this "turns 'it's a metadata issue' into something
the reporter can act on upstream". It does not. Ten years of Google's list traffic
has never turned on someone pointing at a metadata file: a report states a number
or range that does not match expectations, and is then asked for evidence from the
national regulator or an equivalent authority. The pattern is not evidence for that
claim - it is a restatement of the behaviour being disputed.

It is not even the authority on this port's own behaviour. resources/ is a build
input that PhoneNumbers.MetadataBuilder compiles into the embedded binary metadata
the library actually reads, so a scratch test answers "what does it do" in seconds
and the demo comparison in step 2 already settles metadata-vs-port-bug. Reading the
XML adds nothing to either and costs a large part of a context window.

So 3a now says what upstream actually asks for, and says not to go digging in
resources/ to explain an answer. AGENTS.md keeps the extraction one-liner but scopes
it to the one case that needs it - working on the parser and needing to see the
schema - rather than pointing at a skill section that no longer recommends reading
the tree at all.
… marking

CountryCodeToRegionCodeMap.cs and ShortNumbersRegionCodeSet.cs come out of the
linguist list. The first is the reason: nothing regenerates it, AGENTS.md says to
edit it by hand, and marking a hand-maintained file generated collapses its diff
and drops it out of GitHub code search for no gain. ShortNumbersRegionCodeSet.cs
goes with it rather than leaving half a pair marked on a question still open - it
can come back once CountryCodeToRegionCodeMap.cs is settled one way or the other.

resources/phonemetadata.proto and resources/phonenumber.proto are exempted from
the resources/** marking. They are the schema, not the data, and a .proto change
is the one thing in this tree that is never routine: it trips the sync's gate and
stops the run outright because it may need porting by hand. That makes it exactly
the diff a reviewer needs expanded. They are also small - 22 KB for both - and are
the one part of resources/ the hard rule allows reading, so keeping them in code
search is worth something too.

Listed by name rather than by glob, so a new .proto upstream is marked as data by
default and has to be looked at and exempted deliberately.

461 files marked now, down from 465.
…rd rule

resources/carrier/README and resources/timezones/README.md are upstream prose
describing the data, not the data - the same category as the .proto schema, and
2.7 KB between them. An upstream change to either is a readable signal worth
seeing expanded, so they join the exemption list. 459 files marked now.

The resources/ hard rule had grown to twenty lines carrying six separate ideas.
Split: the tree itself in one bullet, the two derived C# tables in another, both
tightened.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant