Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 4 additions & 6 deletions .github/workflows/foresight-evidence-spine.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ name: Foresight Evidence Spine
on:
push:
branches:
- feature/foresight-evidence-spine
- master
pull_request:
paths:
- 'foresight/**'
Expand All @@ -19,13 +19,11 @@ jobs:
name: foresight registry tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@08c6903cd8c0fde910a37f88322edcfb5dd907a8 # v5.0.0
- uses: actions/checkout@08c6903cd8c0fde910a37f88322edcfb5dd907a # v5.0.0
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: '3.12'
- name: Unit tests
run: python -m unittest tests/test_foresight_registry.py -v
- name: Registry smoke validation
run: |
python -m foresight.registry --registry data/foresight/resource-registry.jsonl init
python -m foresight.registry --registry data/foresight/resource-registry.jsonl validate
- name: Registry validation
run: python -m foresight.registry --registry data/foresight/resource-registry.jsonl validate
1 change: 1 addition & 0 deletions data/foresight/resource-registry.jsonl
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"resource_id":"otel-semantic-conventions","item_id":"otel-semantic-conventions-2026-09-27","lane_id":"agent-observability","title":"OpenTelemetry Semantic Conventions","category":"agent-observability","horizon":"H0","evidence_status":"confirmed","summary":"The OpenTelemetry semantic-conventions project publishes a shared schema for semantic attributes; the observed repository identifies specification version v1.61.0 and is licensed under Apache-2.0.","why_it_matters":"A vendor-neutral semantic layer can keep observability evidence portable across telemetry backends.","practical_opportunity":"Use the canonical conventions as an adapter input while retaining raw registry evidence independently.","termux_relevance":"schema-only consumption is lightweight; exporter/runtime portability remains an implementation question","confidence":0.98,"observed_at":"2026-09-27T18:30:00Z","captured_at":"2026-09-27T18:30:00Z","source":{"url":"https://github.com/open-telemetry/semantic-conventions","publisher":"OpenTelemetry","title":"OpenTelemetry Semantic Conventions"},"canonical_source":"https://github.com/open-telemetry/semantic-conventions","source_type":"primary_repository","source_version":"v1.61.0","version_or_commit":"v1.61.0","source_date":null,"evidence":[{"type":"repository_readme","claim":"Repository README identifies specification version v1.61.0."},{"type":"license","claim":"Repository LICENSE is Apache License 2.0."}],"unresolved_questions":["Which GenAI semantic conventions stabilize for long-term interoperability?","Which attributes should be retained in the hub as raw evidence versus normalized projections?"],"tags":["OpenTelemetry","semantic-conventions","agent-observability","FOSS","interoperability"],"procurement":{"license":"Apache-2.0","canonical_source":"https://github.com/open-telemetry/semantic-conventions","maintenance":"active project; verify cadence at ingestion time","portability":["ARM64 Linux","x86_64 Linux","Android/Termux: validate"],"offline_capability":"yes for local schema processing","interoperability":"OpenTelemetry semantic conventions","reproducibility":"pin repository release/tag when ingesting","provenance":"primary GitHub repository; observed 2026-09-27","dependency_risk":"low for schema consumption; runtime dependencies vary by adapter","lock_in_risk":"low at semantic-contract layer; exporter/provider lock-in remains adapter-specific","security_surface":"low for static schema; exporter/runtime integrations vary","resource_cost":"low for registry/schema processing","operational_fit":"agent observability evidence interchange","horizon":"H0","confidence":0.98,"decision_status":"watch"}}
97 changes: 31 additions & 66 deletions docs/ops/FORESIGHT-EVIDENCE-SPINE.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,46 +2,32 @@

## Purpose

Provide one portable evidence/provenance interchange layer for the monorepo's:

- daily research digest;
- Foresight Radar (H0/H1/H2/H3);
- FOSS procurement matrix;
- agent observability/evaluation;
- reproducibility and environment experiments;
- historical evidence and decision records.
Provide one portable evidence/provenance interchange layer for the monorepo's daily research digest, Foresight Radar, FOSS procurement matrix, agent observability/evaluation, reproducibility experiments, and historical evidence.

The spine is deliberately **not** a replacement for SQLite, OpenTelemetry, GitHub Actions, Langfuse, Phoenix, or other adapters. It is the stable, inspectable record exchanged between them.

## Architecture

```
Primary sources / papers / releases / advisories
|
v
retrieval/research agents
|
v
resource-registry.jsonl
| | |
| | +--> procurement export
| +----------> Foresight Radar
+------------------> daily digest
|
v
telemetry / mapper / evaluation adapters
```
## Contract v2

Every normalized evidence record has three stable identities:

- `resource_id`: durable identity for the source/resource being tracked.
- `item_id`: durable identity for a briefing/radar item derived from that resource.
- `lane_id`: stable research lane that owns the interpretation.

For compatibility, normalization derives `item_id` from `resource_id` and `lane_id` from `category` when older records omit them. New producers should write them explicitly.

`canonical_source` is a top-level copy of `source.url`; validation rejects divergence so exports and downstream systems have one canonical source field.

## Evidence semantics

| Horizon | Meaning |
|---|---|
| H0 | confirmed |
| H1 | emerging |
| H0 | confirmed / material |
| H1 | emerging / credible |
| H2 | weak signal |
| H3 | speculative |

Horizon and evidence status are separate. A source can be authoritative while a claim remains an attributed claim or early signal.
Horizon and evidence status are separate. A source can be authoritative while a claim remains an attributed claim, early signal, or disputed finding.

Evidence status values:

Expand All @@ -50,55 +36,34 @@ Evidence status values:
- `attributed_claim`
- `early_signal`
- `speculative`
- `disputed`

## Procurement matrix

Each resource may carry:
## Procurement

license, canonical source, maintenance, portability, offline capability,
interoperability, reproducibility, provenance, dependency risk, lock-in risk,
security surface, resource cost, operational fit, horizon, confidence, and
decision status.
The procurement object carries license, canonical source, maintenance, portability, offline capability, interoperability, reproducibility, provenance, dependency risk, lock-in risk, security surface, resource cost, operational fit, horizon, confidence, and decision status.

Decision status is descriptive state, not a universal ranking:
Decision status is descriptive state, not a universal ranking: `watch`, `investigate`, `prototype`, `adopt`, `reject`, `defer`.

- `watch`
- `investigate`
- `prototype`
- `adopt`
- `reject`
- `defer`
## Integrity

## Integrity model
Records are append-oriented JSONL. Corrections append a newer record with the same `resource_id`; consumers deduplicate to the latest record.

Records are append-oriented JSONL. Corrections append a newer record using the
same `resource_id`; consumers deduplicate to the latest record.
Each normalized record receives a SHA-256 `evidence_hash` over canonical JSON with the hash field excluded. No hosted service is required.

Each normalized record receives a SHA-256 `evidence_hash` over canonical JSON
with the hash field excluded. This detects accidental mutation without requiring
a hosted service.
## Radar query

## Portability
The portable CLI exposes:

The implementation is Python standard-library only and is intended to run on:
`python -m foresight.registry --registry data/foresight/resource-registry.jsonl radar --horizon H1 --lane agent-observability`

- Termux/Android;
- Linux;
- CI runners;
- Docker;
- Codespaces.
This keeps radar extraction deterministic and machine-readable without requiring a dashboard or vendor service.

No network access is required by the registry itself. Retrieval remains an
upstream concern.
## Operational validation

## Security
CI executes the unit suite and validates the committed registry corpus. An empty-registry initialization smoke-test is not sufficient evidence that the contract is exercised.

Never store secrets, cookies, API tokens, browser profiles, or credentials in
the registry. Reference protected artifacts by hash or controlled identifier.
## Portability and security

## Operational rule
Python standard-library only; intended for Termux/Android, Linux, CI, Docker, and Codespaces. The registry requires no network access.

`QUEUED`, `IN_PROGRESS`, and `COMPLETED` are runtime states, not evidence
of correctness. Evidence records should link to the execution/run/attempt and
preserve source SHA, timestamps, artifact identifiers, and validation status
when available.
Never store secrets, cookies, API tokens, browser profiles, or credentials in evidence records. Reference protected artifacts by hash or controlled identifier.
57 changes: 25 additions & 32 deletions foresight/registry.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,28 +10,18 @@
from typing import Any, Iterable

HORIZONS = {"H0", "H1", "H2", "H3"}
EVIDENCE_STATUSES = {
"confirmed",
"research_finding",
"attributed_claim",
"early_signal",
"speculative",
}
EVIDENCE_STATUSES = {"confirmed", "research_finding", "attributed_claim", "early_signal", "speculative", "disputed"}
DECISIONS = {"watch", "investigate", "prototype", "adopt", "reject", "defer"}

PROCUREMENT_FIELDS = [
"license", "canonical_source", "maintenance", "portability",
"offline_capability", "interoperability", "reproducibility",
"provenance", "dependency_risk", "lock_in_risk", "security_surface",
"resource_cost", "operational_fit", "horizon", "confidence",
"decision_status",
"license", "canonical_source", "maintenance", "portability", "offline_capability",
"interoperability", "reproducibility", "provenance", "dependency_risk",
"lock_in_risk", "security_surface", "resource_cost", "operational_fit",
"horizon", "confidence", "decision_status",
]

REQUIRED = [
"resource_id", "title", "category", "horizon",
"evidence_status", "summary", "source", "observed_at",
"resource_id", "item_id", "lane_id", "title", "category", "horizon",
"evidence_status", "summary", "source", "canonical_source", "observed_at",
]

DEFAULT_REGISTRY = Path("data/foresight/resource-registry.jsonl")


Expand All @@ -51,10 +41,14 @@ def fingerprint(record: dict[str, Any]) -> str:

def normalize(record: dict[str, Any]) -> dict[str, Any]:
out = dict(record)
out.setdefault("resource_id", "")
out.setdefault("item_id", out.get("resource_id", ""))
out.setdefault("lane_id", out.get("category", "unclassified"))
out.setdefault("observed_at", now_utc())
out.setdefault("captured_at", now_utc())
out.setdefault("source_type", "unknown")
out.setdefault("source_version", None)
out.setdefault("version_or_commit", out.get("source_version"))
out.setdefault("source_date", None)
out.setdefault("evidence", [])
out.setdefault("unresolved_questions", [])
Expand All @@ -63,6 +57,9 @@ def normalize(record: dict[str, Any]) -> dict[str, Any]:
out.setdefault("termux_relevance", None)
out.setdefault("procurement", {})
out.setdefault("tags", [])
source = out.get("source")
if isinstance(source, dict):
out.setdefault("canonical_source", source.get("url"))
out["horizon"] = str(out.get("horizon", "")).upper()
out["evidence_status"] = str(out.get("evidence_status", "")).lower()
out["confidence"] = float(out.get("confidence", 0.0))
Expand All @@ -75,20 +72,18 @@ def validate(record: dict[str, Any]) -> list[str]:
for key in REQUIRED:
if not record.get(key):
errors.append(f"missing required field: {key}")

if record.get("horizon") not in HORIZONS:
errors.append("horizon must be one of H0/H1/H2/H3")
if record.get("evidence_status") not in EVIDENCE_STATUSES:
errors.append("invalid evidence_status")

confidence = record.get("confidence")
if not isinstance(confidence, (int, float)) or not 0 <= confidence <= 1:
errors.append("confidence must be in [0,1]")

source = record.get("source")
if not isinstance(source, dict) or not source.get("url"):
errors.append("source.url is required")

if record.get("canonical_source") != (source or {}).get("url"):
errors.append("canonical_source must match source.url")
procurement = record.get("procurement", {})
if not isinstance(procurement, dict):
errors.append("procurement must be an object")
Expand All @@ -100,7 +95,6 @@ def validate(record: dict[str, Any]) -> list[str]:
errors.append("procurement.horizon must be H0/H1/H2/H3")
if procurement.get("decision_status") and procurement["decision_status"] not in DECISIONS:
errors.append("invalid procurement.decision_status")

return errors


Expand Down Expand Up @@ -142,7 +136,7 @@ def export_csv(records: Iterable[dict[str, Any]], destination: Path) -> None:
records = list(records)
destination.parent.mkdir(parents=True, exist_ok=True)
columns = REQUIRED + [
"source_type", "source_version", "source_date",
"source_type", "source_version", "source_date", "version_or_commit",
"why_it_matters", "practical_opportunity", "termux_relevance",
"confidence", "evidence_hash",
] + PROCUREMENT_FIELDS
Expand Down Expand Up @@ -191,18 +185,17 @@ def validate_cmd_fn(a):

radar = sub.add_parser("radar")
radar.add_argument("--horizon", action="append", choices=sorted(HORIZONS))
radar.add_argument("--lane")
def radar_cmd(a):
allowed = set(a.horizon or HORIZONS)
for r in sorted(dedupe(load(a.registry)), key=lambda x: (x["horizon"], x["title"].lower())):
if r["horizon"] in allowed:
lane = a.lane
for r in sorted(dedupe(load(a.registry)), key=lambda x: (x["horizon"], x["lane_id"], x["title"].lower())):
if r["horizon"] in allowed and (lane is None or r["lane_id"] == lane):
print(json.dumps({
"resource_id": r["resource_id"],
"horizon": r["horizon"],
"evidence_status": r["evidence_status"],
"confidence": r["confidence"],
"title": r["title"],
"source": r["source"]["url"],
"termux_relevance": r.get("termux_relevance"),
"resource_id": r["resource_id"], "item_id": r["item_id"], "lane_id": r["lane_id"],
"horizon": r["horizon"], "evidence_status": r["evidence_status"],
"confidence": r["confidence"], "title": r["title"],
"source": r["canonical_source"], "termux_relevance": r.get("termux_relevance"),
"decision_status": r.get("procurement", {}).get("decision_status"),
}, ensure_ascii=False))
return 0
Expand Down
16 changes: 8 additions & 8 deletions foresight/schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,27 +2,27 @@
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Termux Monorepo Foresight Evidence Record",
"type": "object",
"required": ["resource_id","title","category","horizon","evidence_status","summary","source","observed_at"],
"required": ["resource_id","item_id","lane_id","title","category","horizon","evidence_status","summary","source","canonical_source","observed_at"],
"properties": {
"resource_id": {"type":"string"},
"resource_id": {"type":"string","minLength":1},
"item_id": {"type":"string","minLength":1},
"lane_id": {"type":"string","minLength":1},
"title": {"type":"string"},
"category": {"type":"string"},
"horizon": {"enum":["H0","H1","H2","H3"]},
"evidence_status": {"enum":["confirmed","research_finding","attributed_claim","early_signal","speculative"]},
"evidence_status": {"enum":["confirmed","research_finding","attributed_claim","early_signal","speculative","disputed"]},
"summary": {"type":"string"},
"why_it_matters": {"type":["string","null"]},
"practical_opportunity": {"type":["string","null"]},
"termux_relevance": {"type":["string","null"]},
"confidence": {"type":"number","minimum":0,"maximum":1},
"observed_at": {"type":"string"},
"captured_at": {"type":"string"},
"source": {
"type":"object",
"required":["url"],
"properties":{"url":{"type":"string"},"publisher":{"type":"string"},"title":{"type":"string"}}
},
"source": {"type":"object","required":["url"],"properties":{"url":{"type":"string","format":"uri"},"publisher":{"type":"string"},"title":{"type":"string"}}},
"canonical_source":{"type":"string","format":"uri"},
"source_type":{"type":"string"},
"source_version":{"type":["string","null"]},
"version_or_commit":{"type":["string","null"]},
"source_date":{"type":["string","null"]},
"evidence":{"type":"array"},
"unresolved_questions":{"type":"array"},
Expand Down
18 changes: 17 additions & 1 deletion tests/test_foresight_registry.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
import unittest
from pathlib import Path

from foresight.registry import append, fingerprint, load, validate
from foresight.registry import append, fingerprint, load, normalize, validate


class ForesightRegistryTests(unittest.TestCase):
Expand All @@ -20,6 +20,12 @@ def record(self):
"procurement": {"horizon": "H0", "decision_status": "watch"}
}

def test_normalize_derives_stable_lane_and_item_identity(self):
r = normalize(self.record())
self.assertEqual(r["lane_id"], "testing")
self.assertEqual(r["item_id"], "test-1")
self.assertEqual(r["canonical_source"], "https://example.org/source")

def test_append_validate_and_hash(self):
with tempfile.TemporaryDirectory() as d:
path = Path(d) / "resource-registry.jsonl"
Expand All @@ -33,6 +39,16 @@ def test_invalid_horizon(self):
record["horizon"] = "H9"
self.assertTrue(validate(record))

def test_invalid_canonical_source(self):
record = normalize(self.record())
record["canonical_source"] = "https://example.org/other"
self.assertTrue(validate(record))

def test_disputed_evidence_is_supported(self):
record = normalize(self.record())
record["evidence_status"] = "disputed"
self.assertEqual(validate(record), [])

def test_hash_changes_when_record_changes(self):
a = self.record()
b = dict(a)
Expand Down
Loading