diff --git a/CHANGELOG.md b/CHANGELOG.md index 062a69412..2af1240de 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -38,6 +38,8 @@ All notable changes to TEPP are documented here. The format follows Keep a Chang ## [Unreleased] +- **Outcome-order analysis-run profile**: `analysis_engine` binds existing `outcome_order::OutcomeKind`, `refuse_reverse_ipo_order`, and `refuse_outcome_of_as_transition` to cutoff-safe `outcome_order_v1` (`tepp.outcome_order.v1`) with inference status `input_process_forward_outcome_of_is_not_transition`. `kind_recovery_rate` stays library-side. Not membership-target, not location-membership, not membership-posterior ICC, not copied-text, not copy-identity, not citation-edge, not GPU, not MCMC, and not topic birth/split/merge. + - `event_core` adds bounded Allen interval-consistency classification, atomic path-consistency closure, contradiction/resource refusals, and an explicit dependency-error fallback without claiming unrestricted global satisfiability. - `psychometric_core` recovers the Driver, Oud, and Voelkle (2017, Table 2, p. 12 `MANIFESTTRAITVAR`; §7.1, p. 19; p. 16 `MANIFESTTRAITVARstd`; footnote 4; 2017-era ctsem `summary.ctsemFit.R`; JSS PDF re-opened 2026-08-27T14:20Z from https://www.jstatsoft.org/index.php/jss/article/download/v077i05/1104) scalar standardised manifest-trait variance on current main after `0ce16e8` dropped the pre-consolidation code while research notes already named the map (register items 83–84). Table 2 names `MANIFESTTRAITVAR` `Ψ_τ` the additional time-invariant variance-covariance on the measurement level and sets it `NULL` when there is no manifest trait. Equation 5 writes `Γ ~ N(τ, Ψ)` and names that covariance the manifest traits. Section 7.1 names manifest traits stable individual differences in indicator levels, distinct from process-level `TRAITVAR` `φ_ξ`. Page 16 prints standardised matrices with the suffix `std` when appropriate. The printed example on p. 16 is `discreteDRIFTstd`, not `MANIFESTTRAITVARstd`. Footnote 4 standardises using only the relevant variance, not the total. The relevant variance for that named indicator-level correlation is `MANIFESTTRAITVAR`, not process-level `TRAITVAR` and not residual `MANIFESTVAR` `θ`. The 2017-era source forms `MANIFESTTRAITVARstd` only when `MANIFESTTRAITVAR != 0`, as `solve(sqrt(diag(MANIFESTTRAITVAR) + ridging)) %&% MANIFESTTRAITVAR` when `verbose = TRUE`. OpenMx `%&%` is `t(A) %*% B %*% A`. Unlike `TRAITVARstd`, that formation adds `diag(c(ridging), n.manifest)`. The default `ridging = FALSE` adds 0, not `0.0001`; that ridge is a numerical hack and is not this exact map. The scalar correlation is `ψ / ψ = 1` after strictly positive `MANIFESTTRAITVAR`. Form strictly positive `ψ` first, then `1 / √ψ`, then `(1 / √ψ) ψ (1 / √ψ)`. Unstandardised `MANIFESTTRAITVAR` is defined for a zero trait; standardised `MANIFESTTRAITVAR` is not. Zero `MANIFESTTRAITVAR` skips forming `MANIFESTTRAITVARstd` in the 2017-era source and fails closed here. Indicator-level trait variance is an event-time structural quantity, so a non-event clock fails closed. `MANIFESTTRAITVAR` does not require stable `a < 0`. Distinct positive `ψ` recover the same 1. `trait / trait = 1` is `TRAITVARstd` and recovers the same number and remains a distinct named quantity. `θ` is `MANIFESTVAR` and is measurement error, not this correlation. Meredith (1993) remains unread (web search 2026-08-27T14:20Z: Springer/Cambridge Core paywalled; Unpaywall historically `is_oa: false`; Springer `content/pdf` is an HTML stub). Mislevy (1991, *Psychometrika, 56*, 177–196) remains unread on the same terms (DOI `10.1007/bf02294457`). Still not a Kalman filter, not a matrix `expm`, not ESEM estimation, not DSEM, and not ctsem estimation. diff --git a/Cargo.lock b/Cargo.lock index 454a7d612..f37dcbc0d 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -74,6 +74,7 @@ dependencies = [ "corpus_split", "event_core", "membership_core", + "outcome_order", "relation_graph", "serde", "serde_json", diff --git a/DOCUMENTATION.md b/DOCUMENTATION.md index 6fa4b9683..c2f853aed 100644 --- a/DOCUMENTATION.md +++ b/DOCUMENTATION.md @@ -71,6 +71,7 @@ TEPP's approved PRD v0.4 and implementation plan are the primary product baselin | Hourly NIM OpenCode doctoring | [`docs/doctoring/hourly-nim-opencode-development.md`](docs/doctoring/hourly-nim-opencode-development.md) | | Analysis engine v1 doctoring | [`docs/doctoring/analysis-engine-v1.md`](docs/doctoring/analysis-engine-v1.md) | | Analysis engine gap-closure doctoring | [`docs/doctoring/analysis-engine-gap-closure.md`](docs/doctoring/analysis-engine-gap-closure.md) | +| Outcome-order analysis-run doctoring | [`docs/doctoring/outcome-order-analysis-run.md`](docs/doctoring/outcome-order-analysis-run.md) | | Corpus-split leakage-audit wire doctoring | [`docs/research/corpus-split-manifest-wire.md`](docs/research/corpus-split-manifest-wire.md) | | Unicode canonical-identity doctoring | [`docs/research/unicode-canonical-identity.md`](docs/research/unicode-canonical-identity.md) | | Change history | [`CHANGELOG.md`](CHANGELOG.md) | diff --git a/crates/analysis_engine/Cargo.toml b/crates/analysis_engine/Cargo.toml index 7322212b2..e190852e8 100644 --- a/crates/analysis_engine/Cargo.toml +++ b/crates/analysis_engine/Cargo.toml @@ -15,6 +15,7 @@ publish = false [dependencies] event_core = { path = "../event_core", version = "0.2.0" } +outcome_order = { path = "../outcome_order", version = "0.2.0" } serde = { workspace = true } serde_json = { workspace = true } sha2 = { workspace = true } diff --git a/crates/analysis_engine/src/lib.rs b/crates/analysis_engine/src/lib.rs index 72bd5854c..527a55f57 100644 --- a/crates/analysis_engine/src/lib.rs +++ b/crates/analysis_engine/src/lib.rs @@ -12,6 +12,7 @@ mod case_deletion_refit; mod lineage_criterion; +mod outcome_order_artifact; mod topic_context_posterior; mod topic_lineage_artifact; @@ -46,6 +47,12 @@ pub use lineage_criterion::{ LineageCriterionFit, LineageCriterionFitError, LineageCriterionObservation, fit_lineage_criterion_posteriors, }; +/// Outcome-order artifact and execution contracts from this engine. +pub use outcome_order_artifact::{ + OUTCOME_ORDER_ARTIFACT_BYTE_LIMIT, OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION, + OUTCOME_ORDER_MODEL_CONTRACT_VERSION, OUTCOME_ORDER_OUTPUT_PROFILE, OutcomeOrderArtifact, + OutcomeOrderEdge, OutcomeOrderExecution, execute_outcome_order_run, +}; /// Bounded posterior topic-context producer contract and record types. pub use topic_context_posterior::{ TOPIC_CONTEXT_POSTERIOR_BYTE_LIMIT, TOPIC_CONTEXT_POSTERIOR_SCHEMA_VERSION, @@ -248,6 +255,8 @@ pub enum AnalysisEngineError { TopicMeasurement(TopicMeasurementError), /// A topic-lineage artifact violated its bounded schema or count invariants. InvalidTopicLineageArtifact, + /// An outcome-order artifact violated its bounded schema or count invariants. + InvalidOutcomeOrderArtifact, } impl fmt::Display for AnalysisEngineError { @@ -262,6 +271,7 @@ impl fmt::Display for AnalysisEngineError { Self::LimitExceeded => "analysis corpus exceeded its execution bound", Self::TopicMeasurement(error) => return error.fmt(formatter), Self::InvalidTopicLineageArtifact => "invalid topic lineage artifact", + Self::InvalidOutcomeOrderArtifact => "invalid outcome-order artifact", }; formatter.write_str(message) } @@ -681,6 +691,10 @@ mod tests { AnalysisEngineError::InvalidTopicLineageArtifact, "invalid topic lineage artifact", ), + ( + AnalysisEngineError::InvalidOutcomeOrderArtifact, + "invalid outcome-order artifact", + ), ]; for (error, message) in messages { assert_eq!(error.to_string(), message); diff --git a/crates/analysis_engine/src/outcome_order_artifact.rs b/crates/analysis_engine/src/outcome_order_artifact.rs new file mode 100644 index 000000000..e2eb25156 --- /dev/null +++ b/crates/analysis_engine/src/outcome_order_artifact.rs @@ -0,0 +1,442 @@ +//! Digest-bound input-process-outcome order refusals as an analysis-run profile. + +use outcome_order::{ + OutcomeKind, OutcomeOrderError, refuse_outcome_of_as_transition, refuse_reverse_ipo_order, +}; +use serde::{Deserialize, Serialize}; +use sha2::{Digest, Sha256}; +use temporal_core::{AvailableTime, KnowledgeCutoff}; +use tepp_api::{ + AnalysisResultSummary, AnalysisRunAccepted, AnalysisRunRequest, AnalysisRunTerminalResult, +}; + +use crate::{ + AnalysisEngineError, MAX_EVIDENCE_UNITS, format_digest, require_receipt_identity, + valid_identifier, +}; + +/// Versioned schema for a completed outcome-order artifact. +pub const OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION: &str = "tepp.outcome_order.v1"; +/// Model contract required by the outcome-order execution path. +pub const OUTCOME_ORDER_MODEL_CONTRACT_VERSION: &str = "outcome_order_v1"; +/// Analysis-run output profile required for an outcome-order artifact. +pub const OUTCOME_ORDER_OUTPUT_PROFILE: &str = "outcome_order_v1"; +/// Maximum canonical artifact JSON size. +pub const OUTCOME_ORDER_ARTIFACT_BYTE_LIMIT: usize = 256 * 1024; +const OUTCOME_ORDER_INFERENCE_STATUS: &str = "input_process_forward_outcome_of_is_not_transition"; + +/// One cutoff-admitted IPO edge with closed kind and opaque event-time ranks. +#[derive(Clone, Debug, Eq, PartialEq)] +pub struct OutcomeOrderEdge { + edge_id: String, + kind: OutcomeKind, + source_rank: u64, + target_rank: u64, + available_time: AvailableTime, +} + +impl OutcomeOrderEdge { + /// Construct a bounded outcome-order edge. + /// + /// # Errors + /// + /// Returns [`AnalysisEngineError::InvalidEvidence`] when the edge identity + /// is empty or oversized. + pub fn new( + edge_id: impl Into, + kind: OutcomeKind, + source_rank: u64, + target_rank: u64, + available_time: AvailableTime, + ) -> Result { + let edge_id = edge_id.into(); + if !valid_identifier(&edge_id) { + return Err(AnalysisEngineError::InvalidEvidence); + } + Ok(Self { + edge_id, + kind, + source_rank, + target_rank, + available_time, + }) + } + + /// Return the opaque edge identity. + #[must_use] + pub fn edge_id(&self) -> &str { + &self.edge_id + } + + /// Return the closed IPO kind. + #[must_use] + pub const fn kind(&self) -> OutcomeKind { + self.kind + } + + /// Return the opaque source event-time rank. + #[must_use] + pub const fn source_rank(&self) -> u64 { + self.source_rank + } + + /// Return the opaque target event-time rank. + #[must_use] + pub const fn target_rank(&self) -> u64 { + self.target_rank + } + + /// Return the availability time used for cutoff eligibility. + #[must_use] + pub const fn available_time(&self) -> AvailableTime { + self.available_time + } +} + +/// Completed, bounded IPO-order census for analysis-run clients. +#[derive(Clone, Debug, Deserialize, PartialEq, Serialize)] +#[serde(deny_unknown_fields)] +pub struct OutcomeOrderArtifact { + /// Exact versioned schema identity. + pub schema_version: String, + /// Opaque accepted-run identity. + pub run_id: String, + /// Immutable source snapshot identity. + pub snapshot_id: String, + /// Historical evidence cutoff used to admit edges. + pub knowledge_cutoff: String, + /// Number of edges admitted at the cutoff. + pub edge_count: u64, + /// Forward `input_to` transitions admitted at the cutoff. + pub input_to_count: u64, + /// Forward `process_to` transitions admitted at the cutoff. + pub process_to_count: u64, + /// `outcome_of` provenance edges admitted at the cutoff. + pub outcome_of_count: u64, + /// Provenance edges refused as state transitions. + pub refused_as_transition_count: u64, + /// Fixed claim boundary for consumer copy. + pub inference_status: String, +} + +impl OutcomeOrderArtifact { + /// Parse and fully validate a bounded artifact JSON payload. + /// + /// # Errors + /// + /// Returns [`AnalysisEngineError::InvalidOutcomeOrderArtifact`] when the + /// schema, identifiers, counts, or claim boundary fail. + pub fn from_json(payload: &str) -> Result { + if payload.len() > OUTCOME_ORDER_ARTIFACT_BYTE_LIMIT { + return Err(AnalysisEngineError::LimitExceeded); + } + let artifact: Self = serde_json::from_str(payload) + .map_err(|_| AnalysisEngineError::InvalidOutcomeOrderArtifact)?; + artifact.validate()?; + Ok(artifact) + } + + /// Serialize canonical validated artifact JSON. + /// + /// # Errors + /// + /// Returns a typed validation, serialization, or size failure. + pub fn to_json(&self) -> Result { + self.validate()?; + let payload = + serde_json::to_string(self).map_err(|_| AnalysisEngineError::SerializationFailure)?; + if payload.len() > OUTCOME_ORDER_ARTIFACT_BYTE_LIMIT { + return Err(AnalysisEngineError::LimitExceeded); + } + Ok(payload) + } + + /// Return the lowercase SHA-256 digest of canonical artifact JSON. + /// + /// # Errors + /// + /// Returns a typed validation or serialization failure. + pub fn sha256(&self) -> Result { + self.to_json() + .map(|json| format_digest(Sha256::digest(json.into_bytes()))) + } + + fn validate(&self) -> Result<(), AnalysisEngineError> { + let transition_sum = self.input_to_count.checked_add(self.process_to_count); + let kind_sum = transition_sum.and_then(|value| value.checked_add(self.outcome_of_count)); + if self.schema_version != OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION + || !valid_identifier(&self.run_id) + || !valid_identifier(&self.snapshot_id) + || KnowledgeCutoff::parse_rfc3339(&self.knowledge_cutoff).is_err() + || self.edge_count < 3 + || self.edge_count > MAX_EVIDENCE_UNITS as u64 + || self.input_to_count == 0 + || self.process_to_count == 0 + || self.outcome_of_count == 0 + || kind_sum != Some(self.edge_count) + || self.refused_as_transition_count != self.outcome_of_count + || self.inference_status != OUTCOME_ORDER_INFERENCE_STATUS + { + return Err(AnalysisEngineError::InvalidOutcomeOrderArtifact); + } + Ok(()) + } +} + +/// One completed outcome-order artifact and its terminal result. +#[derive(Clone, Debug, PartialEq)] +pub struct OutcomeOrderExecution { + /// Digest-bound completed IPO-order census. + pub artifact: OutcomeOrderArtifact, + /// Terminal result carrying the artifact identity, digest, and schema. + pub terminal_result: AnalysisRunTerminalResult, +} + +/// Execute cutoff-safe IPO-order refusals as one analysis-run profile. +/// +/// The executor invokes [`refuse_reverse_ipo_order`] and +/// [`refuse_outcome_of_as_transition`] already on protected main. `input_to` +/// and `process_to` stay forward transitions. `outcome_of` stays provenance. +/// It does not emit `kind_recovery_rate`, a `scientific_acceptance` inspect +/// metric, GPU kernels, MCMC, or topic birth/split/merge events. +/// +/// # Errors +/// +/// Returns a request/receipt/snapshot/cutoff/profile error, empty or +/// single-class corpus, reverse or uncertain IPO order, duplicate edge +/// identity, oversized corpus, or invalid artifact error. +#[allow(clippy::too_many_lines)] +pub fn execute_outcome_order_run( + request: &AnalysisRunRequest, + accepted: &AnalysisRunAccepted, + snapshot_id: &str, + knowledge_cutoff: KnowledgeCutoff, + edges: &[OutcomeOrderEdge], + completed_at: impl Into, +) -> Result { + request.to_json()?; + accepted.to_json()?; + require_receipt_identity(request, accepted)?; + if request.snapshot_id != snapshot_id { + return Err(AnalysisEngineError::SnapshotMismatch); + } + let request_cutoff = KnowledgeCutoff::parse_rfc3339(&request.knowledge_cutoff) + .map_err(|_| AnalysisEngineError::InvalidEvidence)?; + if request_cutoff.instant() != knowledge_cutoff.instant() + || request.model_contract_version != OUTCOME_ORDER_MODEL_CONTRACT_VERSION + || request.output_profile != OUTCOME_ORDER_OUTPUT_PROFILE + { + return Err(AnalysisEngineError::InvalidEvidence); + } + if edges.len() > MAX_EVIDENCE_UNITS { + return Err(AnalysisEngineError::LimitExceeded); + } + + let mut seen = std::collections::BTreeSet::new(); + let mut input_to_count = 0_u64; + let mut process_to_count = 0_u64; + let mut outcome_of_count = 0_u64; + let mut refused_as_transition_count = 0_u64; + for edge in edges { + if !seen.insert(edge.edge_id()) { + return Err(AnalysisEngineError::DuplicateEvidence); + } + if edge.available_time().instant() > knowledge_cutoff.instant() { + continue; + } + match edge.kind() { + OutcomeKind::InputTo => { + refuse_reverse_ipo_order(edge.kind(), edge.source_rank(), edge.target_rank()) + .map_err(map_outcome_order_error)?; + refuse_outcome_of_as_transition(edge.kind()).map_err(map_outcome_order_error)?; + input_to_count = increment(input_to_count)?; + } + OutcomeKind::ProcessTo => { + refuse_reverse_ipo_order(edge.kind(), edge.source_rank(), edge.target_rank()) + .map_err(map_outcome_order_error)?; + refuse_outcome_of_as_transition(edge.kind()).map_err(map_outcome_order_error)?; + process_to_count = increment(process_to_count)?; + } + OutcomeKind::OutcomeOf => { + match refuse_outcome_of_as_transition(edge.kind()) { + Err(OutcomeOrderError::OutcomeOfIsNotTransition) => { + refused_as_transition_count = increment(refused_as_transition_count)?; + } + Ok(()) | Err(_) => return Err(AnalysisEngineError::InvalidEvidence), + } + outcome_of_count = increment(outcome_of_count)?; + } + } + } + + let edge_count = input_to_count + .checked_add(process_to_count) + .and_then(|value| value.checked_add(outcome_of_count)) + .ok_or(AnalysisEngineError::ArithmeticOverflow)?; + if edge_count < 3 + || input_to_count == 0 + || process_to_count == 0 + || outcome_of_count == 0 + || refused_as_transition_count != outcome_of_count + { + return Err(AnalysisEngineError::InvalidEvidence); + } + + let artifact = OutcomeOrderArtifact { + schema_version: OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION.into(), + run_id: accepted.run_id.clone(), + snapshot_id: snapshot_id.to_owned(), + knowledge_cutoff: knowledge_cutoff.to_rfc3339(), + edge_count, + input_to_count, + process_to_count, + outcome_of_count, + refused_as_transition_count, + inference_status: OUTCOME_ORDER_INFERENCE_STATUS.into(), + }; + let digest = artifact.sha256()?; + let summary = AnalysisResultSummary::new("outcome_order", edge_count, 4, "validated")?; + let terminal_result = AnalysisRunTerminalResult::succeeded( + request, + accepted, + format!("outcome_order_artifact_{}", &digest[..16]), + digest, + OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION, + completed_at, + summary, + )?; + Ok(OutcomeOrderExecution { + artifact, + terminal_result, + }) +} + +fn increment(count: u64) -> Result { + count + .checked_add(1) + .ok_or(AnalysisEngineError::ArithmeticOverflow) +} + +fn map_outcome_order_error(error: OutcomeOrderError) -> AnalysisEngineError { + match error { + OutcomeOrderError::ReverseIpoOrder + | OutcomeOrderError::UncertainIpoOrder + | OutcomeOrderError::OutcomeOfIsNotTransition + | OutcomeOrderError::InvalidEdgePayload + | _ => AnalysisEngineError::InvalidEvidence, + } +} + +#[cfg(test)] +mod tests { + use super::{ + OUTCOME_ORDER_ARTIFACT_BYTE_LIMIT, OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION, + OUTCOME_ORDER_INFERENCE_STATUS, OutcomeOrderArtifact, + }; + use crate::AnalysisEngineError; + + fn artifact() -> OutcomeOrderArtifact { + OutcomeOrderArtifact { + schema_version: OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION.into(), + run_id: "run-1".into(), + snapshot_id: "snapshot-1".into(), + knowledge_cutoff: "2026-08-01T00:00:00Z".into(), + edge_count: 3, + input_to_count: 1, + process_to_count: 1, + outcome_of_count: 1, + refused_as_transition_count: 1, + inference_status: OUTCOME_ORDER_INFERENCE_STATUS.into(), + } + } + + fn assert_invalid(artifact: &OutcomeOrderArtifact) { + assert_eq!( + artifact.to_json(), + Err(AnalysisEngineError::InvalidOutcomeOrderArtifact) + ); + } + + #[test] + fn artifact_round_trip_and_size_bounds_fail_closed() { + let artifact = artifact(); + let payload = artifact.to_json().expect("json"); + assert_eq!( + OutcomeOrderArtifact::from_json(&payload), + Ok(artifact.clone()) + ); + assert_eq!(artifact.sha256().expect("digest").len(), 64); + assert_eq!( + OutcomeOrderArtifact::from_json("{}"), + Err(AnalysisEngineError::InvalidOutcomeOrderArtifact) + ); + assert_eq!( + OutcomeOrderArtifact::from_json(&"x".repeat(OUTCOME_ORDER_ARTIFACT_BYTE_LIMIT + 1)), + Err(AnalysisEngineError::LimitExceeded) + ); + } + + #[test] + fn artifact_metadata_tampering_fails_closed() { + let artifact = artifact(); + let invalid_artifacts = [ + { + let mut value = artifact.clone(); + value.schema_version.clear(); + value + }, + { + let mut value = artifact.clone(); + value.run_id.clear(); + value + }, + { + let mut value = artifact.clone(); + value.snapshot_id.clear(); + value + }, + { + let mut value = artifact.clone(); + value.knowledge_cutoff = "invalid".into(); + value + }, + { + let mut value = artifact.clone(); + value.edge_count = 2; + value + }, + { + let mut value = artifact.clone(); + value.input_to_count = 0; + value.edge_count = 2; + value + }, + { + let mut value = artifact.clone(); + value.process_to_count = 0; + value.edge_count = 2; + value + }, + { + let mut value = artifact.clone(); + value.outcome_of_count = 0; + value.refused_as_transition_count = 0; + value.edge_count = 2; + value + }, + { + let mut value = artifact.clone(); + value.refused_as_transition_count = 0; + value + }, + { + let mut value = artifact.clone(); + value.inference_status.clear(); + value + }, + ]; + for invalid in invalid_artifacts { + assert_invalid(&invalid); + } + } +} diff --git a/crates/analysis_engine/tests/outcome_order_cutoff_semantics_contract.rs b/crates/analysis_engine/tests/outcome_order_cutoff_semantics_contract.rs new file mode 100644 index 000000000..7568fcedc --- /dev/null +++ b/crates/analysis_engine/tests/outcome_order_cutoff_semantics_contract.rs @@ -0,0 +1,66 @@ +//! Regression contract for semantic knowledge-cutoff equality in outcome-order runs. + +use analysis_engine::{ + OUTCOME_ORDER_MODEL_CONTRACT_VERSION, OUTCOME_ORDER_OUTPUT_PROFILE, OutcomeOrderEdge, + execute_outcome_order_run, +}; +use outcome_order::OutcomeKind; +use temporal_core::{AvailableTime, KnowledgeCutoff}; +use tepp_api::{AnalysisRunAccepted, AnalysisRunRequest}; + +fn available(stamp: &str) -> AvailableTime { + AvailableTime::parse_rfc3339(stamp).expect("available time") +} + +fn edge( + edge_id: &str, + kind: OutcomeKind, + source_rank: u64, + target_rank: u64, +) -> OutcomeOrderEdge { + OutcomeOrderEdge::new( + edge_id, + kind, + source_rank, + target_rank, + available("2026-07-01T00:00:00Z"), + ) + .expect("edge") +} + +#[test] +fn equivalent_rfc3339_cutoff_offsets_are_admitted() { + let request = AnalysisRunRequest { + contract_version: 1, + idempotency_key: "outcome-order-equivalent-cutoff".into(), + tenant_workspace_id: "tenant-workspace".into(), + snapshot_id: "snapshot-outcome-order".into(), + knowledge_cutoff: "2026-08-01T09:00:00+09:00".into(), + model_contract_version: OUTCOME_ORDER_MODEL_CONTRACT_VERSION.into(), + output_profile: OUTCOME_ORDER_OUTPUT_PROFILE.into(), + }; + let accepted = AnalysisRunAccepted::new("run-outcome-order", "accepted", &request.idempotency_key) + .expect("accepted"); + let execution_cutoff = + KnowledgeCutoff::parse_rfc3339("2026-08-01T00:00:00Z").expect("execution cutoff"); + let edges = vec![ + edge("input-a", OutcomeKind::InputTo, 1, 2), + edge("process-b", OutcomeKind::ProcessTo, 2, 3), + edge("outcome-c", OutcomeKind::OutcomeOf, 9, 1), + ]; + + let execution = execute_outcome_order_run( + &request, + &accepted, + "snapshot-outcome-order", + execution_cutoff, + &edges, + "2026-08-02T00:00:00Z", + ) + .expect("equivalent RFC 3339 spellings denote the same cutoff instant"); + + assert_eq!( + execution.artifact.knowledge_cutoff, + execution_cutoff.to_rfc3339() + ); +} diff --git a/crates/analysis_engine/tests/outcome_order_execution_contract.rs b/crates/analysis_engine/tests/outcome_order_execution_contract.rs new file mode 100644 index 000000000..81b189c44 --- /dev/null +++ b/crates/analysis_engine/tests/outcome_order_execution_contract.rs @@ -0,0 +1,318 @@ +//! End-to-end contract for cutoff-safe input-process-outcome order refusals. + +use analysis_engine::{ + AnalysisEngineError, MAX_EVIDENCE_UNITS, OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION, + OUTCOME_ORDER_MODEL_CONTRACT_VERSION, OUTCOME_ORDER_OUTPUT_PROFILE, OutcomeOrderArtifact, + OutcomeOrderEdge, execute_outcome_order_run, +}; +use outcome_order::OutcomeKind; +use temporal_core::{AvailableTime, KnowledgeCutoff}; +use tepp_api::{AnalysisRunAccepted, AnalysisRunRequest, AnalysisRunTerminalState}; + +fn cutoff() -> KnowledgeCutoff { + KnowledgeCutoff::parse_rfc3339("2026-08-01T00:00:00Z").expect("cutoff") +} + +fn available(stamp: &str) -> AvailableTime { + AvailableTime::parse_rfc3339(stamp).expect("available") +} + +fn request() -> AnalysisRunRequest { + AnalysisRunRequest { + contract_version: 1, + idempotency_key: "outcome-order-idem".into(), + tenant_workspace_id: "tenant-workspace".into(), + snapshot_id: "snapshot-outcome-order".into(), + knowledge_cutoff: "2026-08-01T00:00:00Z".into(), + model_contract_version: OUTCOME_ORDER_MODEL_CONTRACT_VERSION.into(), + output_profile: OUTCOME_ORDER_OUTPUT_PROFILE.into(), + } +} + +fn accepted(request: &AnalysisRunRequest) -> AnalysisRunAccepted { + AnalysisRunAccepted::new("run-outcome-order", "accepted", &request.idempotency_key) + .expect("accepted") +} + +fn edge( + edge_id: &str, + kind: OutcomeKind, + source_rank: u64, + target_rank: u64, + stamp: &str, +) -> OutcomeOrderEdge { + OutcomeOrderEdge::new(edge_id, kind, source_rank, target_rank, available(stamp)).expect("edge") +} + +fn mixed_edges() -> Vec { + vec![ + edge( + "input-a", + OutcomeKind::InputTo, + 1, + 2, + "2026-07-01T00:00:00Z", + ), + edge( + "process-b", + OutcomeKind::ProcessTo, + 2, + 3, + "2026-07-02T00:00:00Z", + ), + edge( + "outcome-c", + OutcomeKind::OutcomeOf, + 9, + 1, + "2026-07-03T00:00:00Z", + ), + ] +} + +fn execute( + request: &AnalysisRunRequest, + edges: &[OutcomeOrderEdge], +) -> Result { + execute_outcome_order_run( + request, + &accepted(request), + "snapshot-outcome-order", + cutoff(), + edges, + "2026-08-02T00:00:00Z", + ) +} + +#[test] +fn mixed_ipo_kinds_emit_digest_bound_refusals_without_recovery_metric() { + let request = request(); + let execution = execute(&request, &mixed_edges()).expect("execution"); + assert_eq!( + execution.artifact.schema_version, + OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION + ); + assert_eq!(execution.artifact.edge_count, 3); + assert_eq!(execution.artifact.input_to_count, 1); + assert_eq!(execution.artifact.process_to_count, 1); + assert_eq!(execution.artifact.outcome_of_count, 1); + assert_eq!(execution.artifact.refused_as_transition_count, 1); + assert_eq!( + execution.artifact.inference_status, + "input_process_forward_outcome_of_is_not_transition" + ); + let payload = execution.artifact.to_json().expect("json"); + assert!(!payload.contains("kind_recovery_rate")); + assert!(!payload.contains("scientific_acceptance")); + assert_eq!( + execution.terminal_result.run_state, + AnalysisRunTerminalState::Succeeded + ); + assert_eq!( + execution + .terminal_result + .summary + .as_ref() + .expect("summary") + .validation_status, + "validated" + ); + assert_eq!( + execution.terminal_result.result_sha256.as_deref(), + Some(execution.artifact.sha256().expect("digest").as_str()) + ); + assert_eq!( + execution.terminal_result.result_schema_version.as_deref(), + Some(OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION) + ); +} + +#[test] +fn compact_oversized_artifact_counts_fail_closed() { + let edge_count = MAX_EVIDENCE_UNITS as u64 + 1; + let outcome_of_count = edge_count - 2; + let artifact = OutcomeOrderArtifact { + schema_version: OUTCOME_ORDER_ARTIFACT_SCHEMA_VERSION.into(), + run_id: "run-compact-oversize".into(), + snapshot_id: "snapshot-compact-oversize".into(), + knowledge_cutoff: "2026-08-01T00:00:00Z".into(), + edge_count, + input_to_count: 1, + process_to_count: 1, + outcome_of_count, + refused_as_transition_count: outcome_of_count, + inference_status: "input_process_forward_outcome_of_is_not_transition".into(), + }; + let raw_payload = serde_json::to_string(&artifact).expect("raw json"); + assert_eq!( + artifact.to_json(), + Err(AnalysisEngineError::InvalidOutcomeOrderArtifact) + ); + assert_eq!( + OutcomeOrderArtifact::from_json(&raw_payload), + Err(AnalysisEngineError::InvalidOutcomeOrderArtifact) + ); +} + +#[test] +fn future_available_edges_are_excluded() { + let request = request(); + let mut with_future = mixed_edges(); + with_future.push(edge( + "future-d", + OutcomeKind::InputTo, + 4, + 5, + "2026-08-02T00:00:00Z", + )); + let execution = execute(&request, &with_future).expect("cutoff"); + assert_eq!(execution.artifact.edge_count, 3); + assert_eq!(execution.artifact.input_to_count, 1); +} + +#[test] +fn empty_or_single_class_reverse_uncertain_and_duplicate_fail_closed() { + let request = request(); + let stamp = "2026-07-01T00:00:00Z"; + assert_eq!( + execute(&request, &[]), + Err(AnalysisEngineError::InvalidEvidence) + ); + let input_only = vec![ + edge("input-a", OutcomeKind::InputTo, 1, 2, stamp), + edge("input-b", OutcomeKind::InputTo, 3, 4, stamp), + edge("input-c", OutcomeKind::InputTo, 5, 6, stamp), + ]; + assert_eq!( + execute(&request, &input_only), + Err(AnalysisEngineError::InvalidEvidence) + ); + let transitions_only = vec![ + edge("input-a", OutcomeKind::InputTo, 1, 2, stamp), + edge("process-b", OutcomeKind::ProcessTo, 2, 3, stamp), + ]; + assert_eq!( + execute(&request, &transitions_only), + Err(AnalysisEngineError::InvalidEvidence) + ); + let provenance_only = vec![ + edge("outcome-a", OutcomeKind::OutcomeOf, 9, 1, stamp), + edge("outcome-b", OutcomeKind::OutcomeOf, 8, 2, stamp), + edge("outcome-c", OutcomeKind::OutcomeOf, 7, 3, stamp), + ]; + assert_eq!( + execute(&request, &provenance_only), + Err(AnalysisEngineError::InvalidEvidence) + ); + let reverse = vec![ + edge("input-a", OutcomeKind::InputTo, 4, 1, stamp), + edge("process-b", OutcomeKind::ProcessTo, 2, 3, stamp), + edge("outcome-c", OutcomeKind::OutcomeOf, 9, 1, stamp), + ]; + assert_eq!( + execute(&request, &reverse), + Err(AnalysisEngineError::InvalidEvidence) + ); + let contemporaneous = vec![ + edge("input-a", OutcomeKind::InputTo, 2, 2, stamp), + edge("process-b", OutcomeKind::ProcessTo, 2, 3, stamp), + edge("outcome-c", OutcomeKind::OutcomeOf, 9, 1, stamp), + ]; + assert_eq!( + execute(&request, &contemporaneous), + Err(AnalysisEngineError::InvalidEvidence) + ); + let duplicates = vec![ + edge("same", OutcomeKind::InputTo, 1, 2, stamp), + edge("same", OutcomeKind::OutcomeOf, 9, 1, stamp), + ]; + assert_eq!( + execute(&request, &duplicates), + Err(AnalysisEngineError::DuplicateEvidence) + ); + assert_eq!( + OutcomeOrderEdge::new("", OutcomeKind::InputTo, 1, 2, available(stamp)), + Err(AnalysisEngineError::InvalidEvidence) + ); +} + +#[test] +fn execution_refuses_snapshot_profile_cutoff_mismatch_and_oversize() { + let request = request(); + let edges = mixed_edges(); + assert_eq!( + execute_outcome_order_run( + &request, + &accepted(&request), + "other-snapshot", + cutoff(), + &edges, + "2026-08-02T00:00:00Z", + ), + Err(AnalysisEngineError::SnapshotMismatch) + ); + let mut mismatched = request.clone(); + mismatched.knowledge_cutoff = "2026-07-01T00:00:00Z".into(); + assert_eq!( + execute_outcome_order_run( + &mismatched, + &accepted(&mismatched), + "snapshot-outcome-order", + cutoff(), + &edges, + "2026-08-02T00:00:00Z", + ), + Err(AnalysisEngineError::InvalidEvidence) + ); + for profile in [ + "trsl_topic_lineage_v1", + "fitted_candidate_k_v1", + "pareto_candidate_k_v1", + "joint_posterior_draws_v1", + "method_effects_v1", + "copy_identity_v1", + "style_source_v1", + "prompt_source_v1", + "modality_source_v1", + "corpus_background_v1", + "citation_edge_v1", + "copied_text_v1", + "lineage_criterion_v1", + "composed_fitted_lineage_v1", + "case_deletion_refit_v1", + "topic_activity_v1", + "location_membership_v1", + "topic_context_posterior_v1", + "membership_posterior_icc_v1", + "membership_target_v1", + ] { + let mut reused = request.clone(); + reused.output_profile = profile.into(); + assert_eq!( + execute_outcome_order_run( + &reused, + &accepted(&reused), + "snapshot-outcome-order", + cutoff(), + &edges, + "2026-08-02T00:00:00Z", + ), + Err(AnalysisEngineError::InvalidEvidence) + ); + } + let oversized: Vec = (0..=MAX_EVIDENCE_UNITS) + .map(|index| { + edge( + &format!("edge-{index}"), + OutcomeKind::InputTo, + u64::try_from(index).expect("index"), + u64::try_from(index).expect("index") + 1, + "2026-07-01T00:00:00Z", + ) + }) + .collect(); + assert_eq!( + execute(&request, &oversized), + Err(AnalysisEngineError::LimitExceeded) + ); +} diff --git a/docs/TRACEABILITY.md b/docs/TRACEABILITY.md index 2b783c2ab..4103bfc38 100644 --- a/docs/TRACEABILITY.md +++ b/docs/TRACEABILITY.md @@ -58,6 +58,7 @@ The full APA 7th standards/literature register remains `docs/research/standards- | versioned service/API contracts and exports | PRD; API contract; ADR 0011/0013 | `tepp_api` analysis-run/export/JSON-LD/GraphML contracts on protected main (PR #21); HTTP service remaining accepted-target | partial | | versioned service/API contracts and exports | PRD; API contract; ADR 0011/0013 | `tepp_api` analysis-run/export/JSON-LD/GraphML contracts on protected main (PR #21); LineageWeave loopback contracts and request-bound terminal result are composed on the active product branch; production TLS remaining | partial | | executable cutoff-safe analysis runs | ADR 0012/0022; temporal research; API terminal-result contract | `analysis_engine` availability cutoff, snapshot binding, multiple-membership aggregation, digest-bound readiness artifact, and `tepp.trsl_topic_lineage.v1` execution through `topic_measurement`; synthetic recovery plus tamper/non-convergence tests and exact coverage on the active product branch | active-PR | +| outcome-order analysis-run profile | ADR 0002/0003/0022/0070; `input_to`/`process_to` stay forward; `outcome_of` is not a transition | `analysis_engine` `outcome_order_v1` binds `refuse_reverse_ipo_order` and `refuse_outcome_of_as_transition`; digest-bound refusals, not `kind_recovery_rate` inspect metric, not membership-target, not location-membership, not GPU, not MCMC, not topic birth/split/merge; not implemented-main | active-PR | | immutable split/run/reproducibility manifests | ADR 0013; ERD | `tepp_api` reproducibility manifest contract on protected main; `persistence_postgres` append-only SQL insert/lookup for `reproducibility_manifest`, `corpus_split_manifest`, `model_run`, and `model_artifact` (migration `0003`); full physical ERD constraints remaining | partial | | multilingual shared latent semantic space | PRD; ADR 0004; ADR 0020 | `semantic_core` span-grounded units (active-PR); concept dictionary and shared latent estimator remaining | active-PR | | TRSL-TM temporal/relational topic posterior and backend compatibility | ADR 0012; ADR 0004 | `topic_measurement` stable ALR/ILR coordinates and bounded CPU `f64` reference estimator on protected main; `model_selection` fitted candidate-`K` scoring on this PR; calibrated posterior promotion, method effects, persistence, and accelerated backends remaining | partial | diff --git a/docs/adr/0070-outcome-order-analysis-run.md b/docs/adr/0070-outcome-order-analysis-run.md new file mode 100644 index 000000000..ffba0e4c3 --- /dev/null +++ b/docs/adr/0070-outcome-order-analysis-run.md @@ -0,0 +1,98 @@ +# ADR 0070 — Input-process-outcome order refusals as an analysis-run output profile + +**Decision status:** Accepted +**Implementation maturity:** active-PR — composed on this branch; not implemented-main +**Date:** 2026-09-01 +**Supersedes:** None; complements ADR 0002 (forward IPO transitions never move backward in event time; `outcome_of` is provenance) and ADR 0022 (cutoff-safe analysis-run execution). Does not reuse ADR 0069 (membership-target), ADR 0068 (topic-context posterior), ADR 0066 (location-membership), ADR 0065 (copied-text residue), ADR 0064 (provenance-is-not-transition / citation-edge), or ADR 0058 (copy-identity / template-copy). +**Figma File ID:** N/A — this increment changes a Rust service crate and has no user-interface surface. +**Storybook inventory:** N/A — no reusable web object or interaction changed. + +## Context + +Protected main already refuses reverse or contemporaneous event-time rank on +`input_to` and `process_to`, and refuses to treat `outcome_of` as a state +transition, via `outcome_order::OutcomeKind`, `refuse_reverse_ipo_order`, and +`refuse_outcome_of_as_transition`. Operators still cannot request that IPO +census as a digest-bound analysis-run output. + +Membership-target (#434 / ADR 0069) binds `MembershipTargetKind`. +Location-membership (#430 / ADR 0066) binds geographic/market assignment. +Copied-text (#427 / ADR 0065) binds unique-content/stopword vocabulary. +Copy-identity (#416 / ADR 0058) binds template-copy identity. +Provenance-is-not-transition (#426 / ADR 0064) binds citation-edge, not +`outcome_of`. + +`kind_recovery_rate` stays library-side. This slice does not put a +`scientific_acceptance` metric on inspect payloads. + +GPU kernels, MCMC, and topic birth/split/merge remain later GAP-004 work +and are not this slice. + +## Decision + +Add the `outcome_order_v1` analysis-run output profile to +`analysis_engine`. The executor: + +- consumes already-validated `OutcomeOrderEdge` rows with closed + `OutcomeKind` values, opaque event-time ranks, and availability time; +- requires the request snapshot and knowledge cutoff to match the offered + input construction; +- excludes edges whose availability is later than the knowledge cutoff; +- invokes `refuse_reverse_ipo_order` and + `refuse_outcome_of_as_transition` without reimplementing the + `input_to` / `process_to` / `outcome_of` vocabulary; +- requires a mixed census of at least one `input_to`, one `process_to`, + and one `outcome_of` after cutoff exclusion; +- emits a canonical SHA-256-digested `tepp.outcome_order.v1` artifact + with per-kind counts, matching `outcome_of` refusal counts, and + inference status `input_process_forward_outcome_of_is_not_transition`; +- does not emit `kind_recovery_rate`, invent MCMC, select GPU backends, + or emit topic birth/split/merge events. + +## Alternatives considered + +1. Duplicate membership-target (#434 / ADR 0069) — rejected because that + profile binds `MembershipTargetKind`, not IPO event-time order. +2. Duplicate location-membership (#430 / ADR 0066) — rejected because + that profile binds `location_membership` LocationKind refusals. +3. Duplicate copied-text (#427) or copy-identity (#416) — rejected + because those profiles bind residue/template identity, not IPO kinds. +4. Duplicate citation-edge (#426 / ADR 0064) — rejected because that + profile binds citation provenance, not `outcome_of`. +5. Put `kind_recovery_rate` on the operator artifact — rejected because + inspect payloads stay metric-free and `tepp.scientific_acceptance.v1` + never appears. +6. Bind the existing outcome-order refusals to ADR 0022's analysis-run + profile — accepted. + +## Consequences + +Operators can request cutoff-safe IPO-order refusals as a digest-bound +terminal result. The artifact does not claim MCMC, GPU parity, +membership-target, location-membership, membership-posterior ICC, +copied-text, copy-identity, citation-edge, corpus-background, +method-effect estimation, or topic birth/split/merge. Snapshot / profile +/ cutoff mismatch, empty or single-class corpora, reverse or uncertain +IPO order, duplicate edge identities, and oversized corpora fail closed. + +## Verification + +The PR includes Rust unit and integration tests for mixed +`input_to` / `process_to` / `outcome_of` corpora, cutoff exclusion, +empty/single-class/duplicate/reverse/uncertain refusal, snapshot / +profile / cutoff mismatch, oversize, and artifact tampering. Run: + +```text +cargo fmt --all -- --check +cargo test -p analysis_engine +cargo clippy -p analysis_engine --all-targets -- -D warnings +python3 scripts/validate_documentation.py +``` + +## Rollback and supersession + +Rollback removes the `outcome_order_v1` profile. No persisted schema +migration is introduced. Supersede only with an ADR that keeps `input_to` +and `process_to` strictly forward in event-time rank, keeps `outcome_of` +out of the transition vocabulary, and keeps `kind_recovery_rate` off +inspect payloads. diff --git a/docs/adr/README.md b/docs/adr/README.md index 1254c8079..f19e60ebb 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -30,6 +30,7 @@ Read [`ADR_POLICY.md`](ADR_POLICY.md) first. **Decision status and implementatio | [0022](0022-deterministic-analysis-run-execution.md) | Deterministic cutoff-safe analysis-run execution | Accepted | active-PR | Closes the first executable product path from accepted run to digest-bound terminal result without claiming estimator authority. | | [0024](0024-lineage-pair-criterion-and-project-journey-posterior.md) | Independent Event Lineage pair criterion and posterior Project Journey | Proposed | active-PR | Strict artifacts preserve criterion/event-time draws, branches, ties, and CPU/GPU receipts without claiming the scientific estimator is complete. | | [0025](0025-macos-native-rust-mlx-metal-boundary.md) | macOS-native Rust-owned MLX Metal execution | Accepted | accepted-target | Compose authenticates to a native host service; Linux never claims Metal, and actual backend/parity receipts fail closed. | +| [0070](0070-outcome-order-analysis-run.md) | Input-process-outcome order refusals as an analysis-run profile | Accepted | active-PR | Complements ADR 0002/0003/0022; `OutcomeKind` + `refuse_reverse_ipo_order` + `refuse_outcome_of_as_transition`, not membership-target, not location-membership. | | [0023](0023-lineage-criterion-anchor-contract.md) | TEPP-owned Event Lineage criterion anchor | Accepted | active-PR | PR #237 publishes the strict accepted/rejected artifact and identities; estimator execution remains fail-closed future work. | | [0024](0024-independent-topic-importance-anchor.md) | Posterior topic-context producer contract | Accepted | contract-only active-PR | Strict DTO/schema only; the current estimator does not emit it. fast-mlsirm owns case-deletion influence. | | [0001](0001-rust-first-modular-msa.md) | Rust-first numerical core and CPU `f64` reference | Accepted | partial | ADR 0011 owns cross-service/MSA authority; 0001 retains numerical/backend authority. | @@ -138,6 +139,7 @@ Use the narrowest owning ADR when decisions overlap: - **project-history wire-size symmetry:** ADR 0019. - **LineageWeave project-history service boundary:** ADR 0021. - **accepted-run execution and terminal artifact production:** ADR 0022. +- **outcome-order analysis-run profile:** ADR 0070. - **independent lineage criterion and posterior Project Journey:** ADR 0023. - **macOS-native Rust-owned MLX Metal execution:** ADR 0024. diff --git a/docs/doctoring/outcome-order-analysis-run.md b/docs/doctoring/outcome-order-analysis-run.md new file mode 100644 index 000000000..874d2f436 --- /dev/null +++ b/docs/doctoring/outcome-order-analysis-run.md @@ -0,0 +1,18 @@ +# Input-process-outcome analysis-run composition + +**Active slice:** ADR 0070 / `outcome_order_v1` +**Protected-main status:** not implemented-main + +`outcome_order` already refuses reverse or contemporaneous event-time +rank on `input_to` and `process_to`, and refuses to treat `outcome_of` +as a state transition. This slice binds `OutcomeKind`, +`refuse_reverse_ipo_order`, and `refuse_outcome_of_as_transition` to a +cutoff-safe analysis-run profile so operators can request a digest-bound +identity artifact. + +The artifact inference status is +`input_process_forward_outcome_of_is_not_transition`. +`kind_recovery_rate` stays library-side. This is not membership-target, +not location-membership, not membership-posterior ICC, not copied-text, +not copy-identity, not citation-edge, not GPU, not MCMC, and not topic +birth/split/merge.