Open Problems
Identified gaps in the current program. A collaborator who fills any of these makes a direct contribution to the architecture. For what the program covers and what it does not — the chain marked stage by stage — see Coverage; that map is what makes this list legible. To find where each gap sits in the corpus, start fromthe Guide and widen via the corpus map.
| # | Gap | Layers | Priority | Notes |
|---|---|---|---|---|
| 1 | Human-subject empirical validation of cohort divergence | L5 | HIGH | All current evidence is LLM-mediated. Instrument sensitivity is confirmed (PRISM-B, Run 15); primacy effects are bounded as model-specific (F1, GPT-only) and domain-specific (F2, brand-only), and Run 15b isolated a JSON-format primacy effect (η² = .217) at the elicitation layer, resolved in sbt-framework v2.3.1 via Latin-square dimension ordering. So the instrument side is in hand and the subject side is not: a coverage sweep of the whole corpus finds no human behavioural data in it anywhere, and the furthest executed reach down the causal chain is the behaviour of a model rather than of a person. A human conjoint or MaxDiff study with elicited dimensional weights is the next milestone. |
| 2 | Long-horizon longitudinal tracking | L2L5 | MEDIUM | H13 closed short-horizon stability across 4 model pairs (cosines > .97), and 2026ba separates apparatus drift from brand signal, which is the precondition for reading any multi-epoch series. Open: tracking the same cohort across 6+ months and through multiple disruption events, to test whether the μ > λ inequality predicts in-vivo trajectories rather than simulated ones. |
| 3 | Real-world agentic deployment | L3L4 | MEDIUM | Exps A/D/Q1 demonstrated compounding and showed constraint framing reduces variance 62% (Q1), and PRISM-C (2026bb) now measures a revealed pick by an LLM agent rather than a stated one. That closes the simulated half. Open: field deployment in live agent workflows with human revealed-purchase outcomes — the agent choosing is still a model, and a model choosing is not a market. |
| 4 | Collapse onset and early-warning indicators | L2L5 | HIGH | R22 (2026ad) closed the recovery side: μ > λ at scale δ restores separability. The symmetric onset problem is open — which observable signals precede a spectral collapse, how early the gap-decay rate can be estimated, and what practitioner-facing lead time exists before separability is lost. |
| 5 | Cohort discovery from raw observation | L1L5 | MEDIUM | Most papers stipulate cohorts (priors, demographics, weight vectors). Open: unsupervised identification of latent cohorts from observation streams without pre-specified weights, and the conditions under which discovered cohorts coincide with the alibi-style invariant structure. |
| 6 | Formal cross-domain operator identification | L4 | MEDIUM | R22 and its OST companion cite independent convergence in capital-markets and DeFi composability work showing the same threshold-inequality and projection-operator structure. Open: a formal identification paper showing that brand perception, organizational verification and composability share one operator-theoretic structure under a specified mapping — as a proof, not as a resemblance. |
| 7 | Causal identification beyond observational designs | L5 | MEDIUM | Revised: it is no longer true that all empirical work here is observational and LLM-mediated. 2026bi is a large-N confirmatory field study on 350 completed transactions, and several LLM studies are randomised and pre-registered. What is still missing is on the perception side and with human observers: a quasi-experimental design — regression discontinuity around a brand event, an instrument, a natural experiment — that identifies the perception-shift effect of a specified disruption rather than its correlation. |
| 8 | Multi-shock and cascade dynamics | L2L3 | LOW | R22 models a single coherence shock. Open: interaction effects of sequenced shocks, simultaneous shocks across cohorts, and contagion across linked brands in a portfolio. The R21 portfolio-immunity result suggests the cascade structure differs for AI and human observers. |
| 9 | Specification-to-perception coupling | L0L3 | MEDIUM | SBT v3.2.0 §5.2.1 formalised the DO/WHAT bridge between organizational specification and observable dimensions. Open: an empirical study mapping a documented specification onto its measured perception cloud and quantifying the coupling strength. Distinct from #13 — this asks whether a specification shows up in how a brand is read, not whether it shows up in what the firm earns. |
| 10 | Capstone synthesis | L6 | PREMATURE | Unified theory review across L0–L5. Premature until human-subject validation (#1) and a longitudinal field result (#2) are in hand; #12 and #13 have since joined that dependency list. |
| 11 | Measurement invariance across observer classes (human vs LLM) | L1L5 | HIGH | Corrected: the cross-architecture agreement reported in R15 (2026v) is not evidence for invariance here, and the paper has since been sharpened to say so. Agreement across architectures is agreement across OBJECTS being read; the open question is agreement across INDIVIDUALS doing the reading, and an architecture-invariance result does not answer an elicitation objection. So the question stands undiminished: does the eight-dimension instrument exhibit configural, metric and scalar invariance across observer classes, or is the LLM aperture a distinct measurement regime? Establishing it, or characterising its bounds, needs calibration against a fixed human reference panel plus convergent-validity tests. Absent that, any cross-observer-class comparison of the same brand rests on an unverified equivalence assumption. Distinct from #1: that asks whether a finding replicates with humans, this asks whether the instrument measures the same construct at all. |
| 12 | Field measurement of the emitted-signal → perception route | L2L3L5 | HIGH | Added after the coverage sweep, and it is the gap the corpus’s own chain claim rests on. Four papers formalise the route from an emitted signal to a perceptual state, bypassing impressions and attention: a campaign as a forcing vector on a forced Ornstein-Uhlenbeck process with a recoverable decay constant, a resonance over-index, a coherence-emission threshold inequality, and order dependence with absorbing states. Every one is theory plus a calibrated demonstration, not one is a field measurement, and each paper says so in its own text. The numbers they carry are model properties rather than measured advertising response. Open: any field measurement of that route. |
| 13 | Field evidence that specification quality moves firm performance | L0L4 | HIGH | Added after the coverage sweep, and stated plainly because the honest version is the useful one: there is no executed study anywhere in this corpus with specification quality as an independent variable and firm performance as a dependent variable. Three near-misses each fall short differently, and none may stand in for the claim. 2026an builds the measure and the full identification template, but the archival estimates have never been run — what exists is a Monte Carlo and a powered placebo that bounds the measure (d = .166, p = .243). 2026bi has real field data on 350 transactions, but tests whether a closing-time gap is NECESSARY for one integration-failure mode: consistency .517, floor .073 with an exact upper bound of .112 — a bounded negative screen, never a predictor of success. 2026aq is the cleanest specification-pays result (d = .314, p = .009; d = .569, p < .001), but its subject is a dyad of language-model agents rather than a firm. |