0 cumulative citations
View corpus contextEngineering archives can give physical-world AI a running start: extracting 'constitutive priors' from design documents lets systems operate from day one and produce the data needed to refine them. The authors formalize a four-world ontology, a legitimacy test for such priors, and a necessary four-layer architecture, and propose five public, falsifiable predictions to validate the framework.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed, and the exception has a name: the artificial physical world. Buildings, industrial facilities, and infrastructure are intentionally constituted and documented: designed artifacts ship with readable archives that precede and constitute their instances; here, norms are promulgated before instances, not averaged from them. Four contributions. (i) From a four-world ontology we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit -- deviation from a constitutive norm is a violation in the world, not a revision of the model. (ii) We establish a layering lower bound: any such framework has at least four layers -- syntax, concept, knowledge, instance -- because four construction goals pair into mutually incompatible carriers. (iii) We register deployment claims across five industrial domains and a 32-class failure-mode vocabulary. (iv) We stake the framework on five falsifiable predictions, the central one checkable on the public engineering record: if it fails, the framework fails. Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and a boundary theorem for certificate-anchored calculi. Large language models find an honored place here -- as readers of the archive, not as the archive. First of three companion works; the companions take up the questions deliberately left open.
Summary
Main Finding
Machine intelligence can legitimately be given a prior (a “constitutive prior”) before large-scale observational data only for a specific subset of the physical world: the artificial physical world (designed artifacts such as buildings, industrial facilities, and infrastructure). Legitimacy requires two conjuncts — (1) the domain is intentionally constituted (design rules exist prior to instances) and (2) those constitutive processes left a readable generative archive (drawings, manuals, certificates). When both hold, an extracted prior binds instances by promulgated norm (failures are violations), enabling immediate, deployable intelligence that then generates data and creates a self-reinforcing data–model flywheel. Such constitutive-prior frameworks must be layered (minimum four layers: syntax, concept, knowledge, instance). The paper formalizes this via a four-world task-relative ontology (phenomenal, basic physical, artificial physical, artificial symbolic), gives semi-formal boundaries and decidability results, enumerates deployment claims and failure modes, and stakes five falsifiable industry-level predictions.
Key Points
- Cold-start deadlock in physical AI: no intelligence without data, no data without deployed intelligence. This deadlock is structural and unevenly distributed across “worlds.”
- Four-world task-relative ontology (derived from three learner-oriented variables: data interface, constraint structure, design rules):
- Artificial symbolic (text, code)
- Artificial physical (designed artifacts with archives)
- Basic physical (physical systems without design archives)
- Phenomenal (raw sensory appearances)
- Legitimacy criterion for extracting priors: legitimate iff (i) the domain is intentionally constituted and (ii) the constitutive process left a readable generative archive. Testable by “direction of fit”: deviation from constitutive norms is a world violation, not only a model update.
- Layering necessity: to meet competing goals (stability, openness, abstraction, concreteness) a prior framework needs at least four layers — syntax, concept, knowledge, instance — because paired goals require incompatible carriers.
- Practical architecture role for LLMs: useful as archive readers (extractors) but not as substitutes for promulgated archives or for the ontology/concept layers.
- Comparative claim: constitutive-prior route differs from purely data-driven, world-model, digital-twin, and ontology-only approaches by enabling intelligence before deployment, low historical-data needs, capturing normative violations, and enabling high cross-instance transfer in designed domains.
- Falsifiable program: five explicit predictions (e.g., continued AI breakthroughs following the worlds order; cross-industry invariance of the concept layer; limits on VLA-only entry into institutions; necessity of ontology layer alongside LLMs). If key predictions fail, framework is rejected.
- Semi-formal technical contributions (in appendices): Gold-style boundary on rule coverage in archiveless worlds, decidability of failure reduction over closed concept layers, explicit boundary theorem for certificate-anchored calculi.
- Practical deliverables: 32-class enumerative vocabulary of failure modes; deployment claims across five industrial domains (existence proof style: claims + public commitments rather than proprietary data release).
Data & Methods
- Nature of work: conceptual/theoretical framework supported by semi-formal arguments and theorems; not an empirical dataset paper.
- Derivation method:
- Start from three learner-facing variables (data interface, constraint strength, design rules).
- Derive the four-world partition as a task-relative instrument for what a learner faces.
- Formulate the legitimacy criterion based on intentional constitution and readable archives; argue sufficiency via direction-of-fit reasoning.
- Prove a layering lower bound (≥4 layers) using incompatibility of paired design goals; sketch a minimal skeleton and enumerate failure modes.
- Formal support: Appendices contain semi-formal proofs (Gold-type coverage limits, decidability results, boundary theorem for certificate-anchored calculi).
- Empirical stance: deployment evidence is presented as “claims + commitments” (publicly testable statements and procedures are specified in appendices) rather than revealing proprietary operational data. The framework is explicitly falsifiable via the five predictions and public engineering records (e.g., test the cross-domain invariance of the concept layer on public engineering documents).
- Role for LLMs and other ML: positioned as tools (archive readers, NLU) within the architecture, not as replacements for the promulgated prior.
Implications for AI Economics
- New investable opportunity: domains with documented, machine-readable archives (buildings, power plants, aviation, manufacturing lines, infrastructure) become high-ROI targets for AI deployment because constitutive priors permit intelligence before accumulating operational data, enabling immediate value capture and faster data flywheels.
- Shift in cost structure: value and costs move upstream — from expensive in-field data collection (especially failure data) to archive curation, digitization, and extraction capability. Upfront investment in archive digitization and standardization can reduce downstream operating and data-acquisition costs.
- Product-market implications:
- Companies that can extract and package constitutive priors (ontology/concept layers) become strategic providers to institutions; LLM-based services will be valuable as archive readers but insufficient alone to meet institutional normative requirements.
- Cross-instance transfer becomes economically feasible (scale effects) within design-governed industries; platform economics favor firms that standardize concept layers across clients.
- Labor and skills: demand rises for archive digitization specialists, ontology engineers, verification/certification tooling, domain knowledge engineers who can translate design archives into machine-interpretable constitutive priors.
- Competitive and regulatory dynamics:
- Incumbents owning detailed archives (engineering firms, asset owners, certificate authorities) may extract rents or exert market power. Access models (licensing, federated reading, regulated disclosure) become central economic and policy questions.
- Promulgated norms and certificate-anchored calculi open new regulatory levers: regulators can require machine-interpretable archives for certification, or mandate provenance standards to reduce asymmetric information and liability.
- Risk & limits:
- Not applicable to archiveless domains (basic physical or phenomenal worlds). Attempts to generalize constitutive priors where archives are absent will fail (Gold-style coverage limit).
- Institutional adoption ceiling: very-large-autoregressive (VLA) models alone will hit limits in institutional settings; ontology/concept layers are economically necessary complements.
- Actionable recommendations for economists, investors, and policy makers:
- Prioritize investment and due diligence on firms and assets where design archives are complete and machine-readable or readily digitizable.
- Fund standards and tooling for archive digitization, schema interoperability, and certificate-machine interfaces (to unlock cross-instance transfer and reduce lock-in).
- Assess firms on two axes: (a) quality & accessibility of constitutive archives and (b) capability to produce and maintain layered priors (syntax/concept/knowledge/instance).
- Monitor the paper’s falsifiable predictions (especially the cross-industry invariance of the concept layer) as empirical signals to validate the framework and to inform allocation of capital.
- Anticipate regulatory interventions around archive access, certification, and liability allocation; design business models that account for potential rent-extraction by archive owners.
Overall economic take: extracting constitutive priors from designed domains promises to unlock persistent value by enabling immediate deployable intelligence, lowering marginal data costs, and creating durable, transferrable assets (layered priors) — but that promise depends on archive availability, standards, and institutional arrangements that shape who captures the gains.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Physical AI faces a cold-start deadlock: intelligence cannot be trained without data, while data is not produced without deployed intelligence. Automation Exposure | negative | Availability of training data and deployable physical-world intelligence |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Buildings, industrial facilities, and infrastructure are intentionally constituted and documented physical domains in which design archives precede and constitute instances. Automation Exposure | positive | Availability of prior design knowledge for physical-world AI deployment |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Prior extraction is legitimate if and only if the object domain is intentionally constituted and the constitutive process has left a readable generative archive. Governance And Regulation | positive | Legitimacy of extracting a prior framework for machine learning |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| When the legitimacy criterion holds, deviations from the constitutive framework should be treated as violations of a promulgated norm rather than as evidence requiring revision of the model. Regulatory Compliance | positive | Detection and interpretation of deviations from prescribed physical-world norms |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| A constitutive-prior framework must contain at least four layers: syntax, concept, knowledge, and instance. Other | positive | Structural completeness of a prior framework |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Constitutive priors are predicted to provide intelligence before deployment, require relatively little historical data, capture normative violations, and support high cross-instance transfer in physical-world applications. Automation Exposure | positive | Pre-deployment intelligence, historical-data dependence, normative-violation detection, and cross-instance transfer |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The framework predicts that the concept layer will be invariant across industrial domains. Organizational Efficiency | positive | Cross-industry invariance of the framework’s concept layer |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Large language models should function as readers of design archives rather than as the archive or the source of constitutive norms. Task Allocation | positive | Role of large language models in physical-world intelligence systems |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The sequence of AI breakthroughs is predicted to correlate with constraint strength and the quality of the available data interface, with symbolic domains such as code preceding robotics. Adoption Rate | positive | Order and timing of AI domain breakthroughs and commercial deployment |
Reading fidelity
high
Study strength
low
|
not reported
|