Many organizational AI failures stem from missing representations of 'when not to follow the rule'; O‑I‑B‑A‑R formalises how to capture exceptions, runtime unknowns, and who must decide, making human–AI handoffs explicit — but the proposal remains conceptual and untested.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
AI systems increasingly enter organizations through policies, procedures, playbooks, prompts, and other explicit representations of work. Yet formal descriptions often differ from situated practice, and captured know-what can omit the contextual know-how experts use when judgments are uncertain. We argue that a recurring class of organizational AI failures arises partly from a knowledge representation problem at the sociotechnical interface: the AI receives the procedure, while the organization operates on the procedure plus negative boundaries, runtime judgments, responsibility assignments, and learning history. We introduce O-I-B-A-R (OPEN, IS, BUT, ACTION, RESULT), a scaffold for externalizing these missing decision boundaries. IS records when a judgment holds. BUT records a concrete failure containing information beyond the logical negation of IS. Comparable success and failure cases are decomposed toward a minimally sufficient changing variable, which becomes a value-bearing decision dimension. A suspension represents the state in which the dimension is known but its current value is unresolved, specifying what must be measured, asked, retrieved, or escalated to a human. RESULT confirms a boundary, shifts a threshold, or exposes a new dimension. Incidents can generate new dimensions, unresolved values can define human-AI handoffs, and feedback can expand the decision space. We also identify a sociotechnical tension: durable and attributable failure histories can suppress the candor on which useful boundary knowledge depends. Externalization must therefore be designed as an organizational intervention with real costs and incentives.
Summary
Main Finding
The paper diagnoses a representational problem at the sociotechnical interface that helps explain recurring organizational AI failures: organizations typically supply AI with positive procedures (what to do), but not with the negative decision boundaries, runtime judgments, responsibility assignments, and failure histories that practitioners use. It proposes O-I-B-A-R (OPEN, IS, BUT, ACTION, RESULT) — a lightweight, contrastive scaffold to externalize decision-boundary knowledge so that AI systems (and organizational processes) can distinguish when a rule applies, what variable actually determines a failure, what must be known at runtime, who must resolve it, and how to update the model after action. The framework also highlights a sociotechnical tension: durable, attributable failure records can suppress candid reporting of boundary knowledge, so externalization must be designed with incentives and organizational protections.
Key Points
- Four representational gaps that produce “lack of judgment” errors:
- Boundary gap: explicit positive rules (IS) but tacit negative boundaries (what makes the rule fail).
- Runtime-value gap: a decision-relevant dimension may be known but its value for the current case is unresolved.
- Responsibility gap: “human in the loop” is underspecified — which variable/value should the human supply or judge?
- Learning gap: organization-level updates often become action-level guardrails instead of expanding or refining the decision space.
- O-I-B-A-R components:
- OPEN: a concrete judgment J within a domain Ω to be tested.
- IS: the positive applicability condition (when J is justified).
- BUT: a concrete failure case; critically NOT merely ¬IS — it should introduce decision-relevant information beyond the IS statement.
- Decision-dimension discovery: compare IS and BUT to isolate minimally sufficient changing variable d (aim for a single variable |∆| = 1).
- Suspension: explicit representation of an unresolved value for known dimension d (qd(x) = ?), making the missing information actionable and identifying human/automation roles.
- ACTION: the committed decision given current model/assumptions.
- RESULT: structured feedback that either confirms the boundary, refines thresholds, or requires adding a new dimension.
- Key theoretical moves:
- Treat IS and BUT asymmetrically — BUTs are generative and can reveal previously omitted decision variables.
- Use domain-relative naming and iterative abstraction (AΩ operator) to produce stable, value-bearing dimension names.
- Represent suspension as a primitive human–AI handoff — who must resolve what.
- Distinguish boundary refinement (threshold moves) from structural expansion (new dimension).
- Sociotechnical caveat: failure externalization carries costs (blame, exposure); organizations may withhold useful failure cases unless incentives, psychological safety, or protections are provided.
Data & Methods
- This is a conceptual, theoretical paper combining:
- Literature synthesis across organizational studies (Suchman; Brown & Duguid; Orr; Feldman & Pentland), cognitive/task elicitation methods (repertory-grid, cognitive task analysis), machine-learning ideas (near-miss learning, version spaces, counterexample-guided abstraction refinement), and organizational learning/behavioral literatures (Argyris & Schön; organizational silence; psychological safety).
- Formalized, operational definitions and heuristics: IS : S ⇒ J; heuristic V(BUT) \ V(IS) ≠ ∅ to screen generative failures; ∆(IS, BUT | Ω) to define minimally sufficient changing variables; representation of suspension queries qd(x).
- Worked examples and conceptual thought experiments illustrating decomposition from action-level failures to decision-dimension attributions (e.g., “empathy was wrong” → “need type misclassified”).
- A proposed loop (OPEN → (IS,BUT) → D → qD → ACTION → RESULT → OPEN′) and notions of local structural stability and “bounded exposure-rate certificate” (warning that absence of new-dimension discoveries over a case stream is not proof of completeness due to rare events).
- No new large-scale empirical dataset or randomized trial is presented. The paper situates O-I-B-A-R relative to existing elicitation and learning methods and lays out operational criteria amenable to empirical implementation and testing.
Implications for AI Economics
- Deployment costs and investment trade-offs:
- Externalizing decision boundaries (collecting BUTs, building suspensions, mapping responsibilities, and maintaining RESULT-driven updates) requires upfront time, coordination, and incentives — raising short-term deployment costs but reducing costly brittle errors long-term.
- Firms that invest in structured externalization may realize higher effective returns to AI by reducing loss from misapplied rules, lowering rework/escalation costs, and accelerating learning about rare-but-important failure modes.
- Changes in the division of labor and task complementarity:
- Representing suspensions clarifies which parts of decision-making remain non-reducible and hence preserves or creates human roles focused on interpretation/ontological judgment rather than rote execution.
- Demand for worker skills will shift toward resolving suspensions, naming decision dimensions, and maintaining boundary knowledge (affecting skill premia and retraining needs).
- Market design and productization:
- AI vendors and platforms can differentiate by embedding O-I-B-A-R primitives: interfaces for IS/BUT capture, suspension types, human-hand-off workflows, and RESULT-driven model updates. This becomes a product feature, potentially increasing switching costs and platform lock-in.
- Predictable human–AI handoffs lower coordination costs and make contractible AI services easier to price and insure.
- Measurement, evaluation, and incentives:
- Standard ML metrics (accuracy/recall) are insufficient. Economic evaluation should include coverage of decision-boundary space, frequency of suspensions, costs of unresolved suspensions, and the nature of RESULT-driven structural expansions.
- Organizations must design incentives (psychological safety, non-punitive reporting, audit protections, or compensation) to encourage reporting of failure cases; otherwise, strategic withholding creates negative externalities (systematic under-disclosure that raises aggregate risk).
- Learning externalities and information spillovers:
- Failure incidents are public goods: single firms may underinvest in documenting BUTs (learning externality). Policy or industry governance (shared incident taxonomies, anonymized failure reporting) can internalize externalities and raise overall system robustness.
- Risk, liability, and regulation:
- Explicitly recorded suspensions and decision dimensions aid auditing and regulatory compliance by making human responsibilities and thresholds auditable — but they also create durable evidence that can affect liability and reputational risk, reinforcing the need for thoughtful organizational protections.
- Empirical research agenda for AI economists:
- Quantify the value of boundary externalization: measure error-avoidance benefits vs. documentation/coordination costs.
- Estimate how suspension frequency and the distribution of discovered dimensions affect automation substitutability and labor demand.
- Test incentive mechanisms (psychological-safety interventions; protected reporting) to increase candid failure reporting and measure downstream effects on system performance and costs.
- Model markets for “decision-boundary-enriched” AI products and insurance products that price exposure to unrepresented dimensions.
- Policy suggestions:
- Encourage or mandate structured failure reporting standards (with legal protections) to mitigate underreporting externalities.
- Promote guidelines that require AI deployments to specify suspensions and human responsibilities in contracts and audits.
Overall, O-I-B-A-R reframes many AI deployment failures as representation and governance problems rather than mere model capability gaps. For AI economists, the framework focuses attention on measurement and incentive problems (who documents what, at what cost, and who benefits), on the economic value of explicit decision-boundary representations, and on policy mechanisms to correct market and organizational frictions that hinder candid externalization of failure knowledge.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Organizational AI failures can arise from a knowledge-representation problem at the sociotechnical interface, in addition to limitations in model capability. Organizational Efficiency | negative | AI performance when applying organizational procedures in real-world contexts |
Reading fidelity
high
Study strength
low
|
not reported
|
| AI systems are often given canonical positive procedures while organizations rely additionally on exceptions, contextual judgments, responsibility assignments, and repair history. Organizational Efficiency | negative | Completeness of the decision knowledge represented to an AI system |
Reading fidelity
high
Study strength
low
|
not reported
|
| The O-I-B-A-R framework externalizes decision boundaries through five stages: OPEN, IS, BUT, ACTION, and RESULT. Organizational Efficiency | positive | Externalization of organizational decision-boundary knowledge |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| A concrete BUT case is not merely the logical negation of IS; it can introduce a decision-relevant variable that is absent from the original positive rule. Decision Quality | positive | Identification of variables explaining when an organizational judgment fails |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Comparing a qualified positive IS case with a negative BUT case can identify a minimally sufficient decision dimension, ideally with one changing decision-relevant variable. Decision Quality | positive | Discovery of decision-relevant variables for distinguishing successful and failed judgments |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Representing an unresolved but known decision dimension as a suspension converts generic uncertainty into an actionable query about what must be measured, retrieved, asked, or escalated. Decision Quality | positive | Identification and resolution of missing runtime information before action |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| A suspension can specify the substantive content of a human–AI handoff rather than merely labeling a case as complex. Task Allocation | positive | Allocation of unresolved decision tasks between AI systems and human workers |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| RESULT should distinguish boundary confirmation, boundary or threshold refinement, and structural dimension expansion after an action produces evidence. Organizational Efficiency | positive | Learning and updating of organizational decision representations after failures or successes |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Observed structural stability of an organizational decision representation is local to the exposed case stream and should not be interpreted as proof that all relevant decision dimensions have been discovered. Organizational Efficiency | mixed | Completeness and stability of the represented decision-variable set |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Durable and attributable failure histories can suppress the candor needed to externalize useful boundary knowledge. Ai Safety And Ethics | negative | Employee willingness to report failures and workarounds |
Reading fidelity
high
Study strength
low
|
not reported
|
| A cited hospital study found that multiple AI tools with strong benchmark-style accuracy failed to meet organizational expectations until managers addressed the gap between experts' know-what and richer know-how. Organizational Efficiency | negative | Organizational acceptance and practical performance of AI tools |
Reading fidelity
high
Study strength
low
|
not reported
|