0 cumulative citations
View corpus contextGenerative models’ competence is jagged: local reliability gaps can block blind adoption but reward users who learn where the model works; increasing scale raises average quality yet can leave sharp local failures and, without better local signals, render improvements invisible.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
4 cumulative citations
View corpus contextGenerative AI systems often display highly uneven performance across tasks that appear ``nearby'': they can be excellent on one prompt and confidently wrong on another with only small changes in wording or context. We call this phenomenon Artificial Jagged Intelligence (AJI). This paper develops a tractable economic model of AJI that treats adoption as an information problem: users care about \emph{local} reliability, but typically observe only coarse, global quality signals. In a baseline one-dimensional landscape, truth is a rough Brownian process, and the model ``knows'' scattered points drawn from a Poisson process. The model interpolates optimally, and the local error is measured by posterior variance. We derive an adoption threshold for a blind user, show that experienced errors are amplified by the inspection paradox, and interpret scaling laws as denser coverage that improves average quality without eliminating jaggedness. We then study mastery and calibration: a calibrated user who can condition on local uncertainty enjoys positive expected value even in domains that fail the blind adoption test. Modelling mastery as learning a reliability map via Gaussian process regression yields a learning-rate bound driven by information gain, clarifying when discovering ``where the model works'' is slow. Finally, we study how scaling interacts with discoverability: when calibrated signals and user mastery accelerate the harvesting of scale improvements, and when opacity can make gains from scaling effectively invisible.
Summary
Main Finding
Generative-AI systems can be systematically "jagged": performance varies sharply across nearby tasks because model knowledge is spatially uneven and users cannot observe local reliability. When knowledge is modelled as scattered support points and truth as a rough process, the user-facing uncertainty (posterior variance) between support points has closed-form structure; longer gaps dominate experienced error (the inspection paradox). As a result, average benchmark scores can understate experienced risk, adoption depends on discoverability as well as scale, and investments in calibration, interface design, and user mastery can be economically complementary to model scaling.
Key Points
- Artificial Jagged Intelligence (AJI): uneven, local model performance—pockets of competence and holes of high error—plus opacity about where those pockets are.
- Parsimonious formalisation:
- Truth Y(x) is a Brownian motion along a task dimension (rough, locally correlated).
- Knowledge/support points {xi} follow a homogeneous Poisson point process with intensity λ (coverage density).
- The AI interpolates between neighbouring knowledge points; posterior variance at x is the irreducible local uncertainty.
- Posterior variance (Brownian bridge) for x in gap [xi, xi+1]:
- σ^2(x) = (x − xi)(xi+1 − x) / (xi+1 − xi).
- Inspection paradox (length bias): a uniformly drawn task is more likely to fall in longer gaps. Gap containing a random task X* ∼ Gamma(2, λ) with mean 2/λ (double the mean gap 1/λ).
- Expected experienced posterior variance:
- E[σ^2] = E[X*]/6 = (2/λ)/6 = 1/(3λ). (Contrast: naive use of mean gap gives 1/(6λ).)
- Economic consequences:
- Blind adoption: a user who cannot observe local σ^2(x) uses the tool only if stakes parameter q is large enough to tolerate expected error; inspection paradox raises the required q.
- Scaling (increasing λ) reduces average error but preserves relative jaggedness — long gaps still dominate experienced exposure.
- Calibration (observing/estimating local uncertainty) can unlock positive expected value even when blind adoption would be irrational: users can selectively use the model where it’s reliable.
- Mastery (learning the reliability map) is not instantaneous: modelled as Gaussian process regression over the latent reliability function, learning rates are bounded by information gain; in high-dimensional or rough landscapes, discovering where the model works can be slow and subject to an "abstention trap".
- Complementarities/Substitutability: below the blind-adoption threshold, scaling and calibration are complements (both needed); above it they are substitutes (marginal calibration value smaller).
- Interface and governance tools (similarity cues, uncertainty estimates, abstention mechanisms, provenance) are strategic complements to model improvements and can be decisive for adoption and productivity.
Data & Methods
- Analytical, tractable economic model of task-level reliability:
- Task space: one-dimensional domain Z (for closed-form results; higher-d generalisations discussed).
- Truth process: driftless Brownian motion (captures roughness; yields Brownian-bridge conditional distributions).
- Knowledge/support points: homogeneous Poisson point process with intensity λ (represents knowledge/coverage density; interpretable as scale).
- AI prediction: posterior mean interpolation between neighbouring support points; posterior variance given by Brownian-bridge formula (equation above).
- Key derived objects:
- Length-biased gap distribution for task-uniform sampling: f_X*(x) = λ^2 x e^{−λx} (Gamma(2,λ)).
- Conditional expected variance in a gap of length X: E[σ^2 | X] = X/6.
- Expected experienced variance E[σ^2] = 1/(3λ).
- Learning/mastery model:
- Users attempting to learn the reliability map are modelled via Gaussian process regression tools; learning-rate bounds are derived that scale with information gain (links to Gaussian process bandit literature).
- Conceptual/analytical techniques: combination of spatial Poisson processes, Brownian bridge variance calculations, Bayesian posterior variance as local error measure, and Gaussian-process learning theory (Rasmussen & Williams; Srinivas et al.) to bound mastery rates.
Implications for AI Economics
- Evaluation and Benchmarks:
- Headline/benchmark averages can mislead: they do not reflect length-biased exposure to gaps. Regulators, evaluators and purchasers should weight task distributions and consider measures of local worst-case/variance, not only mean accuracy.
- Adoption and Productivity:
- Opacity about local reliability can suppress adoption even when average performance is good (blind-adoption threshold). Organisations may underuse otherwise valuable models if rare but severe local failures are hard to detect.
- Experienced error is amplified by the inspection paradox; policy and procurement should account for exposure-weighted risk.
- Product & Interface Design:
- Signals that improve discoverability (uncertainty estimates, similarity cues, provenance, abstention options) are high-return investments; they can convert jagged capability into usable value without changing base model scale.
- Calibration tools and user training (mastery) interact with scaling: below certain thresholds both are needed; above them one can substitute the other. Optimal deployment mixes depend on stakes and λ.
- Model Scaling & Research Priorities:
- Increasing training/coverage density (λ) reduces overall error but does not eliminate jaggedness; scaling alone may leave surprising failures intact.
- Opacity can hide gains from scaling: if users cannot discover where reliability improved, scale benefits may not be harvested.
- Research that increases regularity (smoother interpolation), retrieval/coverage strategies that reduce long gaps, or methods that make local uncertainty observable can be as valuable as raw parameter scaling.
- Policy and Governance:
- Regulation that focuses solely on average performance is insufficient. Policies should require (or incentivise) disclosure of uncertainty, provenance, and tools enabling discoverability, especially for high-stakes domains where users cannot absorb rare catastrophic errors.
- Certification and procurement could condition on exposure-weighted risk metrics or on the existence of reliable calibration/abstention mechanisms.
- Research implications:
- Empirical work should measure task-uniform exposure to errors, not only benchmark averages.
- Further theoretical and empirical work to quantify learning costs (mastery) in real workflows and high-dimensional task spaces is important for understanding adoption dynamics.
Summary: AJI reframes adoption and evaluation of generative-AI systems as an information problem about where a model is locally reliable. Simple stochastic geometry (Poisson support + rough truth) produces closed-form insights (inspection paradox; experienced variance = 1/(3λ)) that change how we should think about scaling, calibration, interfaces, and policy.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Generative AI systems often display highly uneven performance across tasks that appear 'nearby': they can be excellent on one prompt and confidently wrong on another with only small changes in wording or context. We call this phenomenon Artificial Jagged Intelligence (AJI). Other | null_result | local reliability (unevenness of model performance across similar prompts) |
Reading fidelity
high
Study strength
low
|
not reported
|
| The paper develops a tractable economic model of AJI that treats adoption as an information problem: users care about local reliability, but typically observe only coarse, global quality signals. Adoption Rate | null_result | adoption decision / adoption threshold under information frictions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In a baseline one-dimensional landscape, truth is a rough Brownian process, and the model 'knows' scattered points drawn from a Poisson process; the model interpolates optimally, and the local error is measured by posterior variance. Error Rate | null_result | local prediction error (posterior variance) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The authors derive an adoption threshold for a blind user. Adoption Rate | mixed | adoption threshold (user's decision rule under coarse/global signals) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Experienced errors are amplified by the inspection paradox. Error Rate | negative | experienced error rate (user-observed errors) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Scaling laws can be interpreted as denser coverage that improves average quality without eliminating jaggedness. Output Quality | mixed | average model output quality and persistence of local variability (jaggedness) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A calibrated user who can condition on local uncertainty enjoys positive expected value even in domains that fail the blind adoption test. Decision Quality | positive | expected value / utility to a calibrated user from using the model |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Modeling mastery as learning a reliability map via Gaussian process regression yields a learning-rate bound driven by information gain, clarifying when discovering 'where the model works' is slow. Skill Acquisition | mixed | learning rate for discovering reliability (speed of mastery) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| When calibrated signals and user mastery accelerate the harvesting of scale improvements, and when opacity can make gains from scaling effectively invisible. Adoption Rate | mixed | visibility and harvestability of scale-driven quality improvements; effective adoption/benefit realization |
Reading fidelity
high
Study strength
medium
|
not reported
|