0 cumulative citations
View corpus contextAsking more AI advisers makes crowd answers more reliable but also more visibly split — and if individual advisers are under 80% accurate, visible dissent eventually outpaces a correct majority, so disclosure and presentation of verdicts must be designed, not assumed.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the majority more reliable. The theorem, however, has a latent dimension when viewed from the user's vantage: adding advisers also makes disagreement more visible. A binomial model reveals that this ``visible dissent'' becomes nearly inevitable as the number of advisers grows, and that reliability and disagreement both approach certainty but at different convergence rates. The two rates cross at an adviser accuracy of 4/5 (0.8). Below this value, visible dissent approaches certainty faster than reliability and, with enough advisers, becomes more likely than a correct majority. Even ideal panels of independent and competent advisers can be correct in aggregate but appear divided; such disagreement does not by itself indicate aggregation failure. The way advisers split also provides a common basis for predictive multiplicity, reconciliation load, and reliance miscalibration. These results indicate two distinct decisions when using multiple AI advisers: how many advisers to consult and how their verdicts should be presented and interpreted.
Summary
Main Finding
When K independent AI advisers each have equal accuracy p > 1/2 on a binary question, two things both tend to certainty as K grows: (i) the majority is correct (Condorcet) and (ii) the user increasingly observes a visible split among advisers. These two probabilities converge at different exponential rates. The rates cross at adviser accuracy p = 4/5 (0.8). For p < 0.8 there exists a finite panel size K beyond which a visible split becomes at least as likely as a correct majority; for p ≥ 0.8 no finite K exists. Thus even “ideal” panels can be correct in aggregate yet commonly appear divided — and how often this happens is quantitatively predictable.
Key Points
- Base model assumptions: K independent advisers, equal per-item accuracy p with 1/2 < p < 1, binary verdicts, user observes only the vector of votes (not the ground truth).
- Visible dissent (minimal split): defined operationally as at least two dissenters on each side (the smallest split not attributable to a single random error). Closed-form probability: Pvis(K,p) = 1 − q^{K−1}(K p + q) − p^{K−1}(K q + p), where q = 1 − p. Under this definition Pvis → 1 for any fixed imperfect p as K → ∞.
- Aggregate reliability (Pagg, majority correctness) also → 1 as K → ∞ (classical Condorcet result). Both approach 1 exponentially in K but with different exponents.
- Comparison of exponential rates yields a sharp boundary at p = 4/5:
- If p > 0.8, Pagg increases faster and always dominates Pvis (no finite K*).
- If p < 0.8, Pvis increases faster and there is a finite disclosure reference scale K(p) where Pvis ≥ Pagg for the first time (examples: K(0.6)=6, K(0.7)=14, K(0.75)=32; K* diverges as p → 0.8).
- The expected fraction of dissenters m → q (the advisers’ error rate) as K grows. So even when majority is almost certainly correct, users will typically see roughly q fraction dissenting.
- Dependence, unequal accuracy, multi-class outputs, and real-model correlations change quantitative outcomes but do not remove the structural point: adding advisers increases the visibility of dissent. Correlation can reduce observed split rates but can greatly increase probability of unanimous-but-wrong panels (“false consensus”).
- The paper connects this observable split to three user-side concepts:
- Predictive multiplicity (probability any disagreement occurs).
- Reconciliation load (cognitive/work cost to reconcile conflicting advice, increasing with the dissenting fraction).
- Reliance miscalibration (under- or over-trust driven by observing splits or apparent unanimity). The model provides explicit, testable predictions for these stimuli; the behavioral responses remain empirical questions (the authors posit a monotone-response hypothesis).
Data & Methods
- Analytic probability model: binomial voting model (K independent Bernoulli(p) votes) with closed-form expressions for Pvis and standard expressions for Pagg (majority correctness). Derivations use binomial sums and large-deviation / error-exponent comparisons.
- Asymptotic analysis: compare exponential convergence rates of Pvis complement and majority error (rate expressions give the 4/5 crossover).
- Numerical examples: compute Pvis and Pagg across (K,p) to illustrate crossovers and K values. Provide example K values for representative p.
- Robustness checks and extensions (sketched): generalizing minimum-split threshold to a fixed fraction r; simulations of correlated votes using copula models to show dependence effects (dissent deficit, inflated all-wrong probabilities).
- Conceptual linkage to literatures: predictive multiplicity, jury-split models, ensemble diversity, and disagreement-based uncertainty quantification.
Implications for AI Economics
- Two economically relevant decisions: (a) how many advisers to consult (K), and (b) how to present/interpret their verdicts. Both have welfare and cost trade-offs.
- Marginal benefits vs marginal costs:
- Benefit: increasing K reduces majority error at an exponential rate (classical ensemble benefit).
- Cost: increasing K raises the probability of visible dissent and thus expected reconciliation costs (time, expert review, user uncertainty), potential reductions in adoption or reliance, and administrative burdens. These costs can dominate behavioral outcomes for p < 0.8 once K exceeds K*.
- Disclosure and product design economics:
- Presentation choice (show vote distribution vs only majority) affects perceived uncertainty and trust. The paper supplies a quantitative benchmark (K*(p), the 4/5 boundary) to guide when fuller disclosure is likely to produce a majority-correct panel that nonetheless appears divided.
- Platforms can use the independence baseline to detect residual dependence (“dissent deficit”) and to audit ensemble behavior; deviations can indicate model similarity/monoculture risks that have market and systemic implications.
- Pricing and market structure:
- If querying advisers is costly, optimal K must account for both reduced prediction error and increased reconciliation/interpretation costs. A cost-benefit framework could price additional model queries or structure tiered offerings (e.g., small-K “single-shot” products vs larger-K “panel” products with explicit reconciliation services).
- Incentives for diversity: regulators or buyers might subsidize or require model diversity (reduces correlated failures) but diversity also changes perceived disagreement rates — design and disclosure policy must balance these effects.
- Regulation and consumer protection:
- The results argue for policy attention to how ensemble outputs are disclosed: showing raw vote splits can be informative but may lower apparent reliability; hiding splits risks over-trust when unanimity is actually spurious (false consensus). Disclosure rules might condition presentation on estimated p and K relative to K*(p).
- Audit metrics: the independence baseline and the Pvis formula provide testable metrics for auditors to flag suspiciously low or high disagreement relative to what independent equal-accuracy advisers would show.
- Strategic behavior and platform competition:
- Firms might optimize K and presentation to affect customer perceptions (e.g., fewer advisers or aggregated summaries to reduce visible dissent). Regulators should be aware of incentives to understate disagreement.
- Standardization (monoculture) can reduce user-perceived splits but raises systemic risk of correlated failures; market incentives (cost, data access) can push toward or away from such monocultures.
- Empirical agenda for AI economics:
- Measure per-item p and user behavioral responses to split stimuli to estimate reconciliation costs, changes in reliance, and welfare outcomes.
- Calibrate K* for realistic adviser accuracies and query costs to derive policy and pricing guidelines.
Caveats and open limitations - Core results rely on equal-accuracy, independence, and binary outputs; real advisers violate these assumptions. The paper treats the base model as a useful null baseline and documents qualitatively how dependence and heterogeneity change conclusions (notably by inflating false-consensus risks). - Behavioral responses to splits (reconciliation load, trust updating) are not modeled endogenously — these require empirical study to translate split probabilities into economic welfare impacts. - Multi-class outputs, cost per query, and strategic model selection are natural extensions for economic modeling.
Bottom line Adding independent advisers increases aggregate accuracy but also makes visible disagreement more likely; below adviser accuracy p = 0.8, a modest panel size suffices for visible dissent to be as likely as a correct majority. Designers, platform managers, and regulators should therefore treat the number of advisers and the mode of disclosure as separate, economically consequential choices; the paper provides explicit formulas and reference scales (p = 4/5 and K*(p)) to guide those decisions and to form audit baselines.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| For K independent advisers with equal accuracy p greater than 1/2, aggregate majority reliability approaches certainty as the panel grows. Decision Quality | positive | Probability that the majority verdict is correct (Pagg) |
Reading fidelity
high
Study strength
high
|
Pagg(K, p) approaches 1
|
| Under the paper's definition of visible dissent as at least two advisers on each side, the probability of visible dissent approaches one as the number of advisers grows for every imperfect adviser accuracy p less than 1. Decision Quality | positive | Probability that users observe a non-negligible split in adviser verdicts (Pvis) |
Reading fidelity
high
Study strength
high
|
Pvis(K, p) approaches 1
|
| The convergence rates of aggregate reliability and visible dissent cross at adviser accuracy p = 4/5, or 0.8. Decision Quality | mixed | Relative convergence rates of majority reliability and visible-dissent probability |
Reading fidelity
high
Study strength
high
|
p = 4/5 (0.8)
|
| When adviser accuracy is below 0.8, visible dissent eventually becomes at least as likely as a correct majority; when accuracy is at least 0.8, the paper's analytic result rules out such a crossover. Decision Quality | mixed | Relative probability of visible dissent versus a correct majority |
Reading fidelity
high
Study strength
high
|
Finite crossover exists for p < 4/5; none for p ≥ 4/5
|
| At adviser accuracy p = 0.9 and K = 20, the majority is correct with probability above 99.999%, while the user observes a split about 61% of the time. Decision Quality | mixed | Majority accuracy and probability of visible dissent |
Reading fidelity
high
Study strength
high
|
n=20
majority accuracy >99.999%; visible dissent ≈61%
|
| As the panel grows, the expected dissenting fraction converges to the advisers' error probability q = 1−p, even when the majority is correct. Decision Quality | positive | Expected fraction of advisers dissenting from the majority, E[m(K)] |
Reading fidelity
high
Study strength
high
|
m approaches q = 1−p
|
| Under fair tie-breaking, the first dissent–reliability crossover occurs at K = 6 for p = 0.6 and K = 14 for p = 0.7. Decision Quality | negative | Panel size at which visible dissent becomes at least as likely as a correct majority |
Reading fidelity
high
Study strength
high
|
n=14
K* = 6 at p = 0.6; K* = 14 at p = 0.7
|
| For p = 0.7 and a panel vote split of 8–6, the probability that the majority is correct is 0.845 conditional on observing the split, compared with an unconditional majority accuracy of 0.938. Decision Quality | negative | Posterior probability that the majority verdict is correct after observing a narrow 8–6 split |
Reading fidelity
high
Study strength
high
|
n=14
0.938 unconditional versus 0.845 conditional
|
| In simulated correlated panels, positive dependence lowers the split rate, while correlated errors can make all-wrong unanimous panels more common than the q^K probability predicted under independence. Ai Safety And Ethics | mixed | Visible split rate and probability of unanimous but incorrect panels |
Reading fidelity
high
Study strength
medium
|
No numerical effect size reported in the supplied text
|