0 cumulative citations
View corpus contextAn entropy‑based monitoring system flags individual and coalition survey manipulation early and enables dimension‑specific corrections; tested on simulated attacks drawn from 1,233 participants’ data, it perfectly detects straight‑lining and recovers suppression coalitions with 76.5% cluster purity, outperforming standard anomaly detectors.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextInstitutional governance systems increasingly rely on stakeholder surveys in strategic decision-making. Yet survey participants can act as strategic agents who shape organizational outcomes in their favor. By transforming the information content of response distributions, such behavior can systematically distort institutional decisions made under uncertainty. This study proposes a dynamic framework that detects strategic data manipulation in institutional surveys using Shannon entropy and Kullback–Leibler divergence, and converts this detection into decision support. The framework operates in three stages. First, it constructs a robust reference entropy profile from historical data. Second, it processes incoming survey responses as a sequential stream and compares them against this profile. Third, it detects manipulation at both the population and individual levels through a multi-layered anomaly scoring system. The reference profile was built from anonymized real survey data spanning 2021–2025, comprising 1233 participants and 19,728 clean observations. The framework was validated through 600 Monte Carlo scenarios derived from this profile, covering four manipulation types and five intensity levels, and was benchmarked against the Z-score and Isolation Forest methods. The findings are threefold. Straight-lining detection identifies individual suppression and inflation manipulations with perfect accuracy. KL divergence monitoring flags coordinated coalition entries before data collection is complete. Hierarchical clustering recovers a coordinated suppression group that individual scoring fails to isolate, with 76.5% cluster purity and ~59% recall. Policy impact analysis further shows that manipulation distorts dimensions in opposite directions: the gap between raw and verified means is positive in the dimension targeted by coordinated suppression but clearly negative in the dimension targeted by coordinated inflation. This bidirectional distortion shows why dimension-selective detection is necessary, as a single uniform correction cannot resolve it. The study contributes to the literature in two areas, integrating decision-making under uncertainty with strategic agent models, and survey integrity research with information-theoretic metrics.
Summary
Main Finding
A three-stage, information-theoretic framework using Shannon entropy and Kullback–Leibler (KL) divergence can detect strategic manipulation in institutional surveys and convert that detection into decision support. The framework reliably identifies individual straight‑lining (suppression/inflation) with perfect accuracy, flags coordinated coalition entries early via KL monitoring, and—when combined with hierarchical clustering—recovers coordinated suppression groups with substantial cluster purity. Manipulation produces directionally opposite distortions across survey dimensions, implying that dimension‑selective detection and correction are necessary for sound institutional decisions.
Key Points
- Framework architecture (three stages):
- Build a robust reference entropy profile from historical “clean” survey data.
- Process incoming responses as a sequential stream and compare to the reference profile in real time.
- Apply a multi-layer anomaly scoring system to detect manipulation at population and individual levels; translate detections into decision support.
- Information‑theoretic metrics used: Shannon entropy for baseline profiling; KL divergence for divergence monitoring.
- Detection results:
- Straight‑lining detector perfectly identifies individual suppression and inflation manipulations (100% accuracy reported).
- KL divergence monitoring detects coordinated coalition entries before data collection completes.
- Hierarchical clustering recovers coordinated suppression groups with 76.5% cluster purity and ~59% recall when individual scoring alone cannot isolate them.
- Policy impact: manipulation shifts verified vs. raw means in opposite directions depending on targeted dimension (positive gap where coordinated suppression targeted; negative where coordinated inflation targeted). A single uniform correction is insufficient.
- Benchmarking: framework validated against Z‑score and Isolation Forest baselines using synthetic scenarios derived from empirical data.
Data & Methods
- Data:
- Anonymized institutional survey records (clean historical data) covering 2021–2025.
- 1,233 participants and 19,728 clean observations used to construct the reference profile.
- Validation:
- 600 Monte Carlo scenarios generated from the empirical profile.
- Scenarios covered four manipulation types (individual suppression, individual inflation, coordinated suppression coalition, coordinated inflation coalition) and five intensity levels.
- Methods:
- Reference entropy profile: aggregate Shannon entropy measures across dimensions from historical data to establish expected information patterns.
- Sequential monitoring: treat incoming responses as a stream; compute KL divergence between observed and reference distributions for early warning.
- Multi‑layer anomaly scoring:
- Individual-level detectors (including straight‑lining tests) flag per‑respondent deviations.
- Population-level metrics (KL divergence) flag coordinated shifts.
- Hierarchical clustering groups anomalous individuals to reveal coalitions that individual scores miss.
- Benchmark comparisons to Z‑score and Isolation Forest anomaly detection methods to contextualize performance.
Implications for AI Economics
- Strategic agents in data pipelines: Survey responses are endogenous inputs that strategic agents can manipulate; economists and mechanism designers must treat survey data as potentially strategic, not merely noisy.
- Real‑time monitoring for policy decisions: Sequential KL monitoring enables early detection of coordinated entries, allowing adaptive policy or data‑collection responses (e.g., extended sampling, targeted validation) that can reduce decision bias before final aggregation.
- Dimension‑selective correction: Bidirectional distortions across dimensions mean standard uniform de‑biasing is inadequate; econometric adjustment and decision rules should be dimension‑aware and informed by detected manipulation type.
- Integration into automated decision systems: Information‑theoretic anomaly scores can be incorporated into algorithmic governance pipelines (automated alerts, weighted aggregation, validation triggers), improving robustness of AI systems that rely on stakeholder survey inputs.
- Incentive and mechanism design: Detection capability changes the strategic environment — organizations can design incentives or penalties contingent on detected manipulative patterns, reducing returns to manipulation and improving data quality.
- Research directions: Combining information‑theoretic detection with causal identification and incentive-aware mechanism design; exploring privacy-preserving implementations; and extending evaluation to other survey modalities and real adversarial campaigns.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The straight-lining detector perfectly identifies individual suppression and inflation manipulations, with 100% accuracy. Error Rate | positive | Accuracy of individual-level manipulation detection |
Reading fidelity
high
Study strength
medium
|
n=600
100% accuracy
|
| Sequential KL-divergence monitoring detects coordinated coalition entries before data collection is complete. Governance And Regulation | positive | Early detection of coordinated survey manipulation |
Reading fidelity
high
Study strength
medium
|
n=600
|
| Hierarchical clustering recovers coordinated suppression groups with 76.5% cluster purity. Error Rate | positive | Purity of recovered coordinated-suppression clusters |
Reading fidelity
high
Study strength
medium
|
n=600
76.5% cluster purity
|
| Hierarchical clustering recovers approximately 59% of coordinated suppression-group members when individual scoring alone cannot isolate them. Error Rate | positive | Recall of coordinated-suppression group members |
Reading fidelity
high
Study strength
medium
|
n=600
~59% recall
|
| Manipulation shifts the gap between verified and raw survey means in opposite directions depending on the targeted dimension: coordinated suppression produces a positive gap, whereas coordinated inflation produces a negative gap. Decision Quality | mixed | Difference between verified and raw survey means across dimensions |
Reading fidelity
high
Study strength
medium
|
n=600
|
| A single uniform correction is insufficient because manipulation creates dimension-specific, bidirectional distortions. Decision Quality | negative | Adequacy of uniform correction for manipulated survey data |
Reading fidelity
high
Study strength
medium
|
n=600
|
| The reference entropy profile was constructed from 1,233 participants and 19,728 clean observations from anonymized institutional survey records covering 2021–2025. Governance And Regulation | positive | Construction of the baseline survey-information profile |
Reading fidelity
high
Study strength
medium
|
n=1233
|