0 cumulative citations
View corpus contextPersonality matching from public tweets nudges ad attention: in a randomized online study, cross‑trait pairings — especially conscientious ads in neurotic apps — raised click intention and recall, while same‑trait placements reduced effectiveness; the approach offers a privacy‑friendly way to prioritize ad placements before field testing.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextMobile display ads bring in about two-thirds of all app revenue, yet the format often falls short because the ads and the apps they appear in are often poorly matched. As privacy regulations tighten and platforms lose access to user-level data, advertisers are left with superficial signals like app price, category, and rating that carry limited targeting power. We propose a different approach: matching ads to apps on inferred personality, derived from public discourse rather than user-level data. Using 255,531 public tweets from roughly 121,855 unique authors on X (formerly Twitter), we score 45 mobile apps and 53 advertised brands on the Big Five traits. To test whether this matching translates into ad effectiveness, we ran an online experiment with 1,979 participants. Each participant saw one randomly assigned ad-app pair under an app-usability cover story, and we estimated logistic regressions for click intention, brand recall, and category recall. We find that neurotic apps are associated with higher click intention, especially when paired with low-openness ads typical of telecommunications and financial services brands. Conscientious ads (e.g., financial services, healthcare) show higher brand recall as app neuroticism increases, while low-conscientiousness ads (e.g., media, entertainment) perform better in more agreeable apps. Same-trait pairings consistently hurt across both click and recall outcomes; the gains come from cross-trait complementarity, with the Conscientiousness-Ad and Neuroticism-App pairing emerging as the most replicated effect. The paper contributes to research on personality complementarity by extending it to ad-app matching and offers practitioners a privacy-friendly approach for prioritizing promising pairings before subsequent A/B testing.
Summary
Main Finding
Discourse-derived Big Five personality profiles for apps and advertisers—inferred from public tweets—predict ad effectiveness in a randomized online experiment. Complementary (cross-trait) pairings drive gains: apps high in Neuroticism elicit higher click intention (especially when paired with low-Openness ads typical of telecom/financial brands), Conscientious ads yield higher brand recall in neurotic apps, and low-Conscientious ads do better in agreeable apps. Same-trait (matching) pairings tend to reduce click and recall. The Conscientiousness(Ad) × Neuroticism(App) interaction is the most consistent effect. The approach offers a privacy-preserving way to prioritize ad–app pairings before field A/B tests.
Key Points
- Novel unit of analysis: the ad–app dyad, matched by inferred personality (Big Five) rather than by user-level profiling.
- Personality inferred at the entity level from public social-media discourse (X/Twitter), avoiding individual-level tracking.
- Complementarity, not similarity/congruence, explains the strongest ad-performance effects.
- Same-trait pairings consistently suppress outcomes (click intention and recall).
- Most replicated effect: Conscientiousness of the ad interacting positively with Neuroticism of the app.
- Practical value: a low-privacy-cost screening tool to prioritize placements and reduce costly trial-and-error.
Data & Methods
- Entity samples:
- Apps: 45 mainstream mobile apps (drawn from top App Store lists; 15 app categories). Examples: Facebook, Gmail, Google Chrome, GroupMe.
- Ads/brands: 53 advertised brands (11 advertiser categories). Examples: McDonald’s, Nike, T-Mobile, United Airlines, CVS, American Express.
- Social-media corpus:
- Collected all public tweets mentioning each app or advertised brand over a one-month window.
- 255,531 tweets from ~121,855 unique authors.
- Per-entity averages: apps ≈ 2,449 tweets (SD 933), ≈ 25,762 words (SD 10,861); ads ≈ 2,742 tweets (SD 517), ≈ 30,080 words (SD 6,627).
- Personality scoring:
- Aggregated tweets per entity into single documents.
- Scored on Big Five percentiles using IBM Watson Personality Insights (supervised model trained on labeled social-media text). Note: the paper indicates scoring was done while the service was active; IBM retired the service in 2021 but model files were retained.
- Experimental design:
- Online between-subjects experiment with N = 1,979 US participants (MTurk Masters qualification).
- Cover story: app-usability feedback. Each participant viewed one synthetic app (five screens); a randomly assigned display ad appeared on the third screen.
- Outcomes:
- Awareness (noticed ad yes/no).
- Awareness-conditioned click intention (would you have tapped/clicked? yes/no) — primary click outcome used in logistic regressions.
- Brand and category recall measured from free-text responses and validated programmatically.
- Analysis: logistic regressions predicting click intention, brand recall, and category recall from ad and app personality scores and their interactions, controlling for app metadata (price, rating, category) as applicable.
- Robustness/limitations discussed by authors:
- Behavioral tap logs collected but not used as main outcome due to mechanical ambiguity with swipe gestures.
- Personality inference relies on public discourse volume and composition; results depend on coverage and quality of social mentions.
Implications for AI Economics
- Privacy-preserving targeting signal:
- Platforms can use aggregate, public-discourse NLP signals to improve contextual ad allocation without user-level identifiers, helping adapt to stricter privacy regimes (GDPR/CCPA, platform tracking restrictions).
- Platform and market effects:
- Better pre-screening of promising ad–app pairs can reduce A/B testing costs, speed up advertiser learning, and increase developer/app revenue from display ads.
- Introducing personality-based matching into allocation/auction systems could change slot valuation and price differentiation across inventory (apps with certain personality profiles may command premium CPMs for compatible advertisers).
- Model & productization considerations:
- Entity-level personality scoring requires ongoing NLP pipelines and monitoring for drift, sampling bias (Twitter/X user base), and manipulation (brands or bots influencing public discourse).
- Retired third-party tools (e.g., IBM Watson Personality Insights) suggest firms must maintain or validate their own models; transparency about model provenance and stability will matter for adoption.
- Welfare, competition, and regulation:
- The method sidesteps user profiling, reducing some privacy harms, but raises other concerns: differential ad exposure and potential persuasion asymmetries across user groups; platform incentives to favor certain pairings; risks of strategic discourse manipulation to game matching.
- Research and policy agenda in AI economics:
- Quantify economic gains from entity-level personality matching (e.g., revenue uplift, reduced experimentation cost) and distributional impacts across app categories and advertiser types.
- Analyze platform-level optimization: how to incorporate personality scores into auction mechanisms while preserving fairness and preventing gaming.
- Study external validity: replicate in-field (real impression/click) experiments, across platforms and non-Twitter discourse sources, and evaluate long-term effects on user experience and retention.
- Practical recommendations for practitioners:
- Use discourse-derived personality profiles to prioritize candidate ad–app placements for small-scale field tests rather than replacing A/B tests.
- Monitor signal quality (coverage, sentiment shifts), guard against manipulation, and combine personality signals with existing contextual metadata for robustness.
- Estimate potential pricing / revenue impacts before full integration into allocation algorithms.
Limitations to keep in mind - Personality scores come from aggregated public tweets (platform-specific coverage and demographic/skew biases). - Experiment relies on MTurk self-reported click intention and recall under a cover story—not direct measured conversions in naturalistic app usage. - The IBM Personality Insights model used was retired by IBM; reproducibility requires alternative validated models or internal retraining. - The sample of apps and ads is mainstream but limited (45 apps, 53 ads); generalizability to long-tail apps or niche advertisers is uncertain.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study inferred Big Five personality profiles for 45 mobile apps and 53 advertised brands using public discourse on X (formerly Twitter). Other | positive | Discourse-derived Big Five personality scores for apps and advertised brands |
Reading fidelity
high
Study strength
medium
|
n=45
|
| The personality profiles were based on 255,531 public tweets written by approximately 121,855 unique authors. Other | positive | Availability of public-discourse data for entity-level personality inference |
Reading fidelity
high
Study strength
medium
|
n=255531
|
| In an online experiment, neurotic apps were associated with higher participant click intention, particularly when paired with low-openness advertisements typical of telecommunications and financial-services brands. Firm Revenue | positive | Self-reported intention to tap or click on the advertisement after noticing it |
Reading fidelity
high
Study strength
medium
|
n=1979
|
| Advertisements characterized by higher conscientiousness showed higher brand recall as app neuroticism increased. Firm Revenue | positive | Post-exposure recall of the advertised brand |
Reading fidelity
high
Study strength
medium
|
n=1979
|
| Low-conscientiousness advertisements, such as media and entertainment ads, produced better recall in more agreeable apps. Firm Revenue | positive | Recall of the advertised brand and/or product category |
Reading fidelity
high
Study strength
medium
|
n=1979
|
| Same-trait ad-app pairings were associated with lower effectiveness across click-intention and recall outcomes. Firm Revenue | negative | Advertisement click intention, brand recall, and category recall |
Reading fidelity
high
Study strength
medium
|
n=1979
|
| The strongest and most replicated pairing identified by the study was a conscientiousness-oriented advertisement placed in a neuroticism-oriented app. Firm Revenue | positive | Advertisement click intention, brand recall, and category recall |
Reading fidelity
high
Study strength
medium
|
n=1979
|
| The experiment used self-reported click intention rather than logged behavioral taps as the primary click outcome because a tap could be mechanically confused with initiating a swipe gesture. Firm Revenue | mixed | Measurement of advertisement clicking or click intention |
Reading fidelity
high
Study strength
high
|
n=1979
|
| The proposed personality-based ad-app matching approach is intended to help platforms prioritize promising pairings before field A/B testing without relying on individual-level user profiling. Organizational Efficiency | positive | Potential efficiency of ad-placement screening and privacy-preserving targeting |
Reading fidelity
high
Study strength
speculative
|
n=1979
|