0 cumulative citations
View corpus contextCommercial text-to-music models are flattening musical variety: Lyria narrows within-genre sonic diversity while Suno erases genre boundaries, and audio-feature classifiers can almost perfectly tell machine-made tracks from human ones, raising cultural and economic justice concerns about which styles become legible and rewarded at scale.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally distinct homogenizing tendencies. Lyria reduces within-genre acoustic diversity, while Suno collapses the acoustic distinctions between genres without compressing within-genre spread. Neither system follows user prompts faithfully, indicating that the observed patterns reflect learned priors rather than prompt constraints. The two systems do not converge on a common acoustic profile and are more acoustically distant from each other than two random human subsamples would typically be. Nevertheless, a standard classifier distinguishes AI from human tracks near-perfectly on MIR features alone. We argue that these patterns matter not as an aesthetic curiosity but as a justice-relevant condition, shaping which musical styles become legible, valued, and economically rewarded as generated outputs increasingly circulate at scale.
Summary
Main Finding
Generative text-to-music systems exhibit measurable acoustic homogenization relative to matched human-produced music, but in system-specific ways: Lyria 3 compresses within-genre acoustic variation, while Suno tends to collapse acoustic distinctions between genres (genre-blending) without compressing within-genre spread. These patterns reflect learned model priors rather than faithful adherence to user prompts, and a standard classifier can distinguish AI-generated from human tracks near-perfectly using only MIR features. The authors argue this technical homogenization has justice-relevant economic and cultural consequences for which musical styles become legible, valued, and rewarded at scale.
Key Points
- Definitions and focus
- Homogenization operationalized as reduced within-genre variation on standard music information retrieval (MIR) audio features (rhythm/timing, timbre/spectral shape, dynamics).
- Justice lens: recognition (cultural definition), redistribution (who captures value), and epistemic justice (whose musical distinctions are legible/authoritative).
- Systems audited
- Two commercial, black-box text-to-music systems: Suno (Pro v5.5) and Lyria 3.
- Audit period: April–May 2026; default settings; no internal access.
- Genre coverage
- Four genres chosen along geographic / representation axes: Afrobeats (non-Western, low training representation), K-pop (non-Western, higher visibility), Dance Pop (Western, high), Heavy Metal (Western, high).
- Experimental findings (summary)
- Lyria 3: reduces within-genre acoustic diversity (tracks within a genre more similar to each other than human baseline).
- Suno: preserves within-genre spread but compresses distances between genres (genres become acoustically closer).
- Neither system reliably follows detailed MIR-steered prompts; null/genre-only prompts reveal default priors.
- The two systems are more acoustically distant from each other than two random human subsamples would be—i.e., they produce distinct, system-specific signatures rather than converging to a single “AI” profile.
- A classifier trained on MIR features distinguishes AI vs human tracks with near-perfect accuracy.
- Limitations
- Black-box evaluation (no training-data or architecture access).
- Human baseline uses 30-second previews (may omit some variation).
- MIR features (72 features used) may miss musically salient but hard-to-quantify elements (microtiming, groove, regional inflection).
- Results are not automatically generalizable across all models, genres, or future system versions.
Data & Methods
- Data
- Human reference corpora: 100 human tracks per genre (400 total) sampled from public Spotify playlists (release years 2000–2026), deduplicated and cluster-sampled to represent genre acoustic diversity.
- AI-generated data: For each of two experiments, 100 generated tracks per genre per system (Suno + Lyria 3):
- Experiment 1 (MIR-steered prompting): prompts algorithmically derived from MIR descriptors of human tracks to enable pairing; produced 800 AI tracks.
- Experiment 2 (null prompting): genre-name-only prompts (e.g., “Instrumental Afrobeats track. No vocals.”) to surface system priors; produced 800 AI tracks.
- Total AI tracks across both experiments: 1,600 (800 per experiment).
- Features and diagnostics
- Extracted 72 standard MIR features capturing rhythm/timing, timbre/spectral shape, dynamics, etc.
- Diagnostics employed to test homogenization hypotheses:
- Dispersion: within-genre variance/spread in feature space.
- Redundancy: feature correlations and compression of dimensions.
- Separability: distances between genre clusters in feature space.
- Classification analyses: supervised classifier(s) trained on MIR features to separate AI vs human tracks and to examine which acoustic dimensions are most discriminative.
- Key methodological design choices
- Black-box, as-deployed evaluation on platform defaults to reflect realistic user experience.
- Two complementary prompt conditions: MIR-steered (to match human exemplars) vs null (to see defaults).
- Statistical comparisons vs matched human baseline and null-sampling procedures (e.g., comparing system-to-system distances vs human sub-sample distances).
Implications for AI Economics
- Market supply and product differentiation
- Lower marginal cost of generating music can greatly increase supply of produced content, but homogenization concentrates supply into fewer acoustic profiles or blended genres, reducing effective product differentiation for certain musical niches.
- When AI outputs are systematically similar, competition shifts from musical distinctiveness toward platform-level distribution power, branding, and other non-musical signals.
- Price formation, bargaining power, and revenue distribution
- Near-zero-cost AI substitutes for certain production types can depress bargaining power and price for human producers who create acoustically similar material (especially those already underrepresented or operating in monetizable niches).
- Platform and model owners are likely to capture disproportionate economic rents (training-data owners versus model owners, distribution channels), exacerbating existing concentration in music revenues.
- Homogenization at production reduces the scarcity value of stylistic novelty, potentially compressing royalties/licensing value for genres that become easily replicable by AI.
- Demand-side effects and recommender feedback
- If recommender systems and playlists surface AI-generated tracks (because of scale, metadata, or engagement optimization), listener exposure will shift, accelerating cultural salience of the homogenized styles and reinforcing platform incentives to promote AI outputs.
- Collapsed genre boundaries (Suno’s tendency) can alter perceived genre definitions, shifting listener preferences and industry categorizations over time.
- Labor market and creative capital
- Artists whose practices rely on microtiming, regional inflection, or other fine-grained expressive elements that MIR features underrepresent may face eroded cultural visibility and monetization as models fail to reproduce or recognize such distinctions.
- Cultural producers in underrepresented regions (e.g., Afrobeats and other Global South traditions) are particularly vulnerable if models cannot reliably reproduce or attribute their stylistic markers—affecting both recognition and economic returns.
- Market signaling and detection
- The high discriminability of AI vs human outputs on MIR features implies feasible forensic detection, which could be used for labeling, rights enforcement, or new market segmentation (e.g., “human-made” vs “AI-made” markets or pricing).
- Policy and platform design recommendations (economic interventions implied by results)
- Transparency: require disclosure about training-data composition and representation to assess encoding inequalities and support targeted remedies.
- Incentives for diversity: platform curation and recommender objectives should be adjusted to reward acoustic and cultural diversity (not just engagement), mitigating homogenization externalities.
- Revenue-sharing / licensing: explore mechanisms to compensate data-origin communities and creators whose works inform model priors (e.g., collective licensing, training-data royalties).
- Support for public / community datasets: fund representative, high-quality datasets for underrepresented musics to improve model coverage and reduce bias-driven homogenization.
- Labeling & provenance: standardized metadata and provenance signals (including AI-detection methods) to help markets price and allocate value to different production types.
- Auditing & monitoring: ongoing, independent audits of generative systems for acoustic and cultural bias to inform regulation and industry practice.
- Research & market monitoring needs
- Track longitudinal effects on streaming concentration, playlist composition, and royalty flows as AI-generated music scales.
- Study substitution vs complementarity empirically: which human roles are displaced vs augmented by text-to-music tools (session musicians, producers, composers, niche artists).
- Model the welfare trade-offs: consumer surplus from accessible music creation vs producer income losses and cultural-justice externalities.
Overall, the paper provides empirical evidence that current generative music systems do not merely mimic human diversity but introduce systematic, model-specific compressions and shifts in acoustic structure — outcomes with measurable implications for which music is legible, promoted, and economically rewarded in digital markets.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Lyria 3 reduces within-genre acoustic diversity relative to human-produced music. Creativity | negative | Within-genre variation in rhythm, timing, timbre, spectral shape, and dynamics |
Reading fidelity
high
Study strength
high
|
n=100
|
| Suno collapses acoustic distinctions between genres without compressing within-genre acoustic spread. Creativity | negative | Acoustic separability between genres and acoustic dispersion within genres |
Reading fidelity
high
Study strength
high
|
n=100
|
| Neither Suno nor Lyria 3 follows user prompts faithfully under the audit conditions. Output Quality | negative | Agreement between prompted musical characteristics and generated acoustic features |
Reading fidelity
high
Study strength
medium
|
n=100
|
| The observed homogenization patterns are more consistent with learned system priors than with prompt constraints. Creativity | negative | System-level default acoustic tendencies under minimal prompting |
Reading fidelity
high
Study strength
medium
|
n=800
|
| Suno and Lyria 3 do not converge on a common acoustic profile; each system is more acoustically distant from the other than two random human subsamples would typically be. Creativity | mixed | Acoustic distance between generative systems and between human music subsamples |
Reading fidelity
high
Study strength
medium
|
n=100
|
| A standard classifier can distinguish AI-generated tracks from human-produced tracks near-perfectly using MIR features alone. Other | positive | Classification accuracy for distinguishing AI-generated from human-produced music |
Reading fidelity
high
Study strength
medium
|
n=100
near-perfect classification
|
| The reported acoustic homogenization may shape which musical styles become legible, valued, and economically rewarded as AI-generated outputs circulate at scale. Inequality | negative | Potential effects on cultural recognition, economic rewards, and epistemic legitimacy of musical styles |
Reading fidelity
high
Study strength
speculative
|
not reported
|