0 cumulative citations
View corpus contextSpeech-recognition systems systematically sideline low-resource and non-standard language varieties, entrenching colonial language hierarchies; the authors propose a seven-layer diagnostic, a Three-Harms taxonomy, and a community-led audit protocol to make ASR more culturally competent and reduce access harms.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.
Summary
Main Finding
ASR failures for low-resource, Indigenous, and non‑standard language varieties are not merely technical mistakes but embodied linguistic policies that reproduce colonial language hierarchies. The paper reframes ASR design, data, metrics, and deployment choices as policy decisions that produce three classes of harm—Misrecognition, Misalignment, and Mistrust—and proposes conceptual tools (a seven‑layer situatedness model, a 3M taxonomy) plus a participatory audit and governance framework to make culturally competent ASR possible.
Key Points
- Conceptual contributions
- Seven‑layer situatedness model: maps how language support claims collapse situated speech practices (language, nation, region, ethno‑linguistic identity, sociolinguistic ideology, social‑justice correlations, socio‑technical implications).
- Three Harms (3M) taxonomy: Misrecognition (transcription/phonetic errors), Misalignment (wrong pragmatic/intent mapping), Mistrust (reduced uptake, dignity/legitimacy harms).
- Participatory framework and minimum audit protocol: centers affected communities as co‑designers, evaluators, and governance partners.
- ASR design choices are linguistic policies
- Data curation, evaluation metrics, and model priors encode normative choices about which speech is “legible.”
- Training datasets reflect historical and institutional biases (digital language divide, colonial legacies), so “low‑resource” is often a political/economic condition, not an inherent property.
- Limitations of dominant metrics
- WER-centric evaluation is insufficient and can mask meaning‑changing errors (e.g., tonal contrasts in Yoruba, click consonants in Khoisan languages).
- The authors propose augmenting WER/CER with domain‑appropriate measures (Tone Error Rate, click‑specific error reporting), phoneme‑level scores, and community‑weighted harm metrics.
- Model priors and deployment behaviour
- Language models that penalize code‑switching or normalize culturally meaningful forms effectively enforce a standard and exclude other speech ecologies.
- Deployment fallbacks (e.g., English‑only responses) and rigid transcription norms further entrench exclusion.
- Normative stance and positionality
- The paper is explicit about authors’ standpoints and argues for accountability via sustained, community‑led processes rather than token disclosure.
Data & Methods
- Nature of the work: conceptual/theoretical synthesis and normative framework — not an empirical ASR model or new benchmark.
- Methodological inputs:
- Cross‑disciplinary literature synthesis: linguistic capital (Bourdieu), raciolinguistic ideology, language policy/planning, and decolonial/postcolonial computing.
- Diagnostic mapping of pipeline decision sites (training data, metrics, model priors, deployment).
- Development of two instruments: (1) seven‑layer situatedness model to diagnose where ASR flattens sociolinguistic complexity; (2) 3M taxonomy to categorize harms.
- Proposal of a minimum audit protocol covering: evaluator roles, sampling expectations, culturally aware metrics (e.g., TER, click metrics), annotation procedures that accept legitimate transcript variability, and adjudication processes involving community reviewers.
- Illustrative linguistic examples used to motivate metric changes:
- Tonal contrasts (e.g., Yoruba) require tone‑aware measures.
- Click consonants and other phonemes that carry identity/meaning require explicit phoneme‑level reporting and community adjudication.
- No new datasets or numerical experiments are presented; recommendations are actionable design and evaluation practices.
Implications for AI Economics
- Market and welfare implications
- ASR systems encode and reinforce linguistic capital: supporting prestige varieties produces concentrated economic and informational access, while excluding others perpetuates economic marginalization (reduced access to services, employment barriers).
- Mistrust and misalignment can reduce adoption of ASR‑mediated services in large populations, producing negative externalities (lower productivity, lower public‑service reach) and inefficient allocation of digital public goods.
- Data colonialism and value capture
- The paper situates low‑resource status within extractive histories: data collection and commodification of dominant languages perpetuate unequal returns to language communities. Economic models of value capture must account for who supplies data, who benefits, and how governance/ownership is structured.
- Product design, procurement, and regulation
- Procurement rules or regulatory standards could require community‑weighted evaluation metrics and participatory audit evidence for systems used in public services (health, legal, education).
- Investors and firms should treat culturally competent ASR as a compliance and reputational risk‑management metric; neglect creates legal/regulatory, market‑access, and brand‑damage risks.
- Incentives and public policy levers
- Subsidies, grants, or public procurement priorities can correct market failures that disincentivize investment in low‑resource language support (e.g., fund community‑led data collection, language‑specific model adaptation).
- Mandating broader metric regimes (beyond WER) aligns private incentives with social welfare by internalizing social costs of exclusion.
- Labor and platform economics
- Community‑centered data work transforms who is paid and how: participatory design/auditing can create local economic opportunities but also raises questions about compensation, IP, and governance of language assets.
- Platforms that ignore linguistic diversity may face reduced market reach; conversely, investments in culturally competent ASR can unlock underserved markets and public‑sector contracts.
- Measurement and cost–benefit analysis
- Evaluating ASR product investments should include community‑weighted harms (reputational/legal costs, foregone access, downstream economic losses) alongside classic model accuracy. This reweights product ROI calculations and may justify higher upfront costs for inclusive data and participatory processes.
Suggested actionables for AI economists and policy actors - Incorporate community‑weighted metrics and augmented error measures into impact assessments, procurement scoring, and regulatory standards. - Treat language inclusion as a public‑good correction: fund and evaluate community‑led corpora and audits. - Model the long‑term economic benefits of inclusive ASR (increased service uptake, labor market access) against short‑term costs of data collection and participatory governance. - Require transparency about linguistic coverage, fallback behavior, audit results, and community governance agreements in deployed ASR products.
Summary: The paper reframes ASR failures as policy outcomes with economic consequences. For AI economics, this implies that evaluation metrics, governance structures, and funding incentives should be redesigned so that language inclusion (and the avoidance of cultural harms) is internalized into market and policy decision‑making.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| ASR systems routinely fail speakers of low-resource, Indigenous, and non-standard language varieties, affecting access to public services, healthcare, education, and legal processes. Consumer Welfare | negative | Access to and usability of public, healthcare, educational, and legal services through speech interfaces |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A substantial disparity can exist between ASR performance for Standard American English and African American Language: the paper gives an example of 5% WER for Standard American English versus 35% WER for African American Language. Error Rate | negative | Word Error Rate in automatic speech recognition |
Reading fidelity
high
Study strength
medium
|
5% WER versus 35% WER
|
| Language-technology performance correlates more strongly with the socioeconomic power of a language than with its linguistic complexity, and the languages performing worst in NLP systems are spoken by many economically marginalized populations. Inequality | negative | Language technology and NLP system performance across languages |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Word Error Rate is insufficient as the sole or primary measure of culturally competent ASR because it treats word-level errors as equivalent and cannot capture meaning changes, loss of speaker intent, or differential consequences in high-stakes contexts. Decision Quality | negative | Ability of ASR evaluation metrics to capture communicative and socially consequential errors |
Reading fidelity
high
Study strength
medium
|
not reported
|
| For tonal languages such as Yoruba, tone-aware evaluation is necessary because pitch can encode lexical and grammatical contrasts that WER may collapse into a single lexical score. Error Rate | negative | Recognition of tone-dependent lexical and grammatical distinctions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| For languages with click consonants, culturally competent ASR audits should measure click-specific deletion, substitution, insertion, and misclassification, preserve click place and manner distinctions, test meaning-changing minimal pairs, and include community adjudication of harms. Error Rate | positive | Detection and characterization of click-consonant recognition errors and their linguistic or cultural consequences |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| ASR evaluation against standardized reference transcriptions can enforce linguistic conformity when legitimate variation, non-standard orthographies, or code-switching are treated as errors. Ai Safety And Ethics | negative | Representation of legitimate linguistic variation in ASR evaluation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Language-model priors can assign lower confidence to code-switched speech and correct tonal distinctions toward dominant-language approximations, thereby making some speech patterns less recognizable to ASR systems. Error Rate | negative | ASR confidence and recognition of code-switched and tone-bearing speech |
Reading fidelity
high
Study strength
low
|
not reported
|
| The paper's Three Harms taxonomy identifies three forms of harm in ASR and voice interfaces: Misrecognition, Misalignment, and Mistrust. Ai Safety And Ethics | negative | Forms of harm arising from ASR and ASR-mediated voice-interface failures |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The paper proposes a participatory framework in which affected language communities act as co-designers, co-auditors, evaluators, and governance partners with meaningful influence over evaluation criteria, deployment conditions, and repair pathways. Governance And Regulation | positive | Community participation and governance in culturally competent ASR design and evaluation |
Reading fidelity
high
Study strength
speculative
|
not reported
|