The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

An audit of Llama‑3 8B reveals a North‑South information divide: the model provides verifiable-seeming numeric answers in only 11.4% of queries and systematically under-serves lower-income countries, risking uneven AI governance capacity worldwide.

Global AI Bias Audit for Technical Governance
Jason Hung · February 01, 2026
arxiv descriptive low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jason Hung unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jason Hung provider ID
An exploratory global audit of Llama-3 8B finds that technical AI governance knowledge in the model is concentrated in higher-income regions, with only 11.4% of responses presenting number/fact answers and notable information gaps for Global South countries.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper presents the outputs of the exploratory phase of a global audit of Large Language Models (LLMs) project. In this exploratory phase, I used the Global AI Dataset (GAID) Project as a framework to stress-test the Llama-3 8B model and evaluate geographic and socioeconomic biases in technical AI governance awareness. By stress-testing the model with 1,704 queries across 213 countries and eight technical metrics, I identified a significant digital barrier and gap separating the Global North and South. The results indicate that the model was only able to provide number/fact responses in 11.4% of its query answers, where the empirical validity of such responses was yet to be verified. The findings reveal that AI's technical knowledge is heavily concentrated in higher-income regions, while lower-income countries from the Global South are subject to disproportionate systemic information gaps. This disparity between the Global North and South poses concerning risks for global AI safety and inclusive governance, as policymakers in underserved regions may lack reliable data-driven insights or be misled by hallucinated facts. This paper concludes that current AI alignment and training processes reinforce existing geoeconomic and geopolitical asymmetries, and urges the need for more inclusive data representation to ensure AI serves as a truly global resource.

Summary

Main Finding

The exploratory GAID audit of Llama-3 8B (1,704 queries across 213 countries, 8 technical AI metrics) shows a pronounced geographic epistemic gap: the model provided quantitative (number/fact) answers in only 11.4% of queries, admitted ignorance/refused in 44.0%, and gave qualitative/contextual or ambiguous replies in 43.8%. Refusals and ignorance concentrate in lower-income regions (notably Sub‑Saharan Africa, Latin America & Caribbean, South Asia and many Oceanian island states), implying foundation-model knowledge is materially skewed toward higher‑income countries. The paper argues this reinforces geoeconomic asymmetries and poses risks for inclusive AI-driven governance and market decision‑making.

Key Points

  • Dataset & scope
    • Audit used GAID v2 (≈3 million datapoints aggregated from 11 authoritative sources across 227 countries/territories, 1998–2025) as the ground-truth framework.
    • Target: eight 2025 metrics across three pillars — AI safety, fairness, readiness — for up to 227 jurisdictions.
  • Model & queries
    • Model: Llama‑3 8B (open-weight).
    • Queries: 1,704 standardised prompts (8 metrics × 213 countries the model produced responses for). Template example: “What is the [metric value] for [country/territory] in 2025?”
  • Response categories & proportions
    • Factual (numeric output, unverified): 11.4%
    • Refusal / honest ignorance (explicitly states no data): 44.0%
    • Qualitative/contextual (non-numeric description or suggestion): 43.8%
    • Misunderstanding/correction: 0.7%
  • Geographic patterning
    • Highest refusal rates concentrated in Sub‑Saharan Africa, Latin America & Caribbean, South Asia, and many small Oceanian/Caribbean territories.
    • Countries with highest “knowledge rates” included typical high-income states (e.g., Australia, Spain, Austria) and some surprising entries (e.g., Angola, Madagascar, Laos), suggesting uneven and sometimes inconsistent coverage.
  • Metric heterogeneity
    • Private AI investment had the highest knowledge rate (~25%) but still >40% refusal.
    • Government AI readiness index, AI patents granted, and national hardware compute frontier had particularly low knowledge rates (~5%) with refusal rates 50–60%+.
  • Definitions & labeling
    • “Factual” labelled by presence of a numeric answer only; the audit did not verify those numbers against GAID in this exploratory phase.
    • “Refusal” flagged where the model denied the existence of data despite GAID containing it.

Data & Methods

  • Ground truth
    • GAID v2: compiled/cleaned/standardised dataset (DOI provided), integrating OECD.ai, WIPO, UNESCO and other sources for 227 countries/territories.
  • Audit pipeline
    • Metric operationalisation: eight metrics defined and measured (see Table 1 in paper).
    • Automated query generation: colab_generate_audit_queries.py produced standard prompts for each country/metric.
    • Model interrogation: Llama‑3 8B was queried via the standard prompts; responses logged in census_audit_results.csv.
    • Coding/verification (exploratory)
      • Responses coded into four categories (factual numeric, refusal, qualitative/contextual, misunderstanding).
      • Important limitation: numeric answers were not cross-validated against GAID values in this exploratory run.
  • Sample
    • 1,704 queries executed; model returned responses for 213 distinct countries/territories (out of 227 in GAID).
  • Limitations (noted by author)
    • No empirical verification of numeric outputs in this phase.
    • Single model (Llama‑3 8B); not tested across larger or more recent models.
    • Uniform prompt structure may affect retrieval; classification of responses (factual vs. qualitative) may mislabel nuanced outputs.
    • Potential dataset and scraping biases inside GAID and source data may affect what is considered “ground truth.”

Implications for AI Economics

  • Market information asymmetries
    • If foundation models systematically lack or refuse quantitative data about lower‑income countries, market actors (investors, firms, consultants) using LLM outputs risk under‑estimating opportunities or mispricing risk in those jurisdictions. This can reinforce capital concentration in high‑resource regions.
  • Investment allocation and signalling
    • Biased or absent model outputs on private investment, patents, workforce, and infrastructure can distort automated investment screening, deal sourcing, and due‑diligence workflows that rely on LLMs as first‑pass intelligence tools.
  • Policy and regulatory consequences
    • Policymakers in underserved countries who consult such models may receive incomplete or misleading technical evidence, weakening domestic policymaking capacity and potentially entrenching dependence on external consultants or indices.
    • International organizations and donors using model outputs for diagnostics might misallocate resources if models systematically under‑represent Global South capabilities.
  • Competitive dynamics & digital colonisation
    • Model training and alignment practices that privilege high‑resource data will perpetuate a feedback loop: better‑represented countries get more accurate AI outputs → attract more investment and visibility → become even more dominant. This is a form of digital/geoeconomic capture.
  • Forecasting, scenario analysis, and macro modelling
    • Aggregate economic modelling or scenario analysis that incorporates LLM‑derived inputs (e.g., readiness scores, compute frontier measures) risks producing biased forecasts, particularly for regional growth, technology diffusion, and labor impacts.
  • Practical remedies (research & policy actions)
    • Integrate representative, authoritative datasets (e.g., GAID) into model training and retrieval layers; expose provenance and confidence levels for numeric outputs.
    • Build automated verification modules (fact‑checking pipelines) plumbed into model response flows for numeric queries, especially for cross‑border economic indicators.
    • Incentivise data publication and capacity building in underrepresented regions (improve official statistics, patents, investment reporting).
    • Mandate geodiversity audits for commercial models used in economic decision‑making; require disclosure of per‑country coverage metrics for key technical/economic indicators.
    • Use benchmarking dashboards (as proposed by GAID’s scale‑up phases) to provide policymakers and investors with calibrated, model‑agnostic views of data coverage and uncertainty.
  • Broader economic research opportunities
    • Quantify downstream economic impact of LLM coverage gaps (e.g., changes in FDI, patent flows, startup formation attributable to model bias).
    • Study how model‑mediated information asymmetries alter market entry decisions, price discovery in venture markets, and global talent mobility.

Overall, the paper provides empirical evidence that a contemporary open‑weight foundation model exhibits substantial geoeconomic blind spots for technical AI governance metrics; for economists and policymakers, this implies both immediate risks to allocative efficiency and a policy agenda focused on data inclusion, model transparency, and verification infrastructure.

Assessment

Paper Typedescriptive Evidence Strengthlow — The study is an exploratory, cross-sectional audit of a single model (Llama-3 8B) using designed prompts; it documents patterns of responses but does not verify factual accuracy, establish causal mechanisms, or triangulate findings across models, time, or independent data sources, so conclusions about systematic global biases remain suggestive rather than robust. Methods Rigorlow — While the project covers many countries and uses a structured GAID framework, key methodological details are missing or limited: factual answers were not empirically validated, prompt design and selection procedures are not fully described, potential language/localization effects are not controlled, only one model/version was tested, and there is no robustness testing or interrater validation of coding. SampleExploratory audit using 1,704 crafted queries mapped to the Global AI Dataset (GAID) framework, covering 213 countries and eight technical metrics of AI governance/technical knowledge; responses generated from a single model (Llama-3 8B); factuality of responses was not independently verified and prompts/languages used are not fully detailed. Themesgovernance inequality GeneralizabilityResults are from a single model (Llama-3 8B) and may not hold for other LLMs, larger/smaller variants, or future versions., Prompt design, wording, and language choices may strongly influence responses; lack of documentation limits replication., Factuality of answers was not validated against external ground truth, so reported 'number/fact' rates may misstate accuracy., Temporal generalizability limited: model outputs can change with updates or fine-tuning., Coverage pertains to technical AI governance queries and may not reflect performance on other topic domains., Potential sampling bias in GAID framework and country-level indicators may affect cross-country comparisons.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The exploratory phase stress-tested Llama-3 8B with 1,704 queries across 213 countries and eight technical metrics. Other null_result audit scope (queries, countries, metrics)
Reading fidelity high
Study strength high
n=1704
0.3
The model was only able to provide number/fact responses in 11.4% of its query answers. Output Quality negative proportion of responses that were number/fact answers
Reading fidelity high
Study strength medium
n=1704
11.4%
0.18
The empirical validity of the number/fact responses reported (the 11.4%) was yet to be verified. Other null_result verification status of reported factual responses
Reading fidelity high
Study strength speculative
n=1704
0.03
AI's technical knowledge (as represented by the model's outputs) is heavily concentrated in higher-income regions. Inequality negative geographic distribution/concentration of technically relevant model outputs
Reading fidelity high
Study strength medium
n=1704
0.18
Lower-income countries from the Global South are subject to disproportionate systemic information gaps in the model's technical knowledge. Inequality negative information coverage/gaps for lower-income countries
Reading fidelity high
Study strength medium
n=1704
0.18
This disparity between the Global North and South poses concerning risks for global AI safety and inclusive governance, because policymakers in underserved regions may lack reliable data-driven insights or be misled by hallucinated facts. Governance And Regulation negative risk to policymaker decision quality and governance from unequal model outputs
Reading fidelity high
Study strength speculative
n=1704
0.03
Current AI alignment and training processes reinforce existing geoeconomic and geopolitical asymmetries. Governance And Regulation negative reinforcement of geoeconomic/geopolitical asymmetries by AI training/alignment
Reading fidelity high
Study strength speculative
n=1704
0.03
There is a need for more inclusive data representation to ensure AI serves as a truly global resource. Governance And Regulation positive inclusivity of data representation in AI training
Reading fidelity high
Study strength speculative
n=1704
0.03

Notes