0 cumulative citations
View corpus contextMatrAIx simulates an 8.3 billion-person world to test AI products, releasing a 1M-person coreset and an interactive Playground to run thousands of simulated-user trials; controlled validation finds persona agents follow assigned behaviors in 91.5% of trials, though results rely heavily on LLM judges and grounded sources with known demographic skews.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Summary
Main Finding
MatrAIx builds an end-to-end, population-scale simulated-user evaluation infrastructure that represents human diversity with an 8.3 billion–record persona population (Persona 8B, 1,290 categorical dimensions) and a Playground of four interactive environments (Survey, AI Chatbot, Web, App). A released quality-filtered coreset (~1M personas; ~600k human-grounded, 400k synthetic) plus 1,010 reusable task specifications enable scalable evaluation of AI systems and digital products. Across 18,189 trials (and a 400-trial controlled validation), persona agents driven by contemporary LLMs reproduce assigned behaviors at high rates (91.5% adherence) and reveal systematic heterogeneity in outcomes (e.g., price sensitivity, willingness to continue after assistant failure, latency tolerance).
Key Points
- Persona 8B: 8.3 billion persona records over a shared schema of 1,290 categorical dimensions (background, psychology, capability, behavior, lifestyle).
- Public coreset: ~1M curated personas (599,847 human-grounded; 400,000 synthetic) released for research.
- Dual persona generation:
- Synthetic: dependency-aware sampling from a directed acyclic graph (DAG) of conditional relationships and compatibility rules to preserve joint structure.
- Human-grounded: extracted and mapped from sources (Wikipedia bios, Amazon reviews, Stack Overflow Developer Survey, GSS, PRISM, consented MatrAIx survey).
- MatrAIx Playground: four environments for running simulations — Survey, AI Chatbot, Web (browser + CUAs), and App (desktop/iOS simulation).
- MatrAIx Applications: library of 1,010 task specifications across 25+ domains (Commerce, Software, Finance, Healthcare, etc.); each task includes cohort definition, scenario, verifiers, telemetry.
- Empirical evaluation: 18,189 trials across eight representative tasks; LLM-powered persona agents (Claude Opus 4.8, GPT 5.5, Claude Haiku 4.5) show behavior differences across persona groups.
- Validation:
- 400-trial controlled study: assigned behavioral attributes were expressed or correctly suppressed in 366 trials (91.5%).
- Persona extraction judged by humans (mean 4.135/5) and LLM judges; LLM judgments often close to human ratings.
- Built-in telemetry and verifiers for task-level outcome checking and cohort-level reporting.
Data & Methods
- Schema design:
- 1,290 categorical dimensions organized into five top groups: Background (238 dims), Psychology (210), Capability (331), Behavior & Interaction (124), Lifestyle (387).
- Dimensions grounded in public population and survey sources (UN population data, World Bank, ILOSTAT, World Values Survey, Stack Overflow, etc.).
- Synthetic generation:
- Dependency graph (DAG) where nodes = dimensions and edges encode conditional sampling relationships from source-informed conditionals.
- Compatibility rules to prevent implausible combinations (e.g., primary language vs. English proficiency).
- Sampling yields fully specified synthetic personas consistent with joint structure.
- Human-grounded records:
- Attribute extraction from multiple textual/data sources and mapping into the unified schema; de-identification applied (no direct identifiers).
- Curation and release:
- Contradiction checks, deduplication, calibration toward selected real-world demographic distributions.
- Public release: ~1M "coreset" on Hugging Face (link in paper).
- Simulation pipeline:
- Persona sampling according to declared target cohort.
- Persona agent instantiated by pairing persona record with an LLM; agents follow persona-conditioned prompts and internal verifier criteria.
- Environments implement domain-specific interaction modalities (surveys, multi-turn chat, browser automation, app control).
- Verifiers and telemetry collect task outcomes, completion status, rationale texts, timing, and behavioral signals.
- Validation:
- Behavioral adherence measured across 10 attributes and 4 environments (400 trials).
- Extraction quality assessed by LLM judges and human raters on subsets of human-grounded personas.
- Task demonstrations: 18,189 trials across 8 tasks, with cross-persona and cross-model comparisons.
Implications for AI Economics
- Scalable demand-side experimentation:
- MatrAIx enables rapid, low-cost simulation of large and diverse consumer cohorts for pricing, product design, and adoption studies (e.g., price-sensitivity, willingness to pay, dropout rates), reducing reliance on slow/expensive field experiments in early stages.
- Rich heterogeneity and segmentation analysis:
- The 1,290-dim schema supports fine-grained segmentation (demographics, skills, values, constraints). Economists can estimate heterogeneous treatment effects, distributional impacts, and tail behaviors that aggregate A/B tests hide.
- Counterfactual and policy analysis:
- Simulated populations allow systematic counterfactuals (e.g., varying price, latency, personalization rules) to estimate effects on welfare, consumer surplus, and inequality across subgroups before real-world deployment.
- Market design and personalized pricing:
- Use cases include testing personalized recommendations, dynamic pricing strategies, and menu design across realistic personas while measuring acceptance, churn risk, and perceived fairness for different cohorts.
- Labor and automation impact studies:
- By including capability and occupation attributes, MatrAIx can help simulate adoption of developer tools, automation assistants, or productivity features across skill levels — informing forecasts of labor demand, upskilling needs, and wage-pressure risks.
- Cost-benefit and go/no-go decisions:
- Firms and regulators can use simulated evaluations to prioritize feature rollouts, estimate likely user loss/gain, and identify subgroup-specific harms or benefits, informing incremental rollout strategies and compliance checks.
- Limitations & cautions for economic inference:
- Simulation is not a substitute for external validity checks. Estimated elasticities or welfare measures depend on persona fidelity and LLM behavioral realism; mis-specified dependencies or stereotype amplification can bias inferences.
- Necessity of calibration: recommended practice is hybrid evaluation — calibrate simulators against targeted human samples or historical A/B data, report uncertainty, and validate key findings in field trials.
- Distributional and fairness risks: synthetic generation or mapping choices can under- or over-represent vulnerable groups; economic conclusions about inequality or access must account for sample construction and de-biasing steps.
- Research and policy applications:
- Useful for pre-registration of experiments, stress-testing platform policies (privacy, consent, safety trade-offs), and regulatory impact assessments where understanding subgroup responses is critical.
- Operational recommendations for economists:
- Combine MatrAIx with limited human validation and real-world holdouts before deployment.
- Use multiple LLM agent models to assess robustness of behavior-derived estimates.
- Treat result outputs as hypothesis-generating and sensitivity-check key parameters (sampling priors, DAG edges, compatibility rules).
- Transparently report persona sampling designs and coreset calibration choices when making economic claims.
Overall, MatrAIx is a powerful infrastructure for scalable, heterogeneity-aware simulation of user interactions with AI systems and digital products — valuable for economic analysis of market responses, distributional impacts, and policy evaluation — but it requires careful calibration and validation to support causal or welfare-focused conclusions.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Persona 8B contains 8.3 billion persona records represented using a schema with 1,290 categorical dimensions. Other | positive | Scale and dimensionality of the persona population |
Reading fidelity
high
Study strength
medium
|
n=8300000000
8.3 billion records; 1,290 dimensions
|
| The released quality-filtered Persona 8B coreset contains approximately 1 million personas, including 599,847 human-grounded and 400,000 synthetic records. Other | positive | Size and composition of the released persona dataset |
Reading fidelity
high
Study strength
medium
|
n=1000000
approximately 1 million personas; 599,847 human-grounded and 400,000 synthetic
|
| MatrAIx provides four simulated-user evaluation environments: Survey, AI Chatbot, Web, and App. Other | positive | Breadth of supported simulated-user interaction environments |
Reading fidelity
high
Study strength
medium
|
n=4
four environments
|
| The MatrAIx Applications release contains 1,010 reusable tasks spanning more than 25 domains, with 621 Survey, 371 AI Chatbot, 12 Web, and 6 App tasks. Other | positive | Breadth and composition of the application-task library |
Reading fidelity
high
Study strength
medium
|
n=1010
1,010 tasks across more than 25 domains
|
| The authors conducted 18,189 evaluation trials across eight representative tasks and all four environments. Other | positive | Scale of the end-to-end simulated-user evaluation demonstrations |
Reading fidelity
high
Study strength
medium
|
n=18189
18,189 evaluation trials across eight tasks
|
| In a controlled study, persona agents expressed or correctly suppressed the assigned behavior in 366 of 400 trials, corresponding to 91.5% persona adherence. Other | positive | Persona adherence in agent behavior |
Reading fidelity
high
Study strength
medium
|
n=400
91.5% adherence; 366 of 400 trials
|
| Human judges rated the extraction quality of human-grounded personas at a mean score of 4.135 out of 5. Output Quality | positive | Human-rated persona extraction quality |
Reading fidelity
high
Study strength
low
|
n=100
mean 4.135/5
|
| GPT 5.5 scores were within one point of the human mean in 79.2% of comparisons when judging extracted persona quality. Output Quality | positive | Agreement between GPT 5.5 and human persona-quality judgments |
Reading fidelity
high
Study strength
low
|
n=100
79.2% of comparisons within one point of the human mean
|
| Claude Opus 4.8 scores were within one point of the human mean in 93.8% of comparisons when judging extracted persona quality. Output Quality | positive | Agreement between Claude Opus 4.8 and human persona-quality judgments |
Reading fidelity
high
Study strength
low
|
n=100
93.8% of comparisons within one point of the human mean
|
| MatrAIx records persona-agent thoughts, statements, actions, task duration, completion status, and verifier results across its evaluation environments. Organizational Efficiency | positive | Breadth of interaction telemetry and task-level evaluation outcomes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The system is designed to support subgroup-level comparisons by holding the target system and task fixed while varying the simulated-user population. Decision Quality | positive | Ability to compare outcomes across user groups under a fixed task and target system |
Reading fidelity
high
Study strength
medium
|
not reported
|