The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

MatrAIx simulates an 8.3 billion-person world to test AI products, releasing a 1M-person coreset and an interactive Playground to run thousands of simulated-user trials; controlled validation finds persona agents follow assigned behaviors in 91.5% of trials, though results rely heavily on LLM judges and grounded sources with known demographic skews.

MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song · August 04, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Xiaomin Li unresolved corpus identity
  2. Yuexing Hao unresolved corpus identity
  3. Jianheng Hou unresolved corpus identity
  4. Jintao Huang unresolved corpus identity
  5. Qianfeng Wen unresolved corpus identity
  6. Shirley Huang unresolved corpus identity
  7. Yifan Liu unresolved corpus identity
  8. Xiaoyi Liu unresolved corpus identity
  9. Yilan Fan unresolved corpus identity
  10. Yijun Wang unresolved corpus identity
  11. Koutian Wu unresolved corpus identity
  12. Ruoqi Gao unresolved corpus identity
  13. Muhammad Ahmed Mohsin unresolved corpus identity
  14. Jing Tang unresolved corpus identity
  15. Brihi Joshi unresolved corpus identity
  16. Heming Liu unresolved corpus identity
  17. Zheyuan Deng unresolved corpus identity
  18. Zonglin Di unresolved corpus identity
  19. Sankalp Jajee unresolved corpus identity
  20. Jiuyao Lu unresolved corpus identity
  21. Zhiwei Zhang unresolved corpus identity
  22. Saksham Kapoor unresolved corpus identity
  23. Ishan Gupta unresolved corpus identity
  24. Yunhan Zhao unresolved corpus identity
  25. Chanwoo Park unresolved corpus identity
  26. Yucheng Lu unresolved corpus identity
  27. Bing Hu unresolved corpus identity
  28. Weihang Xiao unresolved corpus identity
  29. Aravind Mohan unresolved corpus identity
  30. Hanwen Xing unresolved corpus identity
  31. Runyu Zhang unresolved corpus identity
  32. Mihir Kulshreshtha unresolved corpus identity
  33. Yuanda Xu unresolved corpus identity
  34. Qianyu Zhu unresolved corpus identity
  35. Dianzhuo Wang unresolved corpus identity
  36. Yuxin Xiao unresolved corpus identity
  37. Bowen Jiang unresolved corpus identity
  38. Yongye Su unresolved corpus identity
  39. Wenhao Chai unresolved corpus identity
  40. Zuxin Liu unresolved corpus identity
  41. Lawrence Yunliang Chen unresolved corpus identity
  42. Xuandong Zhao unresolved corpus identity
  43. Ethan Ye unresolved corpus identity
  44. Shivam Patel unresolved corpus identity
  45. Jason Xie unresolved corpus identity
  46. Alex Martin Richmond unresolved corpus identity
  47. Weixiang Ding unresolved corpus identity
  48. Emre Okcular unresolved corpus identity
  49. Diya Mathew unresolved corpus identity
  50. Ziheng Wang unresolved corpus identity
  51. Rana M. Shahroz Khan unresolved corpus identity
  52. Zhejian Peng unresolved corpus identity
  53. Fang Wu unresolved corpus identity
  54. Fan Nie unresolved corpus identity
  55. Xinyang Han unresolved corpus identity
  56. Yubin Kim unresolved corpus identity
  57. Jiawei Zhang unresolved corpus identity
  58. Zhenting Qi unresolved corpus identity
  59. Huangyuan Su unresolved corpus identity
  60. Xu Pan unresolved corpus identity
  61. Abinitha Gourabathina unresolved corpus identity
  62. Hyewon Jeong unresolved corpus identity
  63. Hemanth Neelgund Ramesh unresolved corpus identity
  64. Kumail Alhamoud unresolved corpus identity
  65. Kimia Hamidieh unresolved corpus identity
  66. Zidi Xiong unresolved corpus identity
  67. Samuel Schmidgall unresolved corpus identity
  68. Pengrui Han unresolved corpus identity
  69. Yepeng Huang unresolved corpus identity
  70. Yongheng Wang unresolved corpus identity
  71. Bowen Yang unresolved corpus identity
  72. Alex Gu unresolved corpus identity
  73. Yuchu Wang unresolved corpus identity
  74. Akshay Paruchuri unresolved corpus identity
  75. Brenna Li unresolved corpus identity
  76. Hejie Cui unresolved corpus identity
  77. Jiayuan Ding unresolved corpus identity
  78. Chaosheng Dong unresolved corpus identity
  79. Jiahao Wang unresolved corpus identity
  80. Yixuan He unresolved corpus identity
  81. Chi Wang unresolved corpus identity
  82. Pamela Bhattacharya unresolved corpus identity
  83. Tianyi Peng unresolved corpus identity
  84. Paul Pu Liang unresolved corpus identity
  85. Mitchell Gordon unresolved corpus identity
  86. Yilun Du unresolved corpus identity
  87. Marinka Zitnik unresolved corpus identity
  88. James Zou unresolved corpus identity
  89. Prasanna Tambe unresolved corpus identity
  90. Philip Torr unresolved corpus identity
  91. Emily Fox unresolved corpus identity
  92. Asu Ozdaglar unresolved corpus identity
  93. Dawn Song unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Xiaomin Li provider ID
  2. Yuexing Hao provider ID
  3. Jian Hou provider ID
  4. Jintao Huang provider ID
  5. Qianfeng Wen provider ID
  6. Shirley Huang provider ID
  7. Yifan Liu provider ID
  8. Xiaoyi Liu provider ID
  9. Yi-Mei Fan provider ID
  10. Yijun Wang provider ID
  11. Ko-Hui Wu provider ID
  12. Ru-Ping Gao provider ID
  13. Muhammad Ahmed Mohsin provider ID
  14. Jing Tang provider ID
  15. Brihi Joshi provider ID
  16. Heming Liu provider ID
  17. Zheyuan Deng provider ID
  18. Zonglin Di provider ID
  19. Sankalp Jajee provider ID
  20. Jiuyao Lu provider ID
  21. Zhiwei Zhang provider ID
  22. S. Kapoor provider ID
  23. Isha Gupta provider ID
  24. Yunhan Zhao provider ID
  25. Chanwoo Park provider ID
  26. Yucheng Lu provider ID
  27. Bing Hu provider ID
  28. Wei-Wei Xiao provider ID
  29. Aravind Mohan provider ID
  30. Hanwen Xing provider ID
  31. Run Zhang provider ID
  32. Mihir Kulshreshtha provider ID
  33. Yuanda Xu provider ID
  34. Qian Zhu provider ID
  35. Dianzhuo Wang provider ID
  36. Yuxin Xiao provider ID
  37. Bowen Jiang provider ID
  38. Yongye Su provider ID
  39. Wenhao Chai provider ID
  40. Zuxin Liu provider ID
  41. L. Chen provider ID
  42. Xuandong Zhao provider ID
  43. Ethan Ye provider ID
  44. Shivam Patel provider ID
  45. Jason Xie provider ID
  46. A.M. Richmond provider ID
  47. Wei Ding provider ID
  48. Emre Okçular provider ID
  49. D. Mathew provider ID
  50. Ziheng Wang provider ID
  51. Rana Muhammad Shahroz Khan provider ID
  52. Zhe Peng provider ID
  53. Fang Wu provider ID
  54. Fan Nie provider ID
  55. Xinyang Han provider ID
  56. Y. Kim provider ID
  57. Jiawei Zhang provider ID
  58. Zhenting Qi provider ID
  59. Huangyuan Su provider ID
  60. Xue-Wei Pan provider ID
  61. Abinitha Gourabathina provider ID
  62. H. Jeong provider ID
  63. H. Ramesh provider ID
  64. K. Alhamoud provider ID
  65. K. Hamidieh provider ID
  66. Zidi Xiong provider ID
  67. S. Schmidgall provider ID
  68. Pengrui Han provider ID
  69. Yepeng Huang provider ID
  70. Yonghe Wang provider ID
  71. Bowen Yang unresolved corpus identity
  72. Alex Gu provider ID
  73. Yuchuan Wang provider ID
  74. Akshay Paruchuri provider ID
  75. Brenna Li provider ID
  76. Hejie Cui provider ID
  77. Jiayuan Ding provider ID
  78. Chaosheng Dong provider ID
  79. Jiahao Wang provider ID
  80. Yixuan He provider ID
  81. Ching-Ling Wang provider ID
  82. P. Bhattacharya provider ID
  83. Tianyi Peng provider ID
  84. P. Liang provider ID
  85. Mitchell Gordon provider ID
  86. Yilun Du provider ID
  87. M. Zitnik provider ID
  88. James Zou provider ID
  89. Prasanna B. Tambe provider ID
  90. Philip H. S. Torr provider ID
  91. Emily Fox provider ID
  92. Asu Ozdaglar provider ID
  93. Dawn Song provider ID
MatrAIx provides a population-scale persona dataset (8.3B records, 1,290 categorical dimensions) plus a Playground to run simulated-user evaluations across Survey, Chatbot, Web, and App environments, validated via 18,189 trials and a 400-trial controlled study showing 91.5% persona adherence in expressed/suppressed behaviors.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

Summary

Main Finding

MatrAIx builds an end-to-end, population-scale simulated-user evaluation infrastructure that represents human diversity with an 8.3 billion–record persona population (Persona 8B, 1,290 categorical dimensions) and a Playground of four interactive environments (Survey, AI Chatbot, Web, App). A released quality-filtered coreset (~1M personas; ~600k human-grounded, 400k synthetic) plus 1,010 reusable task specifications enable scalable evaluation of AI systems and digital products. Across 18,189 trials (and a 400-trial controlled validation), persona agents driven by contemporary LLMs reproduce assigned behaviors at high rates (91.5% adherence) and reveal systematic heterogeneity in outcomes (e.g., price sensitivity, willingness to continue after assistant failure, latency tolerance).

Key Points

  • Persona 8B: 8.3 billion persona records over a shared schema of 1,290 categorical dimensions (background, psychology, capability, behavior, lifestyle).
  • Public coreset: ~1M curated personas (599,847 human-grounded; 400,000 synthetic) released for research.
  • Dual persona generation:
    • Synthetic: dependency-aware sampling from a directed acyclic graph (DAG) of conditional relationships and compatibility rules to preserve joint structure.
    • Human-grounded: extracted and mapped from sources (Wikipedia bios, Amazon reviews, Stack Overflow Developer Survey, GSS, PRISM, consented MatrAIx survey).
  • MatrAIx Playground: four environments for running simulations — Survey, AI Chatbot, Web (browser + CUAs), and App (desktop/iOS simulation).
  • MatrAIx Applications: library of 1,010 task specifications across 25+ domains (Commerce, Software, Finance, Healthcare, etc.); each task includes cohort definition, scenario, verifiers, telemetry.
  • Empirical evaluation: 18,189 trials across eight representative tasks; LLM-powered persona agents (Claude Opus 4.8, GPT 5.5, Claude Haiku 4.5) show behavior differences across persona groups.
  • Validation:
    • 400-trial controlled study: assigned behavioral attributes were expressed or correctly suppressed in 366 trials (91.5%).
    • Persona extraction judged by humans (mean 4.135/5) and LLM judges; LLM judgments often close to human ratings.
  • Built-in telemetry and verifiers for task-level outcome checking and cohort-level reporting.

Data & Methods

  • Schema design:
    • 1,290 categorical dimensions organized into five top groups: Background (238 dims), Psychology (210), Capability (331), Behavior & Interaction (124), Lifestyle (387).
    • Dimensions grounded in public population and survey sources (UN population data, World Bank, ILOSTAT, World Values Survey, Stack Overflow, etc.).
  • Synthetic generation:
    • Dependency graph (DAG) where nodes = dimensions and edges encode conditional sampling relationships from source-informed conditionals.
    • Compatibility rules to prevent implausible combinations (e.g., primary language vs. English proficiency).
    • Sampling yields fully specified synthetic personas consistent with joint structure.
  • Human-grounded records:
    • Attribute extraction from multiple textual/data sources and mapping into the unified schema; de-identification applied (no direct identifiers).
  • Curation and release:
    • Contradiction checks, deduplication, calibration toward selected real-world demographic distributions.
    • Public release: ~1M "coreset" on Hugging Face (link in paper).
  • Simulation pipeline:
    • Persona sampling according to declared target cohort.
    • Persona agent instantiated by pairing persona record with an LLM; agents follow persona-conditioned prompts and internal verifier criteria.
    • Environments implement domain-specific interaction modalities (surveys, multi-turn chat, browser automation, app control).
    • Verifiers and telemetry collect task outcomes, completion status, rationale texts, timing, and behavioral signals.
  • Validation:
    • Behavioral adherence measured across 10 attributes and 4 environments (400 trials).
    • Extraction quality assessed by LLM judges and human raters on subsets of human-grounded personas.
    • Task demonstrations: 18,189 trials across 8 tasks, with cross-persona and cross-model comparisons.

Implications for AI Economics

  • Scalable demand-side experimentation:
    • MatrAIx enables rapid, low-cost simulation of large and diverse consumer cohorts for pricing, product design, and adoption studies (e.g., price-sensitivity, willingness to pay, dropout rates), reducing reliance on slow/expensive field experiments in early stages.
  • Rich heterogeneity and segmentation analysis:
    • The 1,290-dim schema supports fine-grained segmentation (demographics, skills, values, constraints). Economists can estimate heterogeneous treatment effects, distributional impacts, and tail behaviors that aggregate A/B tests hide.
  • Counterfactual and policy analysis:
    • Simulated populations allow systematic counterfactuals (e.g., varying price, latency, personalization rules) to estimate effects on welfare, consumer surplus, and inequality across subgroups before real-world deployment.
  • Market design and personalized pricing:
    • Use cases include testing personalized recommendations, dynamic pricing strategies, and menu design across realistic personas while measuring acceptance, churn risk, and perceived fairness for different cohorts.
  • Labor and automation impact studies:
    • By including capability and occupation attributes, MatrAIx can help simulate adoption of developer tools, automation assistants, or productivity features across skill levels — informing forecasts of labor demand, upskilling needs, and wage-pressure risks.
  • Cost-benefit and go/no-go decisions:
    • Firms and regulators can use simulated evaluations to prioritize feature rollouts, estimate likely user loss/gain, and identify subgroup-specific harms or benefits, informing incremental rollout strategies and compliance checks.
  • Limitations & cautions for economic inference:
    • Simulation is not a substitute for external validity checks. Estimated elasticities or welfare measures depend on persona fidelity and LLM behavioral realism; mis-specified dependencies or stereotype amplification can bias inferences.
    • Necessity of calibration: recommended practice is hybrid evaluation — calibrate simulators against targeted human samples or historical A/B data, report uncertainty, and validate key findings in field trials.
    • Distributional and fairness risks: synthetic generation or mapping choices can under- or over-represent vulnerable groups; economic conclusions about inequality or access must account for sample construction and de-biasing steps.
  • Research and policy applications:
    • Useful for pre-registration of experiments, stress-testing platform policies (privacy, consent, safety trade-offs), and regulatory impact assessments where understanding subgroup responses is critical.
  • Operational recommendations for economists:
    • Combine MatrAIx with limited human validation and real-world holdouts before deployment.
    • Use multiple LLM agent models to assess robustness of behavior-derived estimates.
    • Treat result outputs as hypothesis-generating and sensitivity-check key parameters (sampling priors, DAG edges, compatibility rules).
    • Transparently report persona sampling designs and coreset calibration choices when making economic claims.

Overall, MatrAIx is a powerful infrastructure for scalable, heterogeneity-aware simulation of user interactions with AI systems and digital products — valuable for economic analysis of market responses, distributional impacts, and policy evaluation — but it requires careful calibration and validation to support causal or welfare-focused conclusions.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper is an infrastructure/dataset and validation study rather than a causal empirical paper. It reports substantial empirical validation (18,189 trials across eight tasks, a 400-trial controlled adherence study, and judge/human ratings), which supports claims about persona fidelity and evaluation throughput, but not claims about real-world behavioral causality or economic impacts; validation relies heavily on LLM judges and a limited human-evaluation subset. Methods Rigormedium — Design uses principled methods (1,290-dim schema, dependency-aware DAG sampling, multiple grounded sources, task verifiers, and multi-environment execution) and reports controlled validation; however, key limitations include reliance on LLM judges for large-scale evaluation, modest human-annotation coverage (100 human-rated personas; subset human trials), potential grounding/source biases (Wikipedia, Amazon reviews, StackOverflow), and limited ecological validation across only eight executed tasks. SamplePersona 8B: an 8.3 billion-record persona population encoded with a shared 1,290-dimensional categorical schema; a public, quality-filtered coreset of ~1,000,000 personas (599,847 human-grounded from sources including Wikipedia biographies, Amazon reviews, Stack Overflow Developer Survey, General Social Survey, PRISM profiles, and consented MatrAIx surveys; and 400,000 synthetic records generated via dependency-graph sampling and compatibility rules). Validation/runtime experiments: 18,189 evaluation trials across eight representative tasks in four environments (Survey, AI Chatbot, Web, App); 400-trial controlled persona-adherence study; model agents powered by Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5; judge evaluations from LLMs and a small human subset. Themeshuman_ai_collab adoption GeneralizabilityHuman-grounded sources are biased (Wikipedia, Amazon reviews, StackOverflow skew toward particular demographics and professions), limiting population representativeness., Synthetic generation depends on chosen priors, dependency graph, and compatibility rules which may not reflect real joint distributions in all settings., Validation covers eight tasks and a 1M coreset but may not generalize to all product types, cultures, languages, or low-resource populations., Heavy reliance on LLM judges could reproduce model-specific biases; limited human annotation constrains external validity., Behavioral fidelity in simulation may not transfer perfectly to real-user dynamics (e.g., emotional responses, long-term adoption).

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Persona 8B contains 8.3 billion persona records represented using a schema with 1,290 categorical dimensions. Other positive Scale and dimensionality of the persona population
Reading fidelity high
Study strength medium
n=8300000000
8.3 billion records; 1,290 dimensions
0.18
The released quality-filtered Persona 8B coreset contains approximately 1 million personas, including 599,847 human-grounded and 400,000 synthetic records. Other positive Size and composition of the released persona dataset
Reading fidelity high
Study strength medium
n=1000000
approximately 1 million personas; 599,847 human-grounded and 400,000 synthetic
0.18
MatrAIx provides four simulated-user evaluation environments: Survey, AI Chatbot, Web, and App. Other positive Breadth of supported simulated-user interaction environments
Reading fidelity high
Study strength medium
n=4
four environments
0.18
The MatrAIx Applications release contains 1,010 reusable tasks spanning more than 25 domains, with 621 Survey, 371 AI Chatbot, 12 Web, and 6 App tasks. Other positive Breadth and composition of the application-task library
Reading fidelity high
Study strength medium
n=1010
1,010 tasks across more than 25 domains
0.18
The authors conducted 18,189 evaluation trials across eight representative tasks and all four environments. Other positive Scale of the end-to-end simulated-user evaluation demonstrations
Reading fidelity high
Study strength medium
n=18189
18,189 evaluation trials across eight tasks
0.18
In a controlled study, persona agents expressed or correctly suppressed the assigned behavior in 366 of 400 trials, corresponding to 91.5% persona adherence. Other positive Persona adherence in agent behavior
Reading fidelity high
Study strength medium
n=400
91.5% adherence; 366 of 400 trials
0.18
Human judges rated the extraction quality of human-grounded personas at a mean score of 4.135 out of 5. Output Quality positive Human-rated persona extraction quality
Reading fidelity high
Study strength low
n=100
mean 4.135/5
0.09
GPT 5.5 scores were within one point of the human mean in 79.2% of comparisons when judging extracted persona quality. Output Quality positive Agreement between GPT 5.5 and human persona-quality judgments
Reading fidelity high
Study strength low
n=100
79.2% of comparisons within one point of the human mean
0.09
Claude Opus 4.8 scores were within one point of the human mean in 93.8% of comparisons when judging extracted persona quality. Output Quality positive Agreement between Claude Opus 4.8 and human persona-quality judgments
Reading fidelity high
Study strength low
n=100
93.8% of comparisons within one point of the human mean
0.09
MatrAIx records persona-agent thoughts, statements, actions, task duration, completion status, and verifier results across its evaluation environments. Organizational Efficiency positive Breadth of interaction telemetry and task-level evaluation outcomes
Reading fidelity high
Study strength medium
not reported
0.18
The system is designed to support subgroup-level comparisons by holding the target system and task fixed while varying the simulated-user population. Decision Quality positive Ability to compare outcomes across user groups under a fixed task and target system
Reading fidelity high
Study strength medium
not reported
0.18

Notes