The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Generative AI in transport gives different advice to different personas, and some synthetic crash generators distort safety-relevant relationships; combined with stark demographic differences in AI attitudes, these distributional gaps argue for continuous, sensitivity-aware risk metrics instead of binary approval tiers.

Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust
Amir Rafe, Subasish Das · September 10, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Amir Rafe unresolved corpus identity
  2. Subasish Das unresolved corpus identity
This paper develops and implements a Distributional Sociotechnical Audit that shows transport-facing generative AI produces substantial persona-driven variation in advice, some classical synthetic crash-data generators fail to preserve conditional structure, and public AI attitudes vary sharply across demographics, which together motivate a continuous, sensitivity-bounded Sociotechnical Risk Index for governance.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Generative artificial intelligence is entering transportation through traveler-facing advisories, synthetic crash-record generation, and policy decision support. Existing governance frameworks lack transport-specific statistical tools to measure distributional risks across heterogeneous populations. We develop a Distributional Sociotechnical Audit (DSA) that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline. The audit analyzes 5,760 persona-controlled queries to four LLM families across 12 demographic cues and four transport topics, uses two cross-family judges and a Wasserstein-2 Equity Dispersion Index, tests three FARS crash-record generators with conditional projected maximum mean discrepancy (cpMMD), fits a Bayesian ordered-logit model to Pew American Trends Panel Wave 152 (N = 4,538), and combines the signals into a continuous Sociotechnical Risk Index. Congestion-pricing advice has the highest persona-based dispersion (mean EDI = 1.96; highest direct EDI = 2.20). CART synthetic crash records fail all conditional tests (p < 0.001), while the Gaussian copula has borderline conditional stress (p = 0.105) despite passing marginal checks. Attitudes to AI vary across demographic strata. Distributional audits and continuous risk indices with sensitivity reporting offer a more defensible basis for transport GenAI governance than categorical approval tiers, which show a 75% assignment flip rate under weight perturbation.

Summary

Main Finding

Generative AI systems entering transport produce measurable, population-differentiated risks across three linked layers—LLM outputs, synthetic crash-data products, and public attitudes—and these distributional signals can and should be combined into a continuous, sensitivity-bounded governance metric. Key empirical results: (1) persona cues induce sizable variation in LLM transport advice (mean Equity Dispersion Index (EDI) = 1.96; max cell Gemini Flash direct EDI = 2.20), with policy-contested topics (e.g., congestion pricing) showing up to 1.6× more persona-driven variation than low-contest topics (e.g., weather-safety); (2) common synthetic crash-data generators can fail to preserve conditional structure—CART fails conditional distributional tests (p < 0.001) while a Gaussian-copula baseline passes marginal checks but is borderline on conditional checks (p = 0.105); (3) public attitudes toward AI are heterogeneous across strata (Bayesian posterior |β_k| range 0.002–0.860); (4) when combined into a Sociotechnical Risk Index (STRI) with perturbation bounds, continuous risk measurement is more robust than categorical readiness tiers (categorical tiers flipped 75% under weight perturbation).

Key Points

  • Distributional risk matters more than aggregate performance: comparable semantic queries differing only by persona cues produce systematically different LLM outputs on transport topics.
  • The Wasserstein-2 Equity Dispersion Index (EDI) is used to compare full-score distributions across persona groups rather than reliance on means or keyword counts.
  • Synthetic-data validity requires conditional (not only marginal) checks; failing conditional fidelity can produce silent distortions important to safety and regulatory analysis.
  • Public acceptance and trust are stratified; aggregate readiness scores risk masking groups with substantially different attitudes toward GenAI.
  • Continuous composite indices with explicit sensitivity reporting (STRI + Weyl perturbation bounds) provide a more defensible governance signal than rigid categorical tiers.

Data & Methods

  • LLM audit
    • 5,760 persona-controlled queries: 4 LLM families (GPT-5.4 Nano; Claude Haiku 4.5; Gemini 3.1 Flash Lite; Mistral Nemo 12B)
    • 12 demographic persona cues × 4 transport topics × 30 repetitions
    • Responses scored on 8 content axes by two cross-family LLM judges
    • Equity Dispersion Index (EDI): pairwise Wasserstein-2 distances across persona-conditioned rubric-score distributions
  • Synthetic-data audit
    • Authentic FARS crash data vs. three classical generators (Gaussian copula, sequential CART, perturbation baseline)
    • 110,001 synthetic records total
    • Conditional projected maximum mean discrepancy (cpMMD) testing to assess conditional distributional fidelity (marginal checks plus conditional dependence)
    • Findings: CART fails conditional tests (p < 0.001); Gaussian copula borderline (p = 0.105)
  • Public-attitude analysis
    • Pew American Trends Panel Wave 152 (N = 4,538; Aug 12–18, 2024)
    • Bayesian ordered-logit model with horseshoe priors estimating stratum-level effects (age, gender, metropolitan status, race)
    • Posterior |β_k| range: 0.002 to 0.860, indicating substantial heterogeneity
  • Governance synthesis
    • Sociotechnical Risk Index (STRI): continuous composite that integrates EDI (equity), cpMMD (synthetic-data stress), and attitude heterogeneity
    • STRI sensitivity analyzed with Weyl perturbation bounds; categorical Regulatory Readiness Ladder (RRL) tiers shown to be unstable (75% assignment flip under weight perturbation)

Implications for AI Economics

  • Welfare and distributional externalities
    • Persona-differentiated advice implies that AI-mediated information can redistribute welfare (and harm) across demographic groups even without physical infrastructure changes. Economic evaluations (cost-benefit, welfare analysis) must incorporate distributional impacts, not just mean benefits.
  • Market design and demand heterogeneity
    • Heterogeneous attitudes toward AI imply segmented adoption curves and price sensitivities across demographics. Firms and regulators should anticipate non-uniform uptake, affecting demand forecasts, pricing strategies for mobility services, and the design of targeted interventions/subsidies.
  • Insurance, liability, and risk pricing
    • If synthetic data used in safety analyses are conditionally distorted (e.g., CART failures), actuarial estimates and liability models relying on synthetic records may be biased. Insurers and regulators need validated conditional fidelity checks before using synthetic datasets for premium setting or regulatory compliance.
  • Regulatory policy and compliance metrics
    • Continuous, sensitivity-bounded risk indices (STRI) are preferable to categorical pass/fail tiers for regulating GenAI in transport. Economic regulation that uses brittle categorical thresholds risks misclassification, per the 75% flip-rate under perturbation; welfare-improving regulation should incorporate robustness/sensitivity analysis and potentially use continuous taxes/subsidies or graded obligations.
  • Investment and procurement decisions
    • Public agencies procuring GenAI tools should require distributional audits (EDI) and synthetic-data validation (cpMMD) as part of procurement KPIs. From an economic perspective, this raises the transaction cost of deployment and creates new markets for audit/validation services.
  • Research and model-use externalities
    • Policy simulations and transport-economic models that rely on synthetic microdata must incorporate uncertainty from conditional mismatch. Analysts should propagate cpMMD-detected discrepancies into counterfactual estimates and confidence intervals to avoid overconfident policy recommendations.
  • Incentives for model developers
    • Developers face externalities: generating outputs that vary by persona can create reputational, legal, and market risks. Economic incentives (contractual clauses, liability exposure, certification premiums) should be aligned to reward distributional fidelity and robust synthetic-data generation.
  • Measurement recommendations for AI economists
    • Adopt distributional testing (Wasserstein/optimal-transport metrics), conditional two-sample tests for synthetic data, and stratified attitude measures when modeling adoption or welfare; embed sensitivity bounds in policy simulations and cost-benefit analyses.

Limitations to keep in mind: the audit is descriptive (not causal), limited to the selected LLM families, topics, and U.S.-centric data; synthetic generators evaluated are classical baselines (not state-of-the-art GAN/LLM datagenerators), so results indicate risk pathways rather than a universal verdict.

If useful, I can convert these findings into recommended checklist items for transport-sector economic evaluations (e.g., required tests, audit contract language, or templates for embedding STRI-like uncertainty in cost-benefit models).

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper uses large samples (5,760 LLM queries; 110,001 synthetic records; N=4,538 survey respondents) and appropriate statistical tools (Wasserstein EDI, cpMMD, Bayesian ordered-logit). However, it is descriptive rather than causal, evaluates a limited set of LLM families and synthetic baselines, relies on LLM-based cross-family judges for rubric scoring (potential circularity), and is US/FARS-centric, which constrain the strength of policy-generalizable claims. Methods Rigorhigh — The methods combine principled distributional comparisons (Wasserstein-2 EDI), conditional kernel testing (cpMMD), and a Bayesian hierarchical ordered-logit with regularizing horseshoe priors; they further synthesize signals into a perturbation-bounded composite (STRI) with sensitivity bounds. Execution appears careful, but some components (LLM judges, choice of only classical synthetic baselines) introduce practical limitations. SampleThree data streams: (1) LLM audit corpus — 5,760 persona-controlled queries across four LLM families (GPT-5.4 Nano, Claude Haiku 4.5, Gemini 3.1 Flash Lite, Mistral Nemo 12B), 12 demographic persona cues, 4 transport topics, 30 repetitions; responses scored on eight content axes by two cross-family judges. (2) Synthetic-data audit — 110,001 records generated by three classical FARS-like generators (Gaussian copula, sequential CART, perturbation baseline) compared against authentic FARS crash records. (3) Public-attitude survey — Pew American Trends Panel Wave 152, N = 4,538, analyzed with a Bayesian ordered-logit model. Themesgovernance adoption inequality human_ai_collab IdentificationNo causal identification for AI effects; uses a multi-method descriptive/audit design: persona-controlled LLM probes to estimate distributional differences across demographic cues (counterfactual-style comparisons), conditional projected MMD (cpMMD) to test conditional distributional fidelity of synthetic crash generators versus authentic FARS data, and a Bayesian ordered-logit on cross-sectional Pew ATP Wave 152 to describe public-attitude heterogeneity (associational). GeneralizabilityLimited to the four evaluated LLM families and specific model versions — results may not hold for other or future models., Persona-driven probe design may not capture the full complexity of real-world user queries and identities., Synthetic-data findings are tied to the chosen classical generators; modern deep generative models (GANs/VAEs/transformer-based tabular synthesizers) were not evaluated., FARS is U.S.-centric; transport and crash-structure differ internationally., Pew ATP is cross-sectional and observational — associations in attitudes are non-causal and may not generalize beyond the sampled period or population., Reliance on LLM-based judges for rubric scoring risks bias or tautology in evaluating LLM outputs.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Congestion-pricing advice produced the highest distributional dispersion across demographic personas in the LLM audit. Ai Safety And Ethics negative Distributional variation in LLM-generated transport advice across demographic personas
Reading fidelity high
Study strength medium
n=5760
mean EDI = 1.96
0.18
The largest observed persona-related dispersion occurred for Gemini Flash's direct congestion-pricing advice, with an EDI of 2.20. Ai Safety And Ethics negative Persona-based dispersion in direct LLM transport advice
Reading fidelity high
Study strength medium
n=5760
direct EDI = 2.20
0.18
Policy-contested transport topics generated as much as 1.6 times more persona-driven variation than weather-safety advice. Ai Safety And Ethics negative Relative persona-driven variation in LLM advice by transport topic
Reading fidelity high
Study strength medium
n=5760
up to 1.6 times greater persona-driven variation
0.18
The CART-based synthetic crash-record generators failed all conditional distributional validity tests. Ai Safety And Ethics negative Conditional distributional validity of synthetic crash records
Reading fidelity high
Study strength high
n=110001
p < 0.001
0.3
The Gaussian-copula synthetic crash generator passed marginal checks but showed borderline conditional stress. Ai Safety And Ethics mixed Marginal and conditional distributional validity of Gaussian-copula synthetic crash records
Reading fidelity high
Study strength medium
n=110001
conditional stress p = 0.105
0.18
General AI attitudes vary heterogeneously across demographic groups in the Pew American Trends Panel Wave 152. Ai Safety And Ethics mixed Self-reported general attitudes toward AI across demographic strata
Reading fidelity high
Study strength medium
n=4538
posterior |β̂k| range: 0.002 to 0.860
0.18
A continuous Sociotechnical Risk Index provides a more defensible governance signal than categorical approval tiers under weight uncertainty. Governance And Regulation positive Robustness and defensibility of governance risk assessment
Reading fidelity high
Study strength medium
75% assignment flip rate under weight perturbation
0.18
The integrated Distributional Sociotechnical Audit demonstrates that auditing model outputs, synthetic data products, and public attitudes together is feasible and necessary for transport GenAI governance. Governance And Regulation positive Feasibility and governance relevance of integrated distributional auditing
Reading fidelity high
Study strength medium
not reported
0.18

Notes