The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models match experts when ranking policy criteria but disagree on scoring and final rankings in Ghana’s renewable-energy MCDA, implying LLMs are valuable for elicitation but need expert oversight and hybrid workflows to produce reliable policy recommendations.

Evaluating renewable energy development policies: A SWOT-based multi-criteria decision analysis comparison of domain experts and large language models
Jakub Więckowski, Bartosz Paradowski, Bartłomiej Kizielewicz, Mouhamed Bayane Bouraima · September 11, 2026 · Applied Soft Computing
openalex descriptive medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Jakub Więckowski provider ID
  2. Bartosz Paradowski provider ID
  3. Bartłomiej Kizielewicz provider ID
  4. Mouhamed Bayane Bouraima provider ID

Semantic Scholar

Latest observation:

  1. Jakub Wiȩckowski provider ID
  2. Bartosz Paradowski provider ID
  3. Bartłomiej Kizielewicz provider ID
  4. M. Bouraima provider ID
In a SWOT‑based MCDA for Ghanaian renewable energy policy, multiple LLMs closely reproduce expert-derived criteria weights but diverge substantially from experts on decision matrices and final policy rankings, so LLMs are useful for elicitation but require expert validation for final judgments.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Renewable energy development in emerging countries involves complex strategic decisions with long-term economic, environmental, and social implications, making the reliable evaluation of policy responses essential. Such evaluations have traditionally relied on domain expert knowledge, which is often time-consuming and resource-intensive. At the same time, Large Language Models (LLMs) are increasingly explored as potential decision-support tools due to their wide accessibility and ability to rapidly generate structured judgments. This study proposes an expert-driven SWOT-based Multi-Criteria Decision Analysis (MCDA) framework for evaluating renewable energy policy responses in Ghana and introduces a systematic comparison between domain expert judgments and LLM-generated assessments within the same decision-making structure using five selected LLMs: GPT 5.3-mini, Flash 3.0, Grok 4.2, Sonar, and Sonnet 4.6. The analysis is conducted at two stages of the MCDA pipeline: 1) criteria weighting using the Ranking Comparison (RANCOM) method and 2) policy ranking using the Measurement of Alternatives and Ranking according to Compromise Solution (MARCOS) method with aggregation via Compromise Fuzzy Ranking (CFR) into a group consensus ranking. Additionally, repeated LLM evaluations and the Risk-Informed Decision Making (RIDM) method are incorporated to assess output stability and the implications of using LLMs as supportive components in expert-driven decision processes. The results indicate a high level of agreement between LLM-generated and expert-derived criteria weights, while more pronounced differences emerge in decision matrices and resulting policy rankings. Domain experts identified increasing electricity demand as the most critical criterion, while national consensus on renewable energy implementation emerged as the most preferred policy response. LLM-based evaluations partially reproduce these priorities, although with variations in ranking structure depending on the model. The findings suggest that LLMs can effectively approximate expert judgments in selected stages of MCDA, particularly criteria weighting, while their use in complete decision pipelines requires expert validation. These results provide methodological insights into where LLMs can support expert-driven MCDA and where human expertise remains essential to ensure reliable policy recommendations.

Summary

Main Finding

LLMs can effectively approximate domain experts in selected stages of an expert-driven SWOT-based MCDA for renewable energy policy in Ghana—especially in criteria weighting—yet they produce more divergent outputs in decision matrices and final policy rankings. Consequently, LLMs are useful as decision-support tools for elicitation and preliminary scoring, but full MCDA pipelines using LLM outputs require expert validation and hybrid workflows to ensure reliable policy recommendations.

Key Points

  • Study context: renewable energy policy evaluation in Ghana using a SWOT-based Multi-Criteria Decision Analysis (MCDA) framework.
  • Models compared: GPT 5.3‑mini, Flash 3.0, Grok 4.2, Sonar, Sonnet 4.6.
  • Two MCDA stages where LLMs were evaluated:
  • Criteria weighting via Ranking Comparison (RANCOM).
  • Policy ranking via MARCOS (Measurement of Alternatives and Ranking according to Compromise Solution), aggregated using Compromise Fuzzy Ranking (CFR) for group consensus.
  • Stability and robustness checks: repeated LLM evaluations and incorporation of Risk‑Informed Decision Making (RIDM).
  • Empirical findings:
    • High agreement between LLMs and experts on criteria weights (RANCOM stage).
    • Greater divergence between LLMs and experts in decision matrices and resulting policy rankings (MARCOS stage); variation across LLMs.
    • Domain experts: top criterion = increasing electricity demand; top policy = national consensus on renewable energy implementation.
    • LLMs partially reproduce expert priorities but differ in ranking structures depending on model and repeated runs.
  • Methodological contribution: systematic, within‑pipeline comparison of expert vs. LLM judgments, plus use of CFR to form group consensus and RIDM to assess stability.

Data & Methods

  • Decision context: renewable energy policy alternatives for Ghana, evaluated using an expert-driven SWOT-based MCDA.
  • Expert input: domain experts provided criteria, weights, and policy assessments used as a baseline.
  • LLM procedure:
    • Five LLMs prompted to perform identical elicitation and scoring tasks matching the expert MCDA protocol.
    • Repeated prompts used to measure intra-model variability.
  • MCDA mechanics:
    • Criteria weighting: RANCOM (ranking-comparison method) to derive weightings from ordinal inputs.
    • Policy evaluation & ranking: MARCOS to score alternatives against criteria.
    • Aggregation: Compromise Fuzzy Ranking (CFR) to synthesize model/expert outputs into a group consensus ranking.
  • Robustness/risk treatment:
    • Repeated LLM runs and RIDM used to quantify output stability, variability, and decision risk when LLMs are part of the pipeline.
  • Evaluation metrics: correspondence between LLM and expert weights; differences in decision matrices and final policy rankings; sensitivity to repeated runs and model choice.

Implications for AI Economics

  • Practical role of LLMs in MCDA:
    • High utility for rapid, low-cost elicitation of criteria and initial weighting—can reduce expert time and standardize inputs.
    • Less reliable for full, autonomous generation of decision matrices and final rankings without expert oversight.
  • Recommended hybrid workflow:
    • Use LLMs for initial elicitation and RANCOM-based weighting, then have experts validate and adjust decision matrices and MARCOS scores.
    • Use multiple LLMs and ensemble/consensus methods (e.g., CFR) to mitigate single‑model idiosyncrasies.
    • Incorporate repeated runs and RIDM-style risk assessment to detect unstable outputs and quantify decision risk.
  • Policy and methodological safeguards:
    • Require transparent prompts, versioning, and documentation of LLM outputs for auditability.
    • Perform sensitivity analyses and expert validation on stages with high LLM–expert divergence.
    • Calibrate or fine-tune models on domain-specific data where possible to reduce variability.
  • Research directions for AI economics:
    • Test generalizability across countries, policy domains, and more/broader LLMs.
    • Evaluate cost–benefit tradeoffs of LLM-assisted workflows versus pure expert processes.
    • Develop best-practice protocols for integrating LLMs into multi-criteria policy appraisal, including interpretability and bias assessments.
  • Overall: LLMs are promising decision-support components for economic policy appraisal but should be deployed within structured, expert-supervised MCDA pipelines to preserve reliability and legitimacy of policy recommendations.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The study provides a systematic, within-pipeline comparison between multiple LLMs and domain experts with robustness checks (repeated runs, multiple models, RIDM), giving credible internal evidence about stages of MCDA where LLMs align with experts; however, it is not causal, relies on a single country/domain and an unspecified expert sample, and therefore has limited external validity. Methods Rigormedium — Good use of established MCDA methods (RANCOM for weighting, MARCOS for scoring, CFR for aggregation) and explicit robustness checks (repeats, RIDM, multi-model comparison), but key methodological details are missing or limited (size and selection of expert panel, number of alternatives/criteria, prompt specification and prompt-sensitivity analysis, potential model-version drift), reducing reproducibility and the strength of inferences. SampleDecision-context: SWOT-based MCDA evaluating renewable energy policy alternatives for Ghana. Expert baseline: domain experts supplied criteria, ordinal rankings and policy assessments (exact number and selection criteria of experts not specified). LLM sample: five models (GPT-5.3-mini, Flash 3.0, Grok 4.2, Sonar, Sonnet 4.6) prompted to replicate expert elicitation; repeated prompt runs were performed to measure intra-model variability. MCDA mechanics applied to these inputs (RANCOM for weights, MARCOS for policy scoring), with CFR used to form group consensus. Themeshuman_ai_collab governance GeneralizabilitySingle-country (Ghana) and single policy domain (renewable energy) limit geographic and sectoral generalizability, Unclear/limited expert sample size and selection may limit representativeness of the expert baseline, Results depend on the particular LLM versions and prompt designs used; model updates or different prompts could change outcomes, MCDA variant choices (RANCOM, MARCOS, CFR) may affect findings and may not generalize to other decision frameworks

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
LLMs show high agreement with domain experts on criteria weights in the RANCOM stage of the SWOT-based MCDA for renewable energy policy in Ghana. Decision Quality positive Correspondence between LLM-generated and expert-assigned criteria weights
Reading fidelity high
Study strength medium
not reported
0.18
LLM outputs diverge more from expert judgments in decision matrices and final policy rankings than in criteria weighting. Decision Quality negative Agreement between LLM and expert decision matrices and policy rankings
Reading fidelity high
Study strength medium
not reported
0.18
The extent and structure of policy-ranking variation differ across LLMs and across repeated runs of the same model. Decision Quality mixed Stability and variability of policy rankings across models and repeated evaluations
Reading fidelity high
Study strength medium
not reported
0.18
In the expert baseline, increasing electricity demand is the top criterion for evaluating renewable energy policy in Ghana. Decision Quality positive Expert-assigned priority of renewable-energy policy evaluation criteria
Reading fidelity high
Study strength medium
not reported
0.18
In the expert baseline, national consensus on renewable energy implementation is the top-ranked policy alternative. Decision Quality positive Expert ranking of renewable-energy policy alternatives
Reading fidelity high
Study strength medium
not reported
0.18
LLMs are useful for criteria elicitation and preliminary scoring in MCDA, but full MCDA pipelines using LLM outputs require expert validation and hybrid workflows. Decision Quality mixed Reliability of LLM-assisted multi-criteria policy appraisal
Reading fidelity high
Study strength medium
not reported
0.18
Repeated LLM evaluations and Risk-Informed Decision Making can be used to assess output stability, variability, and decision risk when LLMs are included in an MCDA pipeline. Ai Safety And Ethics positive Assessment of LLM-output stability and decision risk
Reading fidelity high
Study strength medium
not reported
0.18

Notes