0 cumulative citations
View corpus contextLarge language models match experts when ranking policy criteria but disagree on scoring and final rankings in Ghana’s renewable-energy MCDA, implying LLMs are valuable for elicitation but need expert oversight and hybrid workflows to produce reliable policy recommendations.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextRenewable energy development in emerging countries involves complex strategic decisions with long-term economic, environmental, and social implications, making the reliable evaluation of policy responses essential. Such evaluations have traditionally relied on domain expert knowledge, which is often time-consuming and resource-intensive. At the same time, Large Language Models (LLMs) are increasingly explored as potential decision-support tools due to their wide accessibility and ability to rapidly generate structured judgments. This study proposes an expert-driven SWOT-based Multi-Criteria Decision Analysis (MCDA) framework for evaluating renewable energy policy responses in Ghana and introduces a systematic comparison between domain expert judgments and LLM-generated assessments within the same decision-making structure using five selected LLMs: GPT 5.3-mini, Flash 3.0, Grok 4.2, Sonar, and Sonnet 4.6. The analysis is conducted at two stages of the MCDA pipeline: 1) criteria weighting using the Ranking Comparison (RANCOM) method and 2) policy ranking using the Measurement of Alternatives and Ranking according to Compromise Solution (MARCOS) method with aggregation via Compromise Fuzzy Ranking (CFR) into a group consensus ranking. Additionally, repeated LLM evaluations and the Risk-Informed Decision Making (RIDM) method are incorporated to assess output stability and the implications of using LLMs as supportive components in expert-driven decision processes. The results indicate a high level of agreement between LLM-generated and expert-derived criteria weights, while more pronounced differences emerge in decision matrices and resulting policy rankings. Domain experts identified increasing electricity demand as the most critical criterion, while national consensus on renewable energy implementation emerged as the most preferred policy response. LLM-based evaluations partially reproduce these priorities, although with variations in ranking structure depending on the model. The findings suggest that LLMs can effectively approximate expert judgments in selected stages of MCDA, particularly criteria weighting, while their use in complete decision pipelines requires expert validation. These results provide methodological insights into where LLMs can support expert-driven MCDA and where human expertise remains essential to ensure reliable policy recommendations.
Summary
Main Finding
LLMs can effectively approximate domain experts in selected stages of an expert-driven SWOT-based MCDA for renewable energy policy in Ghana—especially in criteria weighting—yet they produce more divergent outputs in decision matrices and final policy rankings. Consequently, LLMs are useful as decision-support tools for elicitation and preliminary scoring, but full MCDA pipelines using LLM outputs require expert validation and hybrid workflows to ensure reliable policy recommendations.
Key Points
- Study context: renewable energy policy evaluation in Ghana using a SWOT-based Multi-Criteria Decision Analysis (MCDA) framework.
- Models compared: GPT 5.3‑mini, Flash 3.0, Grok 4.2, Sonar, Sonnet 4.6.
- Two MCDA stages where LLMs were evaluated:
- Criteria weighting via Ranking Comparison (RANCOM).
- Policy ranking via MARCOS (Measurement of Alternatives and Ranking according to Compromise Solution), aggregated using Compromise Fuzzy Ranking (CFR) for group consensus.
- Stability and robustness checks: repeated LLM evaluations and incorporation of Risk‑Informed Decision Making (RIDM).
- Empirical findings:
- High agreement between LLMs and experts on criteria weights (RANCOM stage).
- Greater divergence between LLMs and experts in decision matrices and resulting policy rankings (MARCOS stage); variation across LLMs.
- Domain experts: top criterion = increasing electricity demand; top policy = national consensus on renewable energy implementation.
- LLMs partially reproduce expert priorities but differ in ranking structures depending on model and repeated runs.
- Methodological contribution: systematic, within‑pipeline comparison of expert vs. LLM judgments, plus use of CFR to form group consensus and RIDM to assess stability.
Data & Methods
- Decision context: renewable energy policy alternatives for Ghana, evaluated using an expert-driven SWOT-based MCDA.
- Expert input: domain experts provided criteria, weights, and policy assessments used as a baseline.
- LLM procedure:
- Five LLMs prompted to perform identical elicitation and scoring tasks matching the expert MCDA protocol.
- Repeated prompts used to measure intra-model variability.
- MCDA mechanics:
- Criteria weighting: RANCOM (ranking-comparison method) to derive weightings from ordinal inputs.
- Policy evaluation & ranking: MARCOS to score alternatives against criteria.
- Aggregation: Compromise Fuzzy Ranking (CFR) to synthesize model/expert outputs into a group consensus ranking.
- Robustness/risk treatment:
- Repeated LLM runs and RIDM used to quantify output stability, variability, and decision risk when LLMs are part of the pipeline.
- Evaluation metrics: correspondence between LLM and expert weights; differences in decision matrices and final policy rankings; sensitivity to repeated runs and model choice.
Implications for AI Economics
- Practical role of LLMs in MCDA:
- High utility for rapid, low-cost elicitation of criteria and initial weighting—can reduce expert time and standardize inputs.
- Less reliable for full, autonomous generation of decision matrices and final rankings without expert oversight.
- Recommended hybrid workflow:
- Use LLMs for initial elicitation and RANCOM-based weighting, then have experts validate and adjust decision matrices and MARCOS scores.
- Use multiple LLMs and ensemble/consensus methods (e.g., CFR) to mitigate single‑model idiosyncrasies.
- Incorporate repeated runs and RIDM-style risk assessment to detect unstable outputs and quantify decision risk.
- Policy and methodological safeguards:
- Require transparent prompts, versioning, and documentation of LLM outputs for auditability.
- Perform sensitivity analyses and expert validation on stages with high LLM–expert divergence.
- Calibrate or fine-tune models on domain-specific data where possible to reduce variability.
- Research directions for AI economics:
- Test generalizability across countries, policy domains, and more/broader LLMs.
- Evaluate cost–benefit tradeoffs of LLM-assisted workflows versus pure expert processes.
- Develop best-practice protocols for integrating LLMs into multi-criteria policy appraisal, including interpretability and bias assessments.
- Overall: LLMs are promising decision-support components for economic policy appraisal but should be deployed within structured, expert-supervised MCDA pipelines to preserve reliability and legitimacy of policy recommendations.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| LLMs show high agreement with domain experts on criteria weights in the RANCOM stage of the SWOT-based MCDA for renewable energy policy in Ghana. Decision Quality | positive | Correspondence between LLM-generated and expert-assigned criteria weights |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLM outputs diverge more from expert judgments in decision matrices and final policy rankings than in criteria weighting. Decision Quality | negative | Agreement between LLM and expert decision matrices and policy rankings |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The extent and structure of policy-ranking variation differ across LLMs and across repeated runs of the same model. Decision Quality | mixed | Stability and variability of policy rankings across models and repeated evaluations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In the expert baseline, increasing electricity demand is the top criterion for evaluating renewable energy policy in Ghana. Decision Quality | positive | Expert-assigned priority of renewable-energy policy evaluation criteria |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In the expert baseline, national consensus on renewable energy implementation is the top-ranked policy alternative. Decision Quality | positive | Expert ranking of renewable-energy policy alternatives |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLMs are useful for criteria elicitation and preliminary scoring in MCDA, but full MCDA pipelines using LLM outputs require expert validation and hybrid workflows. Decision Quality | mixed | Reliability of LLM-assisted multi-criteria policy appraisal |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Repeated LLM evaluations and Risk-Informed Decision Making can be used to assess output stability, variability, and decision risk when LLMs are included in an MCDA pipeline. Ai Safety And Ethics | positive | Assessment of LLM-output stability and decision risk |
Reading fidelity
high
Study strength
medium
|
not reported
|