The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Academic research on AI credit scoring delivers strong predictive results but treats fairness and explainability as separate issues, leaving scarce guidance for regulated, real‑world use; the literature rarely evaluates integrated approaches that meet oversight and accountability requirements.

Performance, Fairness, and Explainability in AI-Based Credit Scoring: A Systematic Literature Review
Rashed Bahlool, Nabil Hewahi, Wael Elmedany · February 03, 2026 · Journal of risk and financial management
openalex review_meta n/a evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Rashed Bahlool provider ID
  2. Nabil Hewahi provider ID
  3. Wael Elmedany provider ID

Semantic Scholar

Latest observation:

  1. Rashed Bahlool provider ID
  2. Nabil Hewahi provider ID
  3. Wael Elmedany provider ID
A systematic review of 43 studies finds AI credit scoring research typically treats predictive performance, fairness, and explainability separately, with little integrated work or evidence on regulation‑ready, human‑oversight deployments.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The integration of artificial intelligence (AI) in the financial sector has seen a rapid increase over the past few years, offering new possibilities to streamline processes while ensuring profitability for lending institutions. With its data-driven capability, predicting the creditworthiness of applicants has demonstrated strong predictive performance, particularly for thin-file clients. Despite these advances, growing concerns regarding AI’s fairness, explainability, and regulatory accountability have increasingly limited its adoption in high-stakes credit decision-making. This paper presents a synthesis derived from a systematic literature review (SLR) of 43 peer-reviewed studies published between 2020 and 2025, focusing on AI-based credit scoring and addressing at least one of the performance, fairness, or explainability dimensions. Eligible studies were limited to peer-reviewed journal and conference articles (2020–2025) retrieved from IEEE Xplore, Scopus, Web of Science, and ScienceDirect (last searched: 30 September), examining AI-driven credit scoring in consumer or lending decision contexts. Guided by the Relevance, Rigor, Reproducibility, and Quality (3Rs&Q) appraisal framework, the review analyzes how existing approaches navigate the interplay among performance, fairness, and explainability under regulatory and human oversight considerations. The findings indicate that these dimensions are predominantly addressed in isolation, with limited attention to their joint treatment in regulated deployment settings. By consolidating empirical and conceptual evidence, this review provides actionable guidance for designing and deploying credit scoring models in practice.

Summary

Main Finding

A systematic literature review of 43 peer‑reviewed studies (2020–2025) finds that AI models for credit scoring deliver strong predictive gains—notably for thin‑file applicants—but research and practice largely treat predictive performance, fairness, and explainability as separate objectives. There is limited evidence on methods and governance that jointly optimize these dimensions for deployment in regulated, high‑stakes lending contexts.

Key Points

  • Scope and evidence base
    • 43 peer‑reviewed journal and conference articles (published 2020–2025), retrieved from IEEE Xplore, Scopus, Web of Science, and ScienceDirect (last searched 30 September 2025).
    • Inclusion required treatment of at least one of: performance, fairness, explainability in AI‑based credit scoring.
  • Core empirical finding
    • AI approaches improve predictive accuracy, especially for thin‑file or sparse‑data borrowers, aiding credit access and risk assessment.
  • Governance and adoption limits
    • Concerns over fairness, lack of transparent explanations, and regulatory accountability are major barriers to deployment in high‑stakes decisions.
  • Fragmentation in the literature
    • Most studies address performance, fairness, or explainability in isolation; few evaluate their trade‑offs or integrate solutions in regulated deployment scenarios under human oversight.
  • Quality appraisal
    • Review guided by the Relevance, Rigor, Reproducibility, and Quality (3Rs&Q) framework; both empirical and conceptual contributions were synthesized.
  • Practical takeaway
    • The literature offers actionable recommendations (technical and governance) but empirical validation of integrated approaches in real‑world, regulated settings remains sparse.

Data & Methods

  • Review type: Systematic literature review (SLR).
  • Search and selection
    • Databases: IEEE Xplore, Scopus, Web of Science, ScienceDirect.
    • Time window: 2020–2025; last search conducted 30 September 2025.
    • Article types: Peer‑reviewed journal and conference papers.
    • Inclusion criterion: AI‑driven credit scoring applied to consumer/lending contexts addressing performance, fairness, or explainability.
  • Appraisal framework: Relevance, Rigor, Reproducibility, and Quality (3Rs&Q) used to evaluate study contributions and map gaps.
  • Synthesis approach: Comparative analysis across studies to identify common methods, metrics, regulatory considerations, and evidence gaps; distilled both empirical results and conceptual recommendations.

Implications for AI Economics

  • Market efficiency and credit access
    • Improved predictive performance, especially for thin‑file borrowers, can increase credit supply and reduce information frictions, potentially expanding financial inclusion.
    • However, unaddressed fairness issues risk unequal access and biased pricing, counteracting inclusion gains.
  • Distributional and welfare effects
    • If fairness and explainability are not addressed, AI deployment may exacerbate existing disparities (e.g., by protected attribute proxies), with adverse welfare and reputational costs for lenders.
  • Regulatory and compliance costs
    • Requirements for explainability, auditability, and non‑discrimination create compliance burdens that affect the cost‑benefit calculus of AI adoption—smaller lenders may be disproportionately impacted.
  • Competition and innovation
    • Firms that successfully integrate multi‑objective solutions (performance + fairness + explainability) can gain competitive advantage; fragmentation in research slows diffusion of such integrative practices.
  • Systemic risk and model governance
    • Widespread use of similar black‑box models without robust oversight increases correlated model risk; regulators and market participants should prioritize monitoring and stress testing.
  • Research and policy priorities
    • Need for economic research quantifying trade‑offs (e.g., accuracy vs. fairness vs. mortgage pricing), field experiments on deployment outcomes, and cost‑benefit analyses of regulatory instruments (e.g., disclosure mandates, audit regimes).
  • Practical recommendations for stakeholders
    • Adopt multi‑objective model design (jointly optimizing accuracy, fairness, explainability).
    • Use transparent documentation (model cards, datasheets), rigorous subgroup evaluation, and post‑deployment monitoring.
    • Implement human‑in‑the‑loop processes, audit trails, and regulatory alignment to reduce adoption barriers and mitigate distributional harms.

If you want, I can: (a) extract example metrics and algorithms used across the reviewed studies, (b) draft a short checklist for lenders to operationalize the actionable guidance, or (c) outline a research agenda for economists studying the joint trade‑offs. Which would you prefer?

Assessment

Paper Typereview_meta Evidence Strengthn/a — The paper is a systematic literature review synthesizing results from 43 primary studies rather than producing new causal estimates; it summarizes existing evidence quality and gaps but does not itself generate causal identification or primary empirical effect sizes. Methods Rigorhigh — Authors performed a systematic literature review across four major databases (IEEE Xplore, Scopus, Web of Science, ScienceDirect), applied clear inclusion criteria (peer‑reviewed articles 2020–2025 on AI credit scoring addressing performance, fairness, or explainability), and used a structured appraisal framework (Relevance, Rigor, Reproducibility, and Quality) to assess studies; however the search excludes gray literature and any non‑indexed sources, which the authors acknowledge as a limitation. SampleA corpus of 43 peer‑reviewed journal and conference articles published 2020–2025 (last searched 30 September 2025) retrieved from IEEE Xplore, Scopus, Web of Science, and ScienceDirect, each examining AI‑driven credit scoring in consumer or lending decision contexts and addressing at least one of performance, fairness, or explainability. Themesgovernance inequality GeneralizabilityExcludes gray literature, industry reports, and proprietary vendor documentation, which may omit large-scale deployment evidence, Timebound to 2020–2025; rapid post‑2025 developments not covered, Search limited to four databases and (implicitly) English or indexed publications, risking publication and language bias, Findings synthesize academic research that often focuses on experimental or simulated datasets, so applicability to production, regulated deployments may be limited, Heterogeneity across jurisdictions and lending markets (consumer vs. SME vs. commercial credit) reduces direct transferability of conclusions

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The integration of artificial intelligence (AI) in the financial sector has seen a rapid increase over the past few years. Adoption Rate positive AI adoption rate in financial sector
Reading fidelity high
Study strength medium
not reported
0.24
Predicting the creditworthiness of applicants has demonstrated strong predictive performance, particularly for thin-file clients. Decision Quality positive predictive performance of credit-scoring models (creditworthiness prediction)
Reading fidelity high
Study strength medium
not reported
0.24
Growing concerns regarding AI’s fairness, explainability, and regulatory accountability have increasingly limited its adoption in high-stakes credit decision-making. Governance And Regulation negative limitation of AI adoption in high-stakes credit decision-making due to fairness/explainability/regulatory concerns
Reading fidelity high
Study strength medium
not reported
0.24
This paper presents a synthesis derived from a systematic literature review (SLR) of 43 peer-reviewed studies published between 2020 and 2025. Other null_result number of eligible studies included in the review
Reading fidelity high
Study strength high
n=43
43 studies
0.4
Eligible studies were limited to peer-reviewed journal and conference articles (2020–2025) retrieved from IEEE Xplore, Scopus, Web of Science, and ScienceDirect (last searched: 30 September). Other null_result scope and sources of literature search (databases and date)
Reading fidelity high
Study strength high
not reported
0.4
The review is guided by the Relevance, Rigor, Reproducibility, and Quality (3Rs&Q) appraisal framework. Other null_result critical appraisal framework used for study assessment
Reading fidelity high
Study strength high
not reported
0.4
The review analyzes how existing approaches navigate the interplay among performance, fairness, and explainability under regulatory and human oversight considerations. Governance And Regulation null_result treatment of interactions between performance, fairness, explainability, and oversight in the literature
Reading fidelity high
Study strength high
not reported
0.4
Findings indicate that performance, fairness, and explainability dimensions are predominantly addressed in isolation, with limited attention to their joint treatment in regulated deployment settings. Governance And Regulation negative degree to which literature considers joint treatment of performance, fairness, and explainability in regulated deployment
Reading fidelity high
Study strength medium
not reported
0.24
By consolidating empirical and conceptual evidence, this review provides actionable guidance for designing and deploying credit scoring models in practice. Adoption Rate positive practical guidance for design and deployment of credit scoring models
Reading fidelity high
Study strength speculative
not reported
0.04

Notes