The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Bankers say AI boosts credit-decision accuracy and consistency, but the biggest improvements come when disciplined data pipelines and governance back the models rather than from model sophistication alone.

AI-Assisted Credit Evaluation Models for Improving Risk Assessment Accuracy in U.S. Banking Systems
Mahfuj Ahmed Ruzel · January 01, 2026 · American Journal of Interdisciplinary Studies
openalex correlational low evidence 7/10 relevance Summary only summary available; pdf_status=not_found DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Mahfuj Ahmed Ruzel provider ID

Semantic Scholar

Latest observation:

  1. Mahfuj Ahmed Ruzel provider ID
Banking professionals report that AI-assisted credit evaluation improves perceived risk assessment accuracy and underwriter consistency, with data quality and governance practices driving larger perceived gains than model sophistication alone.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This study investigated a persistent problem in U.S. banking credit decisioning: traditional scorecards and manual underwriting can produce inconsistent judgments and avoidable misclassification, especially when borrower profiles are complex and decision speed is high. The purpose was to quantify how AI-assisted credit evaluation, deployed within enterprise banking environments, improves perceived risk assessment accuracy and which enabling conditions most strongly drive those gains. Using a quantitative, cross-sectional, case-based design, data were collected via a structured 5-point Likert survey from n = 214 eligible banking professionals (usable response rate 71.3%) across underwriting (38.8%), credit analysis (27.1%), risk management (22.0%), and model risk/compliance (12.1%), representing enterprise-grade, cloud-supported decision workflows in the case banks. Key variables included Risk Assessment Accuracy Improvement (dependent) and five predictors: AI Model Capability, Data Quality and Availability, Explainability/Transparency, Governance and Compliance Alignment, and Monitoring and Drift Management. The analysis plan applied reliability testing (Cronbach’s α), descriptive statistics, Pearson correlations, and multiple regression. Measurement reliability was strong (α = .81–.90; DV α = .90). Descriptively, respondents agreed that AI improved accuracy (DV M = 3.97, SD = 0.63), with high ratings for data quality (M = 4.05) and governance (M = 3.94), while monitoring was lower (M = 3.72). Accuracy improvement correlated significantly with all predictors (r = .39–.56, p < .001), strongest for data quality (r = .56) and governance (r = .51). In regression, the model explained substantial variance (R² = .46; Adj. R² = .44; F(5,208) = 35.4, p < .001), with Data Quality (β = .29, p < .001), Governance (β = .22, p = .002), and AI Capability (β = .18, p = .006) as significant drivers; explainability was marginal (β = .11, p = .071) and monitoring was not significant after controls (β = .09, p = .104). Practically, the strongest perceived operational gain was improved underwriter consistency (M = 4.06), alongside reduced false approvals (M = 3.84), implying that banks realize the largest accuracy benefits when enterprise AI is paired with disciplined data pipelines and governance controls rather than model sophistication alone.

Summary

Main Finding

AI-assisted credit evaluation in enterprise banking is perceived to substantially improve risk-assessment accuracy (mean = 3.97/5). The largest drivers of that perceived improvement are Data Quality & Availability and Governance & Compliance Alignment, with AI Model Capability also contributing meaningfully. Explainability/Transparency and Monitoring/Drift Management showed weaker or non-significant effects once controls were included.

Key Points

  • Sample: n = 214 banking professionals (71.3% usable response rate) from underwriting (38.8%), credit analysis (27.1%), risk management (22.0%), and model risk/compliance (12.1%) within enterprise, cloud-supported decision workflows.
  • Measurement reliability: Cronbach’s α ranged .81–.90 for predictors; dependent variable α = .90.
  • Descriptives:
    • Risk Assessment Accuracy Improvement (DV): M = 3.97, SD = 0.63
    • Data Quality & Availability: M = 4.05
    • Governance & Compliance Alignment: M = 3.94
    • Monitoring & Drift Management: M = 3.72 (lowest of the five)
    • Greatest perceived operational gains: improved underwriter consistency (M = 4.06) and reduced false approvals (M = 3.84).
  • Correlations: DV correlated significantly with all predictors (r = .39–.56, p < .001); strongest correlations for Data Quality (r = .56) and Governance (r = .51).
  • Multiple regression (N = 214):
    • Model fit: R² = .46; Adj. R² = .44; F(5,208) = 35.4, p < .001.
    • Significant predictors:
      • Data Quality & Availability: β = .29, p < .001
      • Governance & Compliance Alignment: β = .22, p = .002
      • AI Model Capability: β = .18, p = .006
    • Marginal/not significant:
      • Explainability/Transparency: β = .11, p = .071 (marginal)
      • Monitoring & Drift Management: β = .09, p = .104 (not significant after controls)

Data & Methods

  • Design: Quantitative, cross-sectional, case-based survey study of enterprise banking workflows.
  • Instrument: Structured 5‑point Likert survey measuring perceived Risk Assessment Accuracy Improvement (DV) and five enabling/predictor constructs.
  • Analysis: Reliability testing (Cronbach’s α), descriptive statistics, Pearson correlations, and multiple linear regression to identify independent associations.
  • Strengths: Targeted sample of enterprise practitioners, high measurement reliability, and multivariate controls showing independent contributions.
  • Limitations: Cross-sectional and perception-based (self-report) data limit causal inference; case-based sampling may affect external generalizability; potential common-method bias and omitted-variable confounding.

Implications for AI Economics

  • Complementarities matter: Returns to AI in credit decisioning accrue less from model sophistication alone and more from complementary investments in data quality and governance. This aligns with economic models of technological adoption where complementary organizational investments amplify productivity gains.
  • Investment prioritization: Banks seeking to maximize accuracy gains (and downstream cost savings from reduced misclassification) should prioritize spending on disciplined data pipelines, data availability/quality control, and governance/compliance integration before or alongside advanced model development.
  • Operational value and ROI channels: Perceived benefits center on increased underwriter consistency and fewer false approvals—implying reduced operational variance, lower credit losses, and potential labor productivity gains. These are concrete channels through which AI investment can generate economic value.
  • Regulation and risk management: Strong governance alignment enhances perceived accuracy gains, suggesting that compliance-ready AI implementations can both improve decision quality and ease regulatory adoption—lowering regulatory friction costs.
  • Caution for vendors and policymakers: Explainability and monitoring remain important but showed weaker incremental associations in this study; firms should not neglect them, especially for long-run model safety and regulatory expectations. Policymakers should recognize that simply mandating model transparency without supporting data- and governance-related investments may have limited effect on realized accuracy improvements.
  • Research implications: Future economic evaluations (cost–benefit and causal studies) should quantify actual performance outcomes (misclassification rates, default rates, credit losses) and measure the interaction effects among model capability, data investments, and governance to estimate returns to each investment component.

Assessment

Paper Typecorrelational Evidence Strengthlow — Findings are based on self-reported perceptions from a cross-sectional survey rather than objective performance or experimental data, so associations may reflect common-method bias, selection effects, reverse causality, or unobserved confounding rather than causal effects of AI deployment. Methods Rigormedium — The study uses appropriate psychometric checks (Cronbach's α), adequate sample size and response rate, and standard statistical procedures (correlations, multiple regression), but is limited by cross-sectional design, reliance on subjective measures, limited control set, and lack of behavioral/operational outcome validation. Samplen = 214 eligible U.S. banking professionals (usable response rate 71.3%) across underwriting (38.8%), credit analysis (27.1%), risk management (22.0%), and model risk/compliance (12.1%), drawn from enterprise-grade, cloud-supported decision workflows in the case banks; data are self-reported Likert responses from a cross-sectional survey. Themeshuman_ai_collab governance productivity IdentificationNo causal identification; cross-sectional associations estimated from a structured Likert survey using reliability testing and multivariate OLS regression to identify predictors of perceived accuracy improvement. GeneralizabilityNon-random, case-based sample of enterprise banking professionals limits representativeness of all banks or lenders, Findings reflect perceptions rather than measured underwriting outcomes or portfolio performance, Results may not generalize to small banks, fintech lenders, consumer vs. commercial credit, or non-U.S. jurisdictions, Cross-sectional design prevents inference about effects over time or in production-scale deployments, Potential respondent selection and social desirability biases (professionals working on AI may view it favorably)

Claims (13)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study analyzed responses from n = 214 eligible banking professionals (usable response rate 71.3%) across underwriting (38.8%), credit analysis (27.1%), risk management (22.0%), and model risk/compliance (12.1%). Other null_result Sample composition and response rate
Reading fidelity high
Study strength high
n=214
usable response rate 71.3%; role breakdown: underwriting 38.8%, credit analysis 27.1%, risk management 22.0%, model risk/compliance 12.1%
0.5
Measurement reliability for the scales was strong (Cronbach’s α = .81–.90; dependent variable α = .90). Other null_result Scale reliability (Cronbach's α)
Reading fidelity high
Study strength high
n=214
α = .81–.90; DV α = .90
0.5
Respondents agreed that AI improved risk assessment accuracy (Risk Assessment Accuracy Improvement M = 3.97, SD = 0.63). Decision Quality positive Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
DV M = 3.97, SD = 0.63
0.3
Survey ratings were high for Data Quality and Availability (M = 4.05) and Governance and Compliance Alignment (M = 3.94), while Monitoring and Drift Management was lower (M = 3.72). Other mixed Predictor variable ratings (Data Quality, Governance, Monitoring)
Reading fidelity high
Study strength medium
n=214
Data Quality M = 4.05; Governance M = 3.94; Monitoring M = 3.72
0.3
Risk Assessment Accuracy Improvement correlated significantly with all five predictors (Pearson r = .39–.56, p < .001), strongest for Data Quality (r = .56) and Governance (r = .51). Decision Quality positive Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
r = .39–.56, p < .001; Data Quality r = .56; Governance r = .51
0.3
The multiple regression model explained substantial variance in perceived accuracy improvement (R² = .46; Adj. R² = .44; F(5,208) = 35.4, p < .001). Decision Quality positive Risk Assessment Accuracy Improvement (model fit)
Reading fidelity high
Study strength medium
n=214
R² = .46; Adj. R² = .44; F(5,208) = 35.4, p < .001
0.3
Data Quality and Availability was a significant positive predictor of Risk Assessment Accuracy Improvement (β = .29, p < .001). Decision Quality positive Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
β = .29, p < .001
0.3
Governance and Compliance Alignment was a significant positive predictor of Risk Assessment Accuracy Improvement (β = .22, p = .002). Decision Quality positive Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
β = .22, p = .002
0.3
AI Model Capability was a significant positive predictor of Risk Assessment Accuracy Improvement (β = .18, p = .006). Decision Quality positive Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
β = .18, p = .006
0.3
Explainability/Transparency had a marginal (near-significant) positive association with accuracy improvement in regression (β = .11, p = .071). Decision Quality mixed Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
β = .11, p = .071
0.3
Monitoring and Drift Management was not a significant predictor of perceived accuracy improvement after controls (β = .09, p = .104). Decision Quality null_result Risk Assessment Accuracy Improvement
Reading fidelity high
Study strength medium
n=214
β = .09, p = .104
0.3
The strongest perceived operational gains from AI-assisted credit evaluation were improved underwriter consistency (M = 4.06) and reduced false approvals (M = 3.84). Decision Quality positive Underwriter consistency (improvement) and False approvals (reduction)
Reading fidelity high
Study strength medium
n=214
Improved underwriter consistency M = 4.06; reduced false approvals M = 3.84
0.3
The study concludes that banks realize the largest perceived accuracy benefits when enterprise AI is paired with disciplined data pipelines and governance controls rather than model sophistication alone. Decision Quality positive Risk Assessment Accuracy Improvement (operational drivers)
Reading fidelity high
Study strength medium
n=214
Interpretive conclusion based on β coefficients: Data Quality β = .29; Governance β = .22; AI Capability β = .18
0.3

Notes