0 cumulative citations
View corpus contextA reproducible index for 'agentic' AI in banks promises to separate autonomous execution from advisory AI in disclosure data and enable credible causal analysis; the paper supplies a validated AAII measure and a battery of complementary identification strategies but stops short of applying them to produce outcome estimates.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Summary
Main Finding
This Method Article develops and validates the Agentic AI Intensity Index (AAII), a disclosure‑based, bank‑level measure that distinguishes genuine agentic (delegated‑execution) AI from advisory or generic AI adoption. The paper specifies a fully reproducible five‑step construction and validation protocol (including human + large‑language‑model classification with adjudication), an authority‑weighted scoring rule and normalization to isolate agentic intensity net of overall AI talk. It also provides a matched identification strategy (multiple causal designs and threat mitigations) required to estimate AAII’s effects on bank efficiency, risk detection, operational risk and market reaction. The contribution is methodological: the index and the inferential machinery are presented as a directly executable protocol rather than as applied empirical estimates.
Key Points
- Measurement gap: existing text/term‑frequency AI/digitalisation indices do not separate autonomous execution from decision support and lack psychometric validation.
- Conceptual definition: AAII = extent to which a bank has delegated consequential financial decision and execution authority to autonomous systems (authority, irreversibility and delegated mandate are central).
- Theoretical anchors:
- Dynamic Capabilities Theory — AAII signals a firm’s capacity to sense, seize and transform (governance/maturity moderate outcomes).
- Technology–Organisation–Environment (TOE) — technological, organisational and environmental context (e.g., regulation, fintech pressure, CBDC exposure) condition effects.
- AAII is intended to enable joint evaluation of productivity gains and priced risk (efficiency vs systemic/operational risk trade‑offs), and to support tests of nonlinearity and conditional effects.
Principal methodological design elements - Five‑step AAII construction protocol: 1. Theory‑driven seed keyword pool (appendix lists seeds). 2. Two‑round expert Delphi refinement to finalise items. 3. Inter‑rater reliability assessment with prespecified acceptance thresholds and a coding manual (excerpted in appendix). 4. Exploratory and confirmatory factor analysis to establish dimensional structure. 5. LLM‑assisted document classification with mandatory human adjudication; authority‑weighted scoring rule separates delegated execution language from advisory language; normalization isolates agentic intensity net of general AI‑discussion volume. - Construct validity: convergent, discriminant, criterion and known‑groups tests against independent deployment indicators (external benchmarks). - Threats and mitigations: - Reverse causality (richer banks adopt more AI): addressed via multiple, overlapping identification designs. - AI‑washing and strategic disclosure: explicit detection and robustness checks; human validation of model labels emphasised. - Vendor concentration, explainability and adversarial manipulation of text signals are acknowledged and logged as limitations/controls. - Reproducibility: full protocol, coding manual excerpts, seed keywords and supplementary tables provided; acceptance thresholds specified at each validation step (protocol oriented).
Data & Methods
- Source documents: annual reports, Management Discussion & Analysis (MD&A), and earnings‑call transcripts for banks (panel design described; sample design and variable definitions in appendices).
- Index construction:
- Keyword generation grounded in theory (Dynamic Capabilities + TOE) and refined via Delphi with domain experts.
- Human coders perform reliability checks (inter‑rater metrics and thresholds reported in the protocol).
- Factor analysis (EFA/CFA) to confirm dimensionality and to justify authority weights.
- Classification pipeline: LLMs used to scale text classification; every LLM classification is subject to human adjudication (protocol includes rules for disagreement resolution).
- Scoring: authority‑weighted rule assigns higher weight to language indicating delegated authority/execution; normalization controls for total AI discussion volume.
- Validation strategy:
- Construct validity: convergent (correlation with independent deployment signals), discriminant (distinguish from generic digitalisation indices), criterion (predictive relationships for outcomes), and known‑groups (e.g., early adopters versus non‑adopters).
- Identification strategies for causal inference (detailed, overlapping designs to triangulate effects):
- Two‑way fixed effects and event‑study setups.
- Staggered‑adoption difference‑in‑differences (with careful handling of heterogeneous timing).
- Instrumental variables (instruments and exclusion logic specified).
- System GMM for dynamic panels.
- Propensity‑score matching and synthetic control for case studies / single‑treatment events.
- Robustness: placebo tests, falsification, handling of anticipatory disclosure and measurement error.
Implications for AI Economics
- Provides a validated, replicable bank‑level measure that distinguishes agentic (delegated execution) AI from advisory AI — enabling more precise empirical work on automation, productivity and risk.
- Enables rigorous causal testing of conflicting hypotheses in the literature (e.g., productivity gains vs increased operational/systemic risk; nonlinearity and heterogeneous effects by governance/regulatory context).
- Facilitates joint assessment of market pricing, lending outcomes (including fairness/disparate impact), operational‑risk disclosures and systemic vendor concentration, by providing a continuous, theory‑grounded exposure variable.
- Encourages best practice for text‑based measurement: transparent item generation, expert refinement, psychometric validation, and mandatory human oversight of model classifications to reduce AI‑washing and adversarial manipulation.
- Policymaking and supervision: gives regulators an operational instrument (and a prespecified identification plan) to evaluate agentic AI deployments, test containment measures, and monitor correlated model/vendor risks across the banking sector.
- Research agenda opened: heterogeneous effects across functions/workflows, interactions with governance quality and regulatory regimes, macro‑financial amplification channels from correlated agentic adoption, and evaluation of mitigation policies (e.g., vendor diversification, approval thresholds for delegated authority).
Notes - The article is a methods/protocol paper: it specifies the index and causal designs but does not report AAII application estimates. Supplementary files include Tables 1–4, seed keywords and coding manual excerpts.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The paper proposes the Agentic AI Intensity Index (AAII), a disclosure-based bank-level measure intended to distinguish autonomous execution of consequential financial activities from advisory or generic AI adoption. Automation Exposure | positive | Bank-level agentic AI intensity, specifically the extent of delegated autonomous financial decision and execution authority. |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The AAII construction protocol includes theory-driven keyword-pool generation, two rounds of expert Delphi refinement, inter-rater reliability assessment, exploratory and confirmatory factor analysis, and large-language-model-assisted classification with mandatory human adjudication. Automation Exposure | positive | Reliability and construct measurement quality of the AAII classification protocol. |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The proposed AAII scoring rule weights the authority delegated to the AI system and normalizes agentic intensity relative to the overall volume of general AI discussion. Automation Exposure | positive | Normalized agentic AI intensity and the distinction between delegated execution and decision support. |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper proposes establishing AAII construct validity through convergent, discriminant, criterion, and known-groups tests against independent deployment indicators. Automation Exposure | positive | Construct validity of the AAII. |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper specifies, but does not estimate, relationships between AAII and bank efficiency, risk detection, operational risk, and market reaction. Firm Productivity | null_result | Bank efficiency, risk detection, operational risk, and market reaction as proposed dependent variables. |
Reading fidelity
high
Study strength
low
|
not reported
|
| The proposed empirical identification strategy combines two-way fixed effects, event studies, staggered-adoption difference-in-differences, instrumental variables, system GMM, propensity-score matching, and synthetic control. Firm Productivity | positive | Estimated effects of AAII on bank performance and systemic-risk-related outcomes. |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper argues that fixed-effects-only studies of AI and bank outcomes do not address reverse causality because better-capitalized banks may be able to afford more AI. Firm Productivity | negative | Validity of causal estimates linking AI adoption to bank outcomes. |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The conceptual framework hypothesizes that AAII affects four outcome domains—efficiency, risk detection, operational risk, and market reaction—with the first three paths conditioned by organizational and environmental moderators. Organizational Efficiency | mixed | Efficiency, risk detection, operational risk, and market reaction. |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The paper states that existing disclosure-based technology indices generally do not distinguish autonomous execution from AI adoption in general and lack documented psychometric evidence such as derivation procedures, inter-coder agreement, and dimensional structure. Automation Exposure | negative | Reliability and discriminative validity of existing disclosure-based AI or technology indices. |
Reading fidelity
high
Study strength
medium
|
not reported
|