0 cumulative citations
View corpus contextA practical framework ties model-level testing to system-level harms and shows that simple adversarial nudges to LLM-driven trading recommendations can measurably worsen financial contagion in a stylised RTGS model; the effect is strongest under broad or concentrated AI adoption but the example is illustrative rather than operationally validated.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where harms emerge from interactions between technical, human, and organisational elements. Yet current AI evaluation remains model-centric, offering little insight into how observed behaviours might translate into system-level risk. We propose a framework that links structured hazard analysis, component-level testing, and probabilistic system modelling to bridge this gap. By providing a traceable pathway from model behaviour to system-level outcomes, the framework enables practitioners to answer the "so what?" of AI failures, quantify their systemic impact, and move toward evidence-based and anticipatory governance of AI in complex systems. Applied to the UK's Real Time Gross Settlement (RTGS) system as an illustrative worked example, we derive AI-driven loss scenarios using Systems Theoretic Process Analysis (STPA) and examine adversarial manipulation of LLM-based trading as one such loss scenario. Component-level experiments show that simple adversarial inputs induce measurable behavioural shifts where AI recommendations are followed. Under the component-to-system mapping used here for a financial contagion model, these shifts alter system resilience, increasing bank failures and lowering the threshold at which shocks lead to cascading disruption, particularly under widespread or monopolistic AI adoption.
Summary
Main Finding
The paper develops a practical, traceable framework that links structured hazard analysis (STPA), component-level testing of AI behaviour, and probabilistic system simulation to quantify how AI adoption can generate system-level harms in complex sociotechnical systems. Applying the framework to a worked example (AI-augmented trading interacting with the UK RTGS/financial network), the authors show that simple, low-cost adversarial manipulations of LLM-based trading aids can produce measurable behavioural changes at the component level which, when mapped into a stylised financial contagion model, reduce system resilience — increasing bank failures and lowering the shock threshold for cascading disruption, particularly under widespread or monopolistic AI adoption.
Key Points
- Framework: A three-stage, modular approach for policy makers and CNI stakeholders:
- A — structured hazard analysis (STPA) to identify AI-driven loss scenarios;
- B — component-level testing to measure propensity for concerning behaviours;
- C — simulation-based modelling to quantify system-level consequences.
- STPA results: Applied to an abstracted RTGS control model, STPA produced over 100 unsafe control actions (UCAs) and 8 AI-driven loss scenarios. The analysis focuses on security-style failure modes (e.g., adversarial manipulation) rather than internal model development.
- Selected loss scenario: LS-18.2.2-A — extensive delegation of trading decisions to generative AI, enabling adversaries to amplify fire-sales via indirect prompt-injection-style attacks.
- Component testing: Built a lightweight LLM-based investment assistant that ingests four asset‑type news items (Equities, MBS, Corporate Bonds, Government Bonds) and outputs portfolio allocations (probabilities summing to 1). Baseline (AllNeutralRun) and targeted asymmetry tests (FixedBearishNeutralRun) establish controls.
- Adversarial testing: Low-cost, indirect prompt-injection attacks (inspired by SEO-style manipulations) were placed adjacent to the bearish asset mention. These attacks bias the LLM recommendations toward exacerbating sell pressure on the distressed asset.
- System projection: Measured behavioural shifts were mapped into a stylised financial contagion model; under the component-to-system mapping used, adversarially amplified behaviour increased failure rates and made cascades more likely. Effects are stronger with high penetration or single-provider concentration of AI tools.
- Practical stance: Emphasises a proportionate, iterative approach — analysts can stop at stage A or B if risks are infeasible or mitigations trivial; proceed to C when residual risk merits system-level quantification.
Data & Methods
- Structured hazard analysis:
- Used STPA (five-step process) on an abstracted RTGS + bank network control structure (stakeholders, losses, hazards, constraints, hierarchical control structure, UCAs).
- Mapped plausible AI use-case classes (grounded in industry surveys) onto the control model to derive AI-driven loss scenarios.
- Component-level experimentation:
- Implemented an LLM-based decision component that returns portfolio allocations across four asset classes given four synthetic news articles (sentiments: Bullish/Bearish/Neutral). Each "Episode" = one invocation; runs aggregate multiple Episodes.
- Baseline experiments: AllNeutralRun; asymmetry test: FixedBearishNeutralRun.
- Adversarial attacks: indirect prompt-injection style manipulations targeting the asset already in Bearish sentiment. Attacks are simple, low-cost, and designed to bias preferences rather than evade safety filters.
- Responses were collected and compared across conditions to measure propensity and magnitude of behavioural shift.
- Simulation / system modelling:
- Used a stylised financial contagion model (grounded in macro-financial literature) to map component-level shifts into bank failure probabilities and contagion thresholds.
- The mapping is illustrative (not a high-fidelity RTGS model) and intended to show how measurable component effects can alter system resilience.
- Reproducibility: Code and STPA artefacts are available in the project repository (appendix reference in paper). The authors note the worked example is intentionally high-level to avoid disclosing exploitable system details.
- Limitations explicitly acknowledged: stylised system abstraction, scenario selection (upper-bound adoption), not modelling internal AI training/evolution, and not claiming the RTGS analysis is an operational assessment.
Implications for AI Economics
- Systemic externalities from AI adoption:
- AI decision aids can create correlated behaviour across institutions (herding), producing negative systemic externalities not captured by firm-level benefit/cost analyses.
- Concentration in third-party AI providers magnifies these externalities (single-source failure modes).
- Pricing and regulation:
- Traditional microprudential tools may underprice the systemic risk introduced by widespread automation and correlated model-driven decisions. There is a case for macroprudential-style interventions (stress testing, capital buffers, concentration limits for AI vendors).
- Policy instruments to consider: mandated diversity of models/providers, limits on levels of delegation/automation for systemically important functions, disclosure and audit requirements, minimum resilience standards, and incident-reporting regimes.
- Risk measurement & modelling recommendations:
- Economic models of AI adoption should include endogenous adoption heterogeneity, network topology of exposures, provider concentration, and correlated behavioural responses to shocks and adversarial signals.
- Incorporate adversarial threat models and low-cost manipulation channels (e.g., prompt-injection vectors) into macro-financial stress tests.
- Calibrate welfare analyses to account for tail risks and non-linear amplification (e.g., lower shock thresholds, increased probability of cascades).
- Governance & incentives:
- Market incentives favour early adoption and automation; absent policy, these incentives can increase systemic fragility. Consider aligning incentives (e.g., through regulation, liability, insurance pricing) to internalise system-level risks.
- Encourage operational measures: robust human-in-the-loop design, diversity of decision systems, monitoring for correlated actions, third-party risk management, and quick remediation/rollback mechanisms.
- Research and data needs:
- Empirical estimation of adoption rates, provider concentration, and the ease with which adversarial manipulations can be embedded into real information flows.
- Higher-fidelity, domain-specific system simulations to quantify expected losses and inform policy calibration (e.g., BoE/FSB-style macroprudential analysis).
- Cost–benefit analysis comparing AI productivity gains against potential systemic loss tail risks to inform proportionate regulation.
Suggested next steps for economists and policy analysts: - Extend the framework with calibrated, higher-fidelity contagion models and real-world adoption data to estimate expected social losses. - Run counterfactuals varying adoption patterns, provider concentration, and mitigation measures to identify high-leverage regulatory levers. - Develop metrics and stress-test scenarios that capture adversarial information-channel risks alongside traditional shock scenarios.
Overall, the paper provides a concrete methodology for translating component-level AI behaviours into system-level economic risks — a needed bridge for evidence-based policy in domains where AI adoption can induce correlated failures and systemic harm.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The paper proposes a three-stage framework linking structured hazard analysis, component-level AI testing, and probabilistic system modelling to trace model behaviours through to system-level harms. Governance And Regulation | positive | Traceability and quantification of AI-related system-level risks |
Reading fidelity
high
Study strength
low
|
not reported
|
| Applying STPA to the UK RTGS system identified more than 100 unsafe control actions and produced eight scenarios in which AI adoption could manifest as harm. Governance And Regulation | positive | Number of identified unsafe control actions and AI-enabled loss scenarios |
Reading fidelity
high
Study strength
medium
|
over 100 unsafe control actions; 8 scenarios
|
| Simple adversarial inputs can induce measurable behavioural shifts in an LLM-based trading recommendation component, particularly in settings where AI recommendations are followed. Decision Quality | negative | AI-generated portfolio allocation recommendations and their behavioural shift under adversarial input |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Under the paper's component-to-system mapping, adversarially induced changes in AI trading behaviour reduce financial-system resilience by increasing bank failures. Fiscal And Macroeconomic | negative | Number of bank failures and financial-system resilience |
Reading fidelity
high
Study strength
low
|
not reported
|
| The model indicates that adversarially induced AI trading shifts lower the threshold at which external shocks produce cascading financial disruption. Fiscal And Macroeconomic | negative | Shock threshold for cascading financial disruption |
Reading fidelity
high
Study strength
low
|
not reported
|
| The modelled systemic harms are especially pronounced when AI adoption is widespread or monopolistic across financial institutions. Market Structure | negative | Bank failures, system resilience, and susceptibility to cascading disruption under different AI-adoption structures |
Reading fidelity
high
Study strength
low
|
not reported
|
| Seventy-five percent of UK financial institutions were already deploying AI in 2024. Adoption Rate | positive | Share of UK financial institutions deploying AI |
Reading fidelity
high
Study strength
medium
|
75%
|