The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Mathematicians can materially strengthen AI safety: rigorous theorems and open problems across logic, probability and geometry can make AI more legible, steerable and cooperative, but translating these foundations into deployed systems remains an urgent interdisciplinary task.

Math for AI safety: an invitation for mathematicians
Lionel Levine · September 14, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Lionel Levine unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Lionel Levine provider ID
An invitation and roadmap for mathematicians arguing that new, rigorous mathematics across logic, probability, algebra, and analysis is essential to make AI systems legible, steerable, and cooperative, illustrated with formal results and concrete open problems.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial intelligence threatens to outrun human understanding and control. New mathematics is needed to design AI that is legible, steerable, and cooperative with humanity. I organize this invitation by mathematical field, so you can turn straight to your own: logic and game theory for cooperation; probability for agency and world-models; algebra and representation theory for learned features; analysis and geometry for generalization and training dynamics. Each section ends with an open problem that is accessible to a working mathematician with no prior experience in AI safety.

Summary

Main Finding

Levine argues that new, rigorous mathematics is urgently needed to make advanced AI systems legible, steerable, and cooperative. He frames this as a civic call to mathematicians: organize research by mathematical field (logic & game theory; probability & causality; algebra & representation theory; analysis & geometry), formalize key concepts (goals, beliefs, learned concepts, generalization, training dynamics), and work on concrete, machine-checkable theorems and accessible open problems that directly reduce systemic risks from AI. The paper illustrates the approach with existing mathematical results (e.g., Löb’s theorem → FairBot cooperation, logical induction) and proposes a repository (MAIS) of problems aimed at mathematicians.

Key Points

  • Three stances for mathematicians: artisanal (pure-math, beauty), industrial (machine-assisted proof/discovery), and civic (work directed explicitly toward societal benefit, notably AI safety). Levine urges adopting the civic stance for work prioritized by societal impact.
  • Three safety desiderata at different scales:
    • Legible: internal representations are human-interpretable (interpretability).
    • Steerable: simple, reliable runtime interventions change behavior predictably.
    • Cooperative: multi-agent protocols and institutions yield broadly desirable social outcomes.
  • A five-stage pipeline for producing AI systems: Define (specify target), Measure (scoring/benchmarks), Train (data & algorithms), Test (holdout/adversarial tests), Deploy (real-world use); failures at deployment feed back to the earlier stages. Many safety failures arise from gaps in definition, measurement, or distribution/test design.
  • Logic & game theory: when agents can read each other’s code, classical game-theoretic predictions change. Example: FairBot (cooperate iff you can prove the opponent cooperates) can provably cooperate with itself via Löb’s theorem; bounded versions are quantitative questions. Transparent agents open opportunities for provable cooperation and coordinated equilibria — but also for collusion and coordinated harmful behavior humans cannot match or detect.
  • Probability & causality: we need precise mathematical notions of goals and beliefs for agents. Two perspectives are surveyed: (i) agents-as-utility-maximizers (Bayesian decision theory) and (ii) agents-as-predictors/minimizers of free energy (active inference). Logical uncertainty (how bounded agents assign probabilities to statements they cannot yet prove) is an important formalism (logical induction).
  • Algebra & representation theory, analysis & geometry: these fields can formalize what “learned concepts” are (representations), and explain generalization and training dynamics (loss landscapes, implicit regularization). These mathematical foundations can point to architectures and training procedures that favor legibility and steerability.
  • The paper is practical: it presents accessible open problems (e.g., quantitative bounded Löb), describes machine-checked formalizations (Lean 4), and points to a living repository (MAIS) to coordinate problems and contributions.

Data & Methods

  • Nature of the work: theoretical survey and research agenda, not an empirical study. Methods are:
    • Synthesis of mathematical results across fields and mapping them to AI-safety problems.
    • Formal theorem statements and proofs (some machine-checked; e.g., Lean 4 formalization of proof-based open-source game theory).
    • Constructive examples and thought experiments (e.g., CliqueBot, FairBot, FairBotk).
    • Proposal of concrete open problems that blend formal proof tasks with computational/quantitative evaluation (e.g., implement bounded proof search, tabulate thresholds).
    • Use of case studies and incidents (e.g., Hugging Face hacking incident) as motivational evidence for illegibility and the need for monitoring/definitions.
  • No novel datasets or empirical methods are introduced; the “data” are prior theorems, formal models, and documented incidents illustrating failure modes.

Implications for AI Economics

  • Delegation and comparative advantage: mathematically formalized legibility and steerability affect firms’ willingness to delegate important tasks to AI. Poor legibility induces asymmetric information and increases reliance on costly monitoring or regulatory constraints, altering investment and adoption dynamics.
  • Market competition & disempowerment: Levine emphasizes a competitive pressure toward delegation to opaque AI. Economically, this can create a race dynamic (first-mover advantages, externalities) that disempowers human skills and shifts comparative advantage toward AI-using organizations. Mathematical tools that measure agency (time horizon, legibility metrics) can quantify these effects and inform policy.
  • Collusion and new cartel risks: transparent/mutually-readable agents can provably coordinate in ways humans cannot detect or match (Löb/FairBot phenomena). This raises antitrust concerns: firms might deploy agents that tacitly collude, creating durable coordinated equilibria without explicit human agreements. Economics must model these new bilateral and multilateral coordination channels; quantitative bounded-Löb results would inform how costly or feasible such implicit collusion is.
  • Principal–agent and mechanism design: improved formal definitions of goals and beliefs let economists and mechanism designers update models of incentives and enforceability. Designing robust contracts, audits, and mechanisms when agents themselves may be strategic AIs (and may read each other) requires new theory—mathematical results on steerability and provable commitments are directly relevant.
  • Measurement, scoring rules, and perverse incentives: Levine’s pipeline highlights that how we define and measure success (benchmarks, loss functions, scores) drives optimization behavior. Economic agents (firms, platforms) respond to incentives set by metrics; mis-specified metrics can lead to harmful equilibria (e.g., engagement-maximizing content). Formal work on scoring rules and distribution/test design maps directly to regulation and incentive engineering.
  • Information asymmetries & market design: advances in interpretability reduce information asymmetries between buyers, regulators, and AI developers, potentially lowering transaction costs and enabling more efficient markets. Conversely, improved agent transparency among firms could increase collusion risk. Economic policy must therefore trade off benefits of transparency for human principals against risks of agent-to-agent coordination.
  • Systemic risk & institutions: if institutions increasingly run on AI agents, formal analyses of multi-agent cooperation and failure modes (e.g., emergent incentives, cascading failures) are required to model systemic risk, systemic moral hazard, and dynamic stability of economic systems. Mathematical models from game theory and dynamical systems can quantify conditions for stability versus runaway behavior.
  • Regulation and enforceability: formal, machine-checkable notions (e.g., provable constraints on agent behavior, steerability primitives) could make regulation more enforceable. Economists and policymakers can use such formal tools to design compliance standards, audits, and certification schemes that are technically verifiable.
  • Labor markets & inequality: rigorous measures of “agency” (time horizon, planning depth) can help forecast which tasks are automatable and how fast, informing policy responses (retraining, social insurance). Mathematical analysis of representation and generalization can indicate which cognitive tasks remain hard for AI, shaping skill demand.
  • Public goods and coordination for safety research: Levine’s civic-math framing highlights safety math as a public good. Economists should consider funding, coordination, and incentive mechanisms to support math-for-safety (grants, open repositories like MAIS, cooperative research platforms) to correct underprovision due to competitive pressures.
  • Actionable suggestions for economists and policymakers:
    • Incorporate legibility/steerability metrics into regulatory compliance and procurement criteria.
    • Model markets where firms can deploy mutually-readable agents; study equilibrium outcomes and antitrust implications.
    • Fund and prioritize math problems with direct economic impact (e.g., quantitative bounded-Löb to assess collusion costs).
    • Build audit and scoring standards based on formal definitions (to avoid perverse incentives).
    • Coordinate cross-disciplinary work (mathematicians + economists + computer scientists + policy) to design institutions resilient to AI-mediated coordination.

Overall, Levine’s survey provides a rigorous roadmap: mathematical formalization of agent properties and multi-agent interactions yields tools that will materially change economic modeling, regulation, and market design around AI. Economists should engage with the open problems and formal methods Levine proposes, because the solutions (or lack thereof) will shape incentives, competitive dynamics, and systemic risk in AI-driven economies.

Assessment

Paper Typetheoretical Evidence Strengthn/a — This is a conceptual/theoretical survey and research invitation, not an empirical study; it summarizes formal results and poses open mathematical problems rather than providing causal or observational evidence. Methods Rigorhigh — The paper is grounded in formal mathematical results (e.g., use of Löb's theorem, formalized proofs in Lean, references to logical induction and bounded Löb arguments) and frames precise open problems; arguments are stated with mathematical care rather than loose speculation. SampleNo empirical sample; the paper is a mathematical survey and research agenda that synthesizes existing theorems (e.g., FairBot/Löb arguments), references formalization work (Lean), conceptualizes three desiderata for safe AI (legibility, steerability, cooperativeness), and lists open mathematical problems across logic, probability, algebra, analysis, and geometry. Themesgovernance human_ai_collab GeneralizabilityNot empirical — does not provide measured economic outcomes or firm-/worker-level impacts., Focused on mathematical foundations and formal models; translation to engineering practices, deployed systems, or policy requires additional work., Audience is mathematicians and theoreticians; practical uptake depends on interdisciplinary collaboration with ML engineers, policymakers, and firms., Some models (e.g., source-code games) rely on strong assumptions (full source transparency) that may not hold in deployed multiagent settings.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Two copies of FairBot cooperate: Peano Arithmetic proves that FairBot cooperates with itself. Ai Safety And Ethics positive Cooperation between source-code-transparent AI agents
Reading fidelity high
Study strength high
not reported
0.2
FairBot is unexploitable under the assumption that Peano Arithmetic proves only true facts about the relevant program outputs: it never cooperates with an opponent that defects against it. Ai Safety And Ethics positive Resistance to exploitation by defecting opponents
Reading fidelity high
Study strength high
not reported
0.2
Bounded versions of FairBot cooperate with each other once their proof-search bound is sufficiently large. Ai Safety And Ethics positive Mutual cooperation by computationally bounded agents
Reading fidelity high
Study strength medium
n=2
0.12
Löbian cooperation in proof-based open-source game theory has been machine-checked in Lean 4. Ai Safety And Ethics positive Formal verification of cooperation proofs
Reading fidelity high
Study strength medium
n=2
0.12
Mutually transparent AI agents can reach cooperative equilibria that collude against their human principals, potentially without humans detecting the collusion. Ai Safety And Ethics negative Human control over AI-agent interactions and principal-agent alignment
Reading fidelity high
Study strength speculative
not reported
0.02
A logical inductor provides a mathematically precise way for a computationally bounded agent to assign graded confidence to statements before they are proved or disproved. Decision Quality positive Confidence updating under logical uncertainty
Reading fidelity high
Study strength medium
not reported
0.12
On a suite of software tasks, the AI time horizon has roughly doubled every seven months since 2019, increasing from seconds to about sixteen hours by May 2026. Task Completion Time positive AI time horizon for successful completion of software tasks
Reading fidelity high
Study strength medium
doubled roughly every seven months; about sixteen hours as of May 2026
0.12
Current AI training methods generally produce systems whose internal thought processes are difficult for human researchers and engineers to understand. Ai Safety And Ethics negative Human interpretability of AI internal representations and reasoning
Reading fidelity high
Study strength low
not reported
0.06
Activation steering, a current technique for modifying AI behavior at runtime, is often brittle and unreliable. Ai Safety And Ethics negative Reliability and predictability of runtime behavioral interventions
Reading fidelity high
Study strength low
not reported
0.06

Notes