The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A small team converted open-weight checkpoints into Thomson, a near-frontier foundation model family, using continual learning to match or surpass many proprietary models in legal, tax and research tasks at substantially lower reported compute cost; transparent benchmarking and human evaluation back the claim, but limited disclosure of data and baseline conditions tempers reproducibility.

Thomson: Continual Learning of Frontier Models for SovereignAI
Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofrè, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz · August 27, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Shengzhuang Chen unresolved corpus identity
  2. Jerrod Parker unresolved corpus identity
  3. Yejin Bang unresolved corpus identity
  4. Andrew M. Bean unresolved corpus identity
  5. Nabeel Seedat unresolved corpus identity
  6. Stefan Winzeck unresolved corpus identity
  7. Daniil Glazko unresolved corpus identity
  8. Jannik Zgraggen unresolved corpus identity
  9. Fangyi Yu unresolved corpus identity
  10. Scott Arnott unresolved corpus identity
  11. Dietrich Trautmann unresolved corpus identity
  12. Luca Ciuffreda unresolved corpus identity
  13. Guglielmo Bonifazi unresolved corpus identity
  14. Davide Romano unresolved corpus identity
  15. Bradley Bell unresolved corpus identity
  16. Kirsty Fielding unresolved corpus identity
  17. Daniele Giofrè unresolved corpus identity
  18. Tom Zielund unresolved corpus identity
  19. Ipshita Chatterjee unresolved corpus identity
  20. Sneha Murthy Ghantasala unresolved corpus identity
  21. Manpreet Nanreh unresolved corpus identity
  22. John Scoville unresolved corpus identity
  23. Maciej Sakowicz unresolved corpus identity
  24. Wassim Seifeddine unresolved corpus identity
  25. Lukas Thede unresolved corpus identity
  26. Jonathan Richard Schwarz unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Sheng-Zhuang Chen unresolved corpus identity
  2. Jerrod Parker provider ID
  3. Yejin Bang provider ID
  4. Andrew M. Bean provider ID
  5. Nabeel Seedat provider ID
  6. S. Winzeck provider ID
  7. Daniil Glazko provider ID
  8. Jannik Zgraggen provider ID
  9. Fang-Yi Yu provider ID
  10. Scott Arnott provider ID
  11. Dietrich Trautmann provider ID
  12. Luca Ciuffreda provider ID
  13. Guglielmo Bonifazi provider ID
  14. Davide Romano provider ID
  15. Brad Bell provider ID
  16. Kirsty Fielding provider ID
  17. Daniele Giofr'e provider ID
  18. Tom Zielund provider ID
  19. Ipshita Chatterjee provider ID
  20. S. Ghantasala provider ID
  21. Manpreet Nanreh provider ID
  22. John Scoville provider ID
  23. Maciej Sakowicz provider ID
  24. Wassim Seifeddine provider ID
  25. Lukas Thede provider ID
  26. J. Schwarz provider ID
Using a continual-learning pipeline on open-weight checkpoints, a compact team produced Thomson models that achieve near-frontier benchmark performance across many domains while controlling costs and preserving capabilities.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but offers little concrete advice on how this can be achieved in the short term under a diversity of funding settings. We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models. Unlike limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation of a frozen model, our approach exploits a modern mid- & post-training stack while introducing safeguards that preserve both plasticity and stability at each stage, making the minimal number of high-impact interventions on the parameters. This yields gains comparable to those typically seen across multiple successive model generations, at compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for far more actors. We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work. Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research. Evaluations show a distinctive $π$-shaped pattern: distinct improvements across a wide range of capabilities, including those not explicitly targeted, while almost completely eliminating the forgetting problem common to narrow domain adaptation.

Summary

Main Finding

Thomson demonstrates that “frontier” performance can be achieved by non-hyperscale actors by applying a disciplined Continual Learning (CL) pipeline to existing open-weight foundation models. Using full‑weight mid- and post‑training (not tiny fine‑tuning or distillation), careful data curation, agentic/tool training and bespoke reward design, a small engineering team produced Thomson-1.0 models that approach or match many contemporaneous proprietary frontier models across a broad benchmark suite — at substantially lower GPU and personnel cost than full pretraining.

Key Points

  • Continual Learning as a practical SovereignAI route:
    • CL pipeline preserves stability (no catastrophic forgetting) while enabling plasticity (new capabilities); authors report a distinctive “π-shaped” pattern: strong gains in targeted domains plus improvements across many untargeted capabilities.
  • Base models and update strategy:
    • Built from open-weight checkpoints (Qwen3.5‑397B and Qwen3.6‑35B), performing full‑weight updates (not parameter‑efficient adapters) and avoiding large-scale distillation.
  • Focus domains and strengths:
    • Targeted at high‑stakes professional use (legal, tax, journalism, Deep Research, document/RAG workflows). Strong on instruction following, summarization, multilingualism, citation quality and document processing.
  • Costs, team and timeline:
    • Development run with a small team (≈3 dozen engineers/scientists), max ~368 B200 GPUs during experiments; final three‑week training GPU cost estimated < USD 450k; total development cost ~USD 40M (includes research, infrastructure, domain experts) and a three‑month intensive development timeline from first experiments.
  • Safety/value work:
    • Value‑realignment modules (Snowdon variants) and adversarial testing reported; permissive-licensing emphasis for training data; bespoke reward structures to encourage faithful tool use and citation to reduce hallucinations.
  • Limitations noted:
    • Relative weakness in coding and some abstract reasoning vs top proprietary models; evaluation vs closed APIs may be imperfect because of opaque guardrails in those systems.
  • Reproducibility/portability claim:
    • Authors argue the approach generalizes beyond their exact checkpoints and sizes and can be replicated by other institutions with similar resources.

Data & Methods

  • Overall approach:
    • Three complementary CL modules producing foundation models via mid‑training and post‑training stages. Emphasis on minimal, high‑impact parameter interventions and preserving previously learned skills.
  • Data‑centric practices:
    • Rigorous curation: filtering, rephrasing, de‑duplication, semantic targeting of human subject‑matter collection, and Bayesian optimisation to calibrate post‑training data mixtures.
  • Training methods:
    • Full‑weight updates on open checkpoints (no reliance on a stronger teacher), reward design for tool/agentic behavior, localised error correction and end‑to‑end RL for Deep Research tasks.
  • Tool and agentic training:
    • Designed to integrate private retrieval and tool stacks; Deep Research harness for high‑budget tasks; training encourages faithful citations and conditioned tool usage.
  • Evaluation:
    • Large, mixed suite spanning legal benchmarks (e.g., Stanford LegalBench, Harvey LAB), tax, journalism, summarisation, long context, multilingualism, safety/adversarial tests, and human blind preference tests (~3,000 rated conversations). Benchmarks ran under controlled, identical prompt and inference settings where possible; closed models evaluated as served endpoints (acknowledging unknown intermediate guardrails).
  • Key experimental numbers:
    • Thomson-1.0-Small: 35B parameter open-weight release. Thomson‑1.0‑Large: full‑weight updated large model based on Qwen3.5‑397B checkpoint.
    • Reported GPU budget peak: 368 B200 GPUs; final run GPU cost < USD 450k; total development ≈ USD 40M.

Implications for AI Economics

  • Lowers technical & economic barriers to building near‑frontier models:
    • Demonstrates that competent teams with modest compute (vs full pretraining budgets) can produce competitive domain‑specialist or even broadly capable models by investing in CL, data quality and tooling.
  • Reduces exclusive advantage of hyperscalers:
    • Shifts some value capture from scale-of-pretraining to data, tooling, bespoke reward design, and continual refinement — enabling more institutions to claim SovereignAI (model/data/governance control).
  • New cost structure & business models:
    • Organizations can prioritize investment in mid/post‑training pipelines, evaluation suites, and private tool integration rather than massive pretraining runs; this supports differentiated, verticalized AI products (e.g., legal/tax agents) with lower incremental cost.
  • Competitive dynamics:
    • Expect increased decentralization and competition in specialized/high-value segments (enterprises, governments, research labs). Proprietary incumbents may retain advantages in raw scale, hardware access, and opaque system‑level safety measures.
  • Regulatory & governance implications:
    • Wider access to powerful models raises policy relevance: licensing/data provenance, auditability of continual updates, and institutional responsibility for safety/alignment become central. Dependence on open‑weight checkpoints and accelerator hardware concentrated supply chains remain systemic risks.
  • Strategic guidance for institutions:
    • Invest in CL pipelines, data‑centric processes, evaluation and tool sovereignty to achieve high impact at bounded cost.
    • Secure reliable hardware access and plan for ongoing update cycles; treat model sovereignty as a stack (training + data + tools + governance).
  • Risks and caveats for markets:
    • Outcome depends on access to permissively licensed base models; closed‑API guardrails and upstream availability of open checkpoints could constrain adoption.
    • Potential for uneven replication: results are encouraging but may be sensitive to team skill, research infrastructure, and the exact base checkpoints used.

Summary judgment: Thomson provides a compelling, empirically supported blueprint that materially changes the economics of achieving frontier or near‑frontier capabilities for a broad class of institutions. The key economic insight is that focused continual improvement, rigorous data design, and tool integration can substitute for orders‑of‑magnitude pretraining scale in many commercially important domains — shifting where value and competitive advantage reside in the AI ecosystem.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The report provides extensive empirical benchmarking (automated suites and a 3,000+ example blind human preference study), ablation claims, and resource/cost accounting, supporting its main claims about model performance and development cost; however, there are important transparency limitations (details of some datasets, precise training recipes, and full evaluation code are not included in the supplied text), potential conflicts of interest (authors are from the developer organisation), and comparisons against proprietary systems may be affected by unknown guardrail layers and inaccessible baseline settings. Methods Rigormedium — The team uses a broad, consistent evaluation harness across many benchmarks, human preference testing, adversarial safety testing, and reports ablations and cost estimates; but crucial reproducibility details are missing or high-level (exact datasets, full training hyperparameters, and detailed evaluation scripts), and some cross-model comparisons rely on served endpoints where provider-side interventions cannot be ruled out. SampleDevelopment built on open-weight base checkpoints (Qwen3.5-397B and Qwen3.6-35B) producing Thomson-1.0-Small (35B) and Thomson-1.0-Large; training used a continual learning pipeline with mid- and post-training data curation (permissively licensed sources described at high level), domain-focused corpora for legal, tax, journalism and Deep Research, automated benchmark suites (e.g., Stanford LegalBench, Harvey Legal Agent, reasoning, summarisation, multilingual, coding, etc.), adversarial safety tests, and a blind human preference study over ~3,000 rated conversations; compute resources reported up to 368 B200 GPUs, final training run ~3 weeks; cost estimates given (GPU run <$450k; total program ~$40M). Themesorg_design adoption innovation productivity governance GeneralizabilityComparisons to proprietary models may be biased by unknown provider-side guardrails or routing when accessed as served endpoints., Report focuses on two specific open-weight base checkpoints; performance and pipeline effectiveness may vary with different bases or architectures., Domain-targeting focused on legal, tax, journalism and research tasks; gains may not fully generalise to coding or other non-targeted domains (coding showed some forgetting)., Reproducibility limited until full datasets, hyperparameters, and training/evaluation code are released., Hardware availability and engineering expertise requirements still non-trivial; smaller organisations may need adaptations to match reported timelines/costs.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The Thomson family was developed by a technical team of no more than three dozen engineers and scientists, using a compute cluster with no more than 368 B200 GPUs available at any stage of experimentation. Organizational Efficiency positive Development personnel and compute resources required to build the model
Reading fidelity high
Study strength medium
no more than three dozen engineers; no more than 368 B200 GPUs
0.18
The development timeline for Thomson-1.0-Small and Thomson-1.0-Large, measured from the first experiments with Qwen3.5-397B, was three months. Organizational Efficiency positive Model development time
Reading fidelity high
Study strength medium
three months
0.18
The final training run for Thomson-1.0-Large cost under USD 450,000 in GPU costs over three weeks of training. Organizational Efficiency positive GPU cost of the final model training run
Reading fidelity high
Study strength medium
under USD 450,000
0.18
Thomson-1.0-Large achieved an overall benchmark average of 78.5%, compared with 73.6% for its Qwen3.5-397B base model, a difference of 4.9 percentage points. Output Quality positive Overall model benchmark performance
Reading fidelity high
Study strength medium
78.5% vs 73.6% (4.9 percentage-point difference)
0.18
Thomson-1.0-Large outperformed Gemini 3.1 Pro, Sonnet 5, and GLM-5.2 on the reported overall average benchmark score. Output Quality positive Overall comparative benchmark performance
Reading fidelity high
Study strength medium
78.5% vs 78.0%, 75.7%, and 77.2%
0.18
On the legal-domain benchmark average, Thomson-1.0-Large scored 78.4%, compared with 72.7% for Qwen3.5-397B. Output Quality positive Legal-domain benchmark performance
Reading fidelity high
Study strength medium
78.4% vs 72.7% (5.7 percentage points)
0.18
Thomson-1.0-Small achieved an overall benchmark average of 74.6%, compared with 71.7% for Qwen3.6-35B. Output Quality positive Overall small-model benchmark performance
Reading fidelity high
Study strength medium
74.6% vs 71.7% (2.9 percentage-point difference)
0.18
Thomson-1.0-Large scored 97.3% on political neutrality, compared with 85.3% for Qwen3.5-397B. Ai Safety And Ethics positive Political neutrality score
Reading fidelity high
Study strength medium
97.3% vs 85.3% (12.0 percentage points)
0.18
Coding was the only reported domain showing mild forgetting for Thomson-1.0-Large relative to its base model, while general mathematical and abstract reasoning remained within the Qwen performance level. Skill Obsolescence mixed Retention of capabilities during continual learning, especially coding and reasoning
Reading fidelity high
Study strength medium
coding: 39.9% vs 40.9%
0.18
In blind human evaluations covering more than 3,000 preference-rated conversations, Thomson-1.0-Large won between 53% and 62% of comparisons against the listed external frontier systems, with ties ranging from 10% to 17%. Output Quality positive Human preference in pairwise system comparisons
Reading fidelity high
Study strength medium
n=3000
53%-62% Thomson wins; 10%-17% ties
0.18

Notes