0 cumulative citations
View corpus contextA small team converted open-weight checkpoints into Thomson, a near-frontier foundation model family, using continual learning to match or surpass many proprietary models in legal, tax and research tasks at substantially lower reported compute cost; transparent benchmarking and human evaluation back the claim, but limited disclosure of data and baseline conditions tempers reproducibility.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but offers little concrete advice on how this can be achieved in the short term under a diversity of funding settings. We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models. Unlike limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation of a frozen model, our approach exploits a modern mid- & post-training stack while introducing safeguards that preserve both plasticity and stability at each stage, making the minimal number of high-impact interventions on the parameters. This yields gains comparable to those typically seen across multiple successive model generations, at compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for far more actors. We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work. Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research. Evaluations show a distinctive $π$-shaped pattern: distinct improvements across a wide range of capabilities, including those not explicitly targeted, while almost completely eliminating the forgetting problem common to narrow domain adaptation.
Summary
Main Finding
Thomson demonstrates that “frontier” performance can be achieved by non-hyperscale actors by applying a disciplined Continual Learning (CL) pipeline to existing open-weight foundation models. Using full‑weight mid- and post‑training (not tiny fine‑tuning or distillation), careful data curation, agentic/tool training and bespoke reward design, a small engineering team produced Thomson-1.0 models that approach or match many contemporaneous proprietary frontier models across a broad benchmark suite — at substantially lower GPU and personnel cost than full pretraining.
Key Points
- Continual Learning as a practical SovereignAI route:
- CL pipeline preserves stability (no catastrophic forgetting) while enabling plasticity (new capabilities); authors report a distinctive “π-shaped” pattern: strong gains in targeted domains plus improvements across many untargeted capabilities.
- Base models and update strategy:
- Built from open-weight checkpoints (Qwen3.5‑397B and Qwen3.6‑35B), performing full‑weight updates (not parameter‑efficient adapters) and avoiding large-scale distillation.
- Focus domains and strengths:
- Targeted at high‑stakes professional use (legal, tax, journalism, Deep Research, document/RAG workflows). Strong on instruction following, summarization, multilingualism, citation quality and document processing.
- Costs, team and timeline:
- Development run with a small team (≈3 dozen engineers/scientists), max ~368 B200 GPUs during experiments; final three‑week training GPU cost estimated < USD 450k; total development cost ~USD 40M (includes research, infrastructure, domain experts) and a three‑month intensive development timeline from first experiments.
- Safety/value work:
- Value‑realignment modules (Snowdon variants) and adversarial testing reported; permissive-licensing emphasis for training data; bespoke reward structures to encourage faithful tool use and citation to reduce hallucinations.
- Limitations noted:
- Relative weakness in coding and some abstract reasoning vs top proprietary models; evaluation vs closed APIs may be imperfect because of opaque guardrails in those systems.
- Reproducibility/portability claim:
- Authors argue the approach generalizes beyond their exact checkpoints and sizes and can be replicated by other institutions with similar resources.
Data & Methods
- Overall approach:
- Three complementary CL modules producing foundation models via mid‑training and post‑training stages. Emphasis on minimal, high‑impact parameter interventions and preserving previously learned skills.
- Data‑centric practices:
- Rigorous curation: filtering, rephrasing, de‑duplication, semantic targeting of human subject‑matter collection, and Bayesian optimisation to calibrate post‑training data mixtures.
- Training methods:
- Full‑weight updates on open checkpoints (no reliance on a stronger teacher), reward design for tool/agentic behavior, localised error correction and end‑to‑end RL for Deep Research tasks.
- Tool and agentic training:
- Designed to integrate private retrieval and tool stacks; Deep Research harness for high‑budget tasks; training encourages faithful citations and conditioned tool usage.
- Evaluation:
- Large, mixed suite spanning legal benchmarks (e.g., Stanford LegalBench, Harvey LAB), tax, journalism, summarisation, long context, multilingualism, safety/adversarial tests, and human blind preference tests (~3,000 rated conversations). Benchmarks ran under controlled, identical prompt and inference settings where possible; closed models evaluated as served endpoints (acknowledging unknown intermediate guardrails).
- Key experimental numbers:
- Thomson-1.0-Small: 35B parameter open-weight release. Thomson‑1.0‑Large: full‑weight updated large model based on Qwen3.5‑397B checkpoint.
- Reported GPU budget peak: 368 B200 GPUs; final run GPU cost < USD 450k; total development ≈ USD 40M.
Implications for AI Economics
- Lowers technical & economic barriers to building near‑frontier models:
- Demonstrates that competent teams with modest compute (vs full pretraining budgets) can produce competitive domain‑specialist or even broadly capable models by investing in CL, data quality and tooling.
- Reduces exclusive advantage of hyperscalers:
- Shifts some value capture from scale-of-pretraining to data, tooling, bespoke reward design, and continual refinement — enabling more institutions to claim SovereignAI (model/data/governance control).
- New cost structure & business models:
- Organizations can prioritize investment in mid/post‑training pipelines, evaluation suites, and private tool integration rather than massive pretraining runs; this supports differentiated, verticalized AI products (e.g., legal/tax agents) with lower incremental cost.
- Competitive dynamics:
- Expect increased decentralization and competition in specialized/high-value segments (enterprises, governments, research labs). Proprietary incumbents may retain advantages in raw scale, hardware access, and opaque system‑level safety measures.
- Regulatory & governance implications:
- Wider access to powerful models raises policy relevance: licensing/data provenance, auditability of continual updates, and institutional responsibility for safety/alignment become central. Dependence on open‑weight checkpoints and accelerator hardware concentrated supply chains remain systemic risks.
- Strategic guidance for institutions:
- Invest in CL pipelines, data‑centric processes, evaluation and tool sovereignty to achieve high impact at bounded cost.
- Secure reliable hardware access and plan for ongoing update cycles; treat model sovereignty as a stack (training + data + tools + governance).
- Risks and caveats for markets:
- Outcome depends on access to permissively licensed base models; closed‑API guardrails and upstream availability of open checkpoints could constrain adoption.
- Potential for uneven replication: results are encouraging but may be sensitive to team skill, research infrastructure, and the exact base checkpoints used.
Summary judgment: Thomson provides a compelling, empirically supported blueprint that materially changes the economics of achieving frontier or near‑frontier capabilities for a broad class of institutions. The key economic insight is that focused continual improvement, rigorous data design, and tool integration can substitute for orders‑of‑magnitude pretraining scale in many commercially important domains — shifting where value and competitive advantage reside in the AI ecosystem.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The Thomson family was developed by a technical team of no more than three dozen engineers and scientists, using a compute cluster with no more than 368 B200 GPUs available at any stage of experimentation. Organizational Efficiency | positive | Development personnel and compute resources required to build the model |
Reading fidelity
high
Study strength
medium
|
no more than three dozen engineers; no more than 368 B200 GPUs
|
| The development timeline for Thomson-1.0-Small and Thomson-1.0-Large, measured from the first experiments with Qwen3.5-397B, was three months. Organizational Efficiency | positive | Model development time |
Reading fidelity
high
Study strength
medium
|
three months
|
| The final training run for Thomson-1.0-Large cost under USD 450,000 in GPU costs over three weeks of training. Organizational Efficiency | positive | GPU cost of the final model training run |
Reading fidelity
high
Study strength
medium
|
under USD 450,000
|
| Thomson-1.0-Large achieved an overall benchmark average of 78.5%, compared with 73.6% for its Qwen3.5-397B base model, a difference of 4.9 percentage points. Output Quality | positive | Overall model benchmark performance |
Reading fidelity
high
Study strength
medium
|
78.5% vs 73.6% (4.9 percentage-point difference)
|
| Thomson-1.0-Large outperformed Gemini 3.1 Pro, Sonnet 5, and GLM-5.2 on the reported overall average benchmark score. Output Quality | positive | Overall comparative benchmark performance |
Reading fidelity
high
Study strength
medium
|
78.5% vs 78.0%, 75.7%, and 77.2%
|
| On the legal-domain benchmark average, Thomson-1.0-Large scored 78.4%, compared with 72.7% for Qwen3.5-397B. Output Quality | positive | Legal-domain benchmark performance |
Reading fidelity
high
Study strength
medium
|
78.4% vs 72.7% (5.7 percentage points)
|
| Thomson-1.0-Small achieved an overall benchmark average of 74.6%, compared with 71.7% for Qwen3.6-35B. Output Quality | positive | Overall small-model benchmark performance |
Reading fidelity
high
Study strength
medium
|
74.6% vs 71.7% (2.9 percentage-point difference)
|
| Thomson-1.0-Large scored 97.3% on political neutrality, compared with 85.3% for Qwen3.5-397B. Ai Safety And Ethics | positive | Political neutrality score |
Reading fidelity
high
Study strength
medium
|
97.3% vs 85.3% (12.0 percentage points)
|
| Coding was the only reported domain showing mild forgetting for Thomson-1.0-Large relative to its base model, while general mathematical and abstract reasoning remained within the Qwen performance level. Skill Obsolescence | mixed | Retention of capabilities during continual learning, especially coding and reasoning |
Reading fidelity
high
Study strength
medium
|
coding: 39.9% vs 40.9%
|
| In blind human evaluations covering more than 3,000 preference-rated conversations, Thomson-1.0-Large won between 53% and 62% of comparisons against the listed external frontier systems, with ties ranging from 10% to 17%. Output Quality | positive | Human preference in pairwise system comparisons |
Reading fidelity
high
Study strength
medium
|
n=3000
53%-62% Thomson wins; 10%-17% ties
|