Known limitations
The things this system gets wrong, stated plainly. Read this before citing anything here.
This is the honest list. It is here because a graded, confident-looking interface invites more trust than the underlying process has earned.
The labels are machine guesses
Study design, evidence grade, effect direction and outcome category all come from a language model. It can misread a design, invert a direction, or file a claim under the wrong category. There is no human verification step anywhere in the pipeline. If a label matters to your argument, check it against the paper.
Most of the evidence is observational
Only a minority of the corpus uses causal or quasi-experimental identification; randomised trials are a small fraction of that, and only a small share of claims are graded high. The rest is correlational, descriptive or theoretical. The site shows weaker studies rather than hiding them, because context decides how much weight a finding deserves. The methods page carries the current figures.
Tensions are not confirmed contradictions
The tension rule flags opposing directions within an outcome category. It cannot check whether the two claims share a population, a context or an outcome definition. Treat every tension as a prompt to read both papers, never as a settled disagreement.
Effect sizes are sparse and local
Where a specific number appears, it comes from one study in one setting. Effects vary widely across populations, sectors and periods. A single extracted figure is not the field's answer, and averaging extracted figures across incomparable studies would be worse than quoting one.
Coverage is partial, and starts in late 2025
The pipeline indexes a slice of English-language academic work from a handful of sources. Working papers arrive through preprint feeds; published articles can lag by months. Good research will sometimes simply be missing.
More surprising than the breadth is the start date: collection began in late 2025, so most of the earlier literature in this field is absent, including work by authors you would expect to find. Because ranking is semantic, a search always returns a full page of nearest neighbours, so nothing in the interface signals that absence. See the corpus start date.
Search always answers, even when it should not
Site search ranks semantically, which means it returns the closest papers it has rather than only papers that match. That is usually what you want, and it is occasionally misleading: a query with no good answer in the corpus produces a confident-looking page of loosely related work rather than an empty result. Check that the papers returned actually address your question before concluding the literature says something.
Syntheses age
A synthesis reflects the evidence available the day it was written. New papers arrive constantly, so pages are marked as having an update pending when fresh evidence has landed since. Even a current synthesis only covers what cleared the relevance bar at generation time.
Author identities are corpus records
Exact provider IDs and unambiguous ORCIDs support conservative links, but they do not prove personhood. Names alone never merge identities, so unresolved homonyms stay separate. Provider coverage differs, which changes paper and collaborator counts when the provider selector changes.
A reference-check miss proves nothing
The reference checker searches an economics corpus that starts in 1995. A miss means unverified, not fabricated. See what a miss actually means.
Found something wrong?
Mislabelled papers and bad grades are the most useful thing you can report, because they are invisible from the inside. Get in touch.