Method and limitations
This page sets out how the corpus is built, on what criteria a document gets in, how retrieval works and what known defects it carries. The last part matters as much as the first: a documentary tool that does not state its limits asks you to trust it, and the point here is the opposite.
Figures read from the graph itself when the page loads. The rest of the data is as of 7 September 2026.
Where it comes from, and what gets in
The corpus is fed by three automated routes and one manual one, all of them daily. None of them searches for \"whatever exists about drugs\": each walks forty-four axes defined by discipline, so the whole does not end up being pharmacology with trimmings.
| Route | Documents | What it contributes, and why |
|---|---|---|
| Crossref | 39.374 | Everything with a DOI: humanities, law, botany |
| PubMed / Europe PMC | 5.385 | Peer-reviewed biomedical literature |
| OpenAlex | 2.538 | Social sciences and fields with no biomedical coverage |
| Own library and archive | 3.286 | Grey literature: reports, proceedings, theses, work without a DOI |
| Cannabis Magazine | 700 | Own back catalogue since 1997 |
Inclusion criteria. From the automated routes, only work with a deposited abstract gets in — without text there is nothing to index — and only what passes a relevance filter of field terms applied to title and abstract. Anything already held is dropped by comparing DOI, address and normalised title. There is no impact or journal-prestige criterion: a report from a risk-reduction organisation weighs the same as a paper in Nature, and the reader sees which one each claim comes from.
A document gets in if it deals with a psychoactive substance or the field around it: pharmacology, drug policy, harm reduction, addiction, forensic toxicology, the anthropology and history of use. In September 2026 everything collected up to then was reviewed against that criterion and 62,929 documents were set aside: they had got in on a single generic word in their abstracts — “ritual”, “conservation”, “taxonomy”, “drug” — and were not about psychoactives at all: philosophy of law, agronomy, materials chemistry. Nothing was deleted, but they are neither searched nor counted. Material added by hand (library, archive, books and the magazine back catalogue) does not go through that filter: there a person made the selection.
How it is indexed
Segmentation
Each document is split into overlapping passages of a manageable size. The overlap keeps a claim from being cut in half and lost to retrieval.
Dual index
Each passage is indexed by meaning, using a multilingual embedding model running on our own server, and by literal text. The first finds what is said in other words; the second finds the acronyms and proper nouns an embedding dilutes.
Graph extraction
Concepts are extracted from the passages — substances, receptors, brain regions, effects, clinical conditions — along with the relations between them and the textual evidence that supports each one. That is what lets a question about a receptor also surface what has been written about the plant that activates it.
Authors and full text
Each document’s authors are collected from OpenAlex, Crossref and PubMed on ingestion, and a nightly pass fills in any that are missing (97% of the corpus has them today). Full text is stored only when the article is open access: every night the ones that are, and were still abstract-only, are fetched from Europe PMC. Anything not open access stays as an abstract, out of respect for the rights of whoever published it.
How retrieval works, and why that bounds invention
A question triggers not one search but several. The query is rewritten and translated into the corpus languages; retrieval runs in parallel by meaning and by literal match; the two rankings are fused; and a local reranking model judges the candidates again on content rather than language, because an embedding scores text in the question's own language higher and that was skewing answers toward the Spanish corpus.
Only then is the answer written, and only from the retrieved passages. The model contributes no knowledge of its own: if the corpus does not hold the answer, the instruction is to say so. That does not eliminate the risk of a claim slipping through, but it bounds it to what is in the cited documents, which are linked below every answer so they can be checked. It is slower than a search engine, and that is the reason.
There are two more modes besides the written answer. “Search documents” returns the list of documents without writing anything, with filters by year, evidence level, source, author and journal, and can be exported to BibTeX, RIS or CSV: it is the same retrieval as above, minus the final step. And the “extended answer” (with an account) retrieves fourteen sources instead of six and writes section by section, for someone doing a review rather than settling a quick doubt.
Evidence level: how it is assigned
Every answer states what kind of studies it was built from, and the search can be restricted to whichever levels you want. The scale is the usual evidence-based medicine pyramid, in five levels: high-quality synthesised evidence (systematic reviews and meta-analyses), experimental studies (randomised controlled trials), analytical observational (cohort, case-control), descriptive and series (cross-sectional, surveys, case series and reports) and, at the base, preclinical and opinion (in vitro, animal models, editorials and narrative reviews).
Assignment happens in two steps, from most reliable to least, and never by eye. First we take the PublicationType tag that indexers at the National Library of Medicine assign to every PubMed article; we fetch it by identifier and, for articles arriving through other routes, by matching the DOI against PubMed. That classification is the work of human librarians, not ours and not a model's. Where the tag is missing — and it is missing for close to half of articles, which declare only «Journal Article» — we fall back on how the abstract describes the study itself: «randomised, double-blind», «we report the case of», «in mice». The route used for each document is recorded.
A large share of the corpus ends up with no level assigned, and it is shown that way on purpose. This is not a flaw in the method: a substantial part of these documents come from analytical chemistry, botany, ethnography, law or history, fields where the clinical evidence pyramid means nothing. Forcing them into a bucket would manufacture a rigour that isn't there, and hiding how many they are would inflate the apparent solidity of every answer. Better to leave the gap visible.
Two warnings for anyone leaning on this. The tighter the filter, the fewer documents remain: at the higher levels, and on little-studied topics, the system can run out of material — and that absence is itself a finding. And the level describes the study's design, not how well it was carried out nor how relevant it is to your question: a badly run randomised trial is still level 2.
What coverage there is
Known limitations
What follows are not courtesy warnings: they are measured defects, with their magnitude. Anyone using this tool for research should know them before citing anything.
How to cite it
Cite the corpus rather than the answer, and cite as well the specific sources the answer showed you, which are what support each claim.
Every query is linkable: the «Link to this query» button under each answer copies an address that reruns the same question with the same evidence filter (for example ?q=…&niveles=1,2). Anyone opening that link watches the search run again and can check where each figure comes from. Worth including when you cite: the corpus grows every night, so today’s answer may rest on more documents tomorrow, and the link records exactly what was asked.
Under every answer you will find the date it was generated and the size of the corpus that day (for example, “from a corpus of 105,911 documents, last updated 2026-09-03”). That is the figure worth giving when citing: it says which version of the corpus answered.
Psiconáutica (2026). Corpus documental multidisciplinar sobre sustancias psicoactivas (Noosphere). https://brain.psiconautica.org/corpus.html
If you spot an error, say so
A wrong metadata field, an impossible attribution, a document that should not be there: write to noosphere@psiconautica.org. We are equally interested in literature indexed nowhere that is being lost: grey material, proceedings, work that was not published in English.