Method and limitations
This page sets out how the corpus is built, on what criteria a document gets in, how retrieval works and what known defects it carries. The last part matters as much as the first: a documentary tool that does not state its limits asks you to trust it, and the point here is the opposite.
Figures read from the graph itself when the page loads. The rest of the data is as of 3 August 2026.
Where it comes from, and what gets in
The corpus is fed by three automated routes and one manual one, all of them daily. None of them searches for \"whatever exists about drugs\": each walks forty-four axes defined by discipline, so the whole does not end up being pharmacology with trimmings.
| Route | Documents | What it contributes, and why |
|---|---|---|
| Crossref | 14.472 | Everything with a DOI: humanities, law, botany |
| PubMed / Europe PMC | 11.128 | Peer-reviewed biomedical literature |
| OpenAlex | 2.542 | Social sciences and fields with no biomedical coverage |
| Own library and archive | 3.276 | Grey literature: reports, proceedings, theses, work without a DOI |
| Cannabis Magazine | 700 | Own back catalogue since 1997 |
Inclusion criteria. From the automated routes, only work with a deposited abstract gets in — without text there is nothing to index — and only what passes a relevance filter of field terms applied to title and abstract. Anything already held is dropped by comparing DOI, address and normalised title. There is no impact or journal-prestige criterion: a report from a risk-reduction organisation weighs the same as a paper in Nature, and the reader sees which one each claim comes from.
How it is indexed
Segmentation
Each document is split into overlapping passages of a manageable size. The overlap keeps a claim from being cut in half and lost to retrieval.
Dual index
Each passage is indexed by meaning, using a multilingual embedding model running on our own server, and by literal text. The first finds what is said in other words; the second finds the acronyms and proper nouns an embedding dilutes.
Graph extraction
Concepts are extracted from the passages — substances, receptors, brain regions, effects, clinical conditions — along with the relations between them and the textual evidence that supports each one. That is what lets a question about a receptor also surface what has been written about the plant that activates it.
How retrieval works, and why that bounds invention
A question triggers not one search but several. The query is rewritten and translated into the corpus languages; retrieval runs in parallel by meaning and by literal match; the two rankings are fused; and a local reranking model judges the candidates again on content rather than language, because an embedding scores text in the question's own language higher and that was skewing answers toward the Spanish corpus.
Only then is the answer written, and only from the retrieved passages. The model contributes no knowledge of its own: if the corpus does not hold the answer, the instruction is to say so. That does not eliminate the risk of a claim slipping through, but it bounds it to what is in the cited documents, which are linked below every answer so they can be checked. It is slower than a search engine, and that is the reason.
What coverage there is
Known limitations
What follows are not courtesy warnings: they are measured defects, with their magnitude. Anyone using this tool for research should know them before citing anything.
How to cite it
Cite the corpus rather than the answer, and cite as well the specific sources the answer showed you, which are what support each claim.
Psiconáutica (2026). Corpus documental multidisciplinar sobre sustancias psicoactivas (Noosphere). https://brain.psiconautica.org/corpus.html
If you spot an error, say so
A wrong metadata field, an impossible attribution, a document that should not be there: write to noosphere@psiconautica.org. We are equally interested in literature indexed nowhere that is being lost: grey material, proceedings, work that was not published in English.