For researchers

Method and limitations

This page sets out how the corpus is built, on what criteria a document gets in, how retrieval works and what known defects it carries. The last part matters as much as the first: a documentary tool that does not state its limits asks you to trust it, and the point here is the opposite.

32,205documents
317,864indexed passages
124,676concepts
170,764relations

Figures read from the graph itself when the page loads. The rest of the data is as of 3 August 2026.

Where it comes from, and what gets in

The corpus is fed by three automated routes and one manual one, all of them daily. None of them searches for \"whatever exists about drugs\": each walks forty-four axes defined by discipline, so the whole does not end up being pharmacology with trimmings.

Route Documents What it contributes, and why
Crossref14.472Everything with a DOI: humanities, law, botany
PubMed / Europe PMC11.128Peer-reviewed biomedical literature
OpenAlex2.542Social sciences and fields with no biomedical coverage
Own library and archive3.276Grey literature: reports, proceedings, theses, work without a DOI
Cannabis Magazine700Own back catalogue since 1997

Inclusion criteria. From the automated routes, only work with a deposited abstract gets in — without text there is nothing to index — and only what passes a relevance filter of field terms applied to title and abstract. Anything already held is dropped by comparing DOI, address and normalised title. There is no impact or journal-prestige criterion: a report from a risk-reduction organisation weighs the same as a paper in Nature, and the reader sees which one each claim comes from.

How it is indexed

  1. Segmentation

    Each document is split into overlapping passages of a manageable size. The overlap keeps a claim from being cut in half and lost to retrieval.

  2. Dual index

    Each passage is indexed by meaning, using a multilingual embedding model running on our own server, and by literal text. The first finds what is said in other words; the second finds the acronyms and proper nouns an embedding dilutes.

  3. Graph extraction

    Concepts are extracted from the passages — substances, receptors, brain regions, effects, clinical conditions — along with the relations between them and the textual evidence that supports each one. That is what lets a question about a receptor also surface what has been written about the plant that activates it.

How retrieval works, and why that bounds invention

A question triggers not one search but several. The query is rewritten and translated into the corpus languages; retrieval runs in parallel by meaning and by literal match; the two rankings are fused; and a local reranking model judges the candidates again on content rather than language, because an embedding scores text in the question's own language higher and that was skewing answers toward the Spanish corpus.

Only then is the answer written, and only from the retrieved passages. The model contributes no knowledge of its own: if the corpus does not hold the answer, the instruction is to say so. That does not eliminate the risk of a claim slipping through, but it bounds it to what is in the cited documents, which are linked below every answer so they can be checked. It is slower than a search engine, and that is the reason.

What coverage there is

By decadeFrom the 1950s to today, very unevenly: 32 documents from the fifties, 1,418 from the nineties, 8,503 from the 2010s and 14,930 from this decade. The oldest with documentary value is from 1949, Moruzzi and Magoun on the reticular formation.
By disciplineForty-four axes with 17,536 documents tagged by discipline on entry, between 550 and 750 per axis: mycology, risk reduction, ethnobotany, analytical chemistry, conservation, phytochemistry, agronomy, religious studies, sociology, pharmacognosy, clinical psychology and toxicology leading.
By journalMore than 6,000 distinct journals after merging the variant spellings each source uses for the same title. None accounts for more than 1% of the whole.
Depth5,291 documents with full or extensive text; the remaining 26,914 with abstract and metadata. The proportion is a limitation rather than a choice: see below.

Known limitations

What follows are not courtesy warnings: they are measured defects, with their magnitude. Anyone using this tool for research should know them before citing anything.

Severe language bias30,635 documents in English against 1,236 in Spanish and token figures in Portuguese, German and French. 95% of the corpus comes from the anglophone tradition, with everything that implies about which questions get treated as researchable. Retrieval is multilingual, but it cannot retrieve what is not there.
Reliance on abstracts84% of documents enter with title, abstract and metadata only, because third-party full text is not redistributed. An abstract omits the method, the effect sizes and almost always the study's own limitations. Judging a piece of work means opening the original, which is why every answer links to it.
Metadata that arrives wrong at sourceA 1960 paper on psilocybin is listed as published in Corrosion Science. We checked it against the Crossref API: the error is in the publisher's deposit, not in our matching. There are a handful of cases like this and they cannot be fixed from here.
Our own mismatchesWhen assigning DOIs to archive documents that lacked one, twelve came out wrong: two papers on Banisteriopsis caapi point to a course on vegetative propagation. They are identified and flagged for manual review.
Inclusion is by keywordA term filter does not understand what a text is about. Measured against the body of the documents, around 15% mention no substance or discipline from the field; many are neuroscience method papers that serve as background, but others simply do not belong. The bias runs both ways: relevant material written in another vocabulary is also discarded.
Publication and recency biasThe corpus inherits what the literature publishes: positive results over negative ones, and almost half the documents dated in this decade. A field that reawakened fifteen years ago produces a corpus that looks more like the present than like its own history.
The answer is not reproducibleIt is generated on the spot and two identical queries may be worded differently. What is reproducible are the sources, and those are what belongs in a citation. Citing the answer would be citing a paraphrase.
No peer review and no expert curationNobody reviews what comes in document by document. The controls are automated — relevance, duplicates, integrity — and audited periodically, but they do not amount to the curation of an institutional repository. This project is not run by a university.

How to cite it

Cite the corpus rather than the answer, and cite as well the specific sources the answer showed you, which are what support each claim.

Psiconáutica (2026). Corpus documental multidisciplinar sobre sustancias psicoactivas (Noosphere). https://brain.psiconautica.org/corpus.html

If you spot an error, say so

A wrong metadata field, an impossible attribution, a document that should not be there: write to noosphere@psiconautica.org. We are equally interested in literature indexed nowhere that is being lost: grey material, proceedings, work that was not published in English.