Technical sheet

The Noosphere corpus

Noosphere does not search the internet: it answers from a document base of its own. This page states exactly what that base contains, where each part came from and how it was built, so anyone can judge the worth of an answer before trusting it.

8,802documents
266,844indexed passages
49,369concepts
62,998relations

Live figures, read from the graph itself when the page loads. The breakdown below is as of 31 July 2026.

Where it comes from

No single source dominates, and that is deliberate. Biomedical literature brings experimental rigour; the project's own archive brings what rarely reaches an indexed journal — harm-reduction reports, conference proceedings, dissertations, specialist press; OpenAlex brings the fields a biomedical search never reaches, such as anthropology or law.

Origin Documents What it contributes
PubMed / Europe PMC3.973Peer-reviewed biomedical literature
Own library2.017Reference books and monographs
Document archive1.268Reports, proceedings, theses and grey literature
OpenAlex753Social sciences, botany, law
Cannabis Magazine700Own back catalogue, 1997 onwards
Open and reference web80Public bodies and organisations

Scope

LanguagesEnglish (7,550), Spanish (1,160), Portuguese, Catalan, German and Dutch. Ask in any language: retrieval runs across all of them and the answer is written in the language of the question.
Time spanFrom 1953 — the year of Howard Becker's "Becoming a Marihuana User" — to material published ahead of print for 2027. 6,243 documents carry a stated year; the rest are materials with no editorial date, such as internal reports or archive holdings.
DisciplinesTwenty classified fields, led by psychology, medicine, social sciences, agricultural and biological sciences, pharmacology and toxicology, biochemistry, neuroscience and the humanities.
SubstancesEvery psychoactive substance, not only the psychedelics: cannabis, opioids, stimulants, depressants, dissociatives, plants and fungi of traditional use, and preparations of animal origin.
JournalsMore than 1,600 distinct journals, counted after merging the variant spellings each source uses for the same title. None accounts for more than 3% of the corpus: the best represented is Psychopharmacology, with 169 papers.

How it is built

Automated harvesters query PubMed, Europe PMC and OpenAlex every week across forty-four thematic axes, from receptor pharmacology to botanical taxonomy and drug policy. What they find is filtered for relevance, dropped if already held, and anything new is split into passages, indexed by meaning and by word, and analysed to extract the concepts it mentions and the relations between them. That graph is what lets a question about a receptor also surface what has been written about the plant that activates it.

Answering is not a single search: the question is rewritten, retrieval runs along two routes — meaning and literal match — the results are fused, and a local reranking model decides which passages earn a place in the answer. It is slower than a search engine, and it is why the citations match what is claimed.

What you can and cannot read here

The index, the bibliographic metadata and the citations are freely consultable. Noosphere does not republish the full text of third-party works: when an answer rests on a document, it links to the original so it can be read at source. Full text is kept only when the work is open access or belongs to the project itself.

How to cite it

If you use Noosphere in your work, cite the corpus rather than the answer: answers are generated on the spot and are not reproducible word for word. The sources they cite are, and those are what belongs in an academic citation.

Psiconáutica (2026). Corpus documental multidisciplinar sobre sustancias psicoactivas (Noosphere). https://brain.psiconautica.org/corpus.html

Contributing documents

The corpus grows mainly by contribution. If you hold relevant literature — theses, agency reports, conference proceedings, harm-reduction material, back catalogue — write to noosphere@psiconautica.org and we will review it before taking it in. We are especially interested in what is indexed nowhere and gets lost: the grey, the local, and what was not published in English.

Ask it something you already know

It is the only honest way to measure a corpus: ask a question whose answer you already know, and check where it says it got it.