PaperAtlas

PaperAtlas is an atlas of the computational methods and software described in the open-access biomedical literature, built in a single automated pass over the PubMed Central full-text archive.

By the numbers

Articles screened6,446,741
Papers where computation is a central result1,074,191
Schema-valid structured records1,074,140
Clusters in the atlas1,000
Papers in the atlas165,432
Distinct extracted name strings121,844
Of those, with no strict name match in bio.tools, PyPI, CRAN, Bioconductor or Bioconda93,762 (77.0%)
Name strings from software package and web server papers with no strict match19,031 (61.0% of 31,180)
Distinct dataset accessions285,615
Distinct Git repository links77,382

The atlas is free to browse. The published files are the September 2026 release, the version described in the manuscript. Releases are versioned by date and earlier releases are kept.

How it is built

The PaperAtlas pipeline: JATS XML parsing, abstract screening, guided-JSON full-text extraction, embedding and density-based clustering, restriction to biomedical topics, and static-site serving; with the acquisition funnel, publication-year distribution and screening outcome.
The pipeline and the corpus. Every article in the PubMed Central bulk archive is parsed, every abstract is screened for whether computation is a central result, and a structured record is extracted from each retained full text. The funnel runs from 7,283,594 articles parsed to the 165,432 retained in the atlas.
The atlas: a UMAP projection of tool-paper embeddings coloured by parent topic, the cluster-size distribution, EDAM ontology alignment per cluster, and curated EDAM topic concentration, and internal validation measures.
The atlas and its validation. Papers that contribute an algorithm, software package or web server are embedded and clustered, and 1,000 clusters are retained as biomedical. Where a cluster links to at least five bio.tools entries, 61.7% of them share its most frequent curated EDAM topic, against 26.4% when clusters are permuted. Of 31,180 name strings from software package and web server papers, 61.0% have no strict name match in five registry snapshots.

Citation

Shah N, Vippagunta A, Bhargava Y. PaperAtlas: an automatically constructed atlas of computational methods and software from 6.4 million open-access articles. 2026.