AutoReLUNMF — PubMed corpora and topic factors
收藏官方服务:
资源简介:
SQLite databases and pre-computed sentence-transformer embeddings for the five corpora used in A Fast Factorization Pipeline for Topic Modeling on Sentence-Transformer Embeddings (Berber, Karayağız, Siyah). Includes the four PubMed slices (nutri ~530k, oncology ~50k, cardio ~100k, precision ~400k) and the 20 Newsgroups embedding cache. Title and abstract text are public-domain PubMed records fetched through the NCBI E-utilities and stored in the schema produced by experiments/fetch_pubmed.py in the companion code repository.
提供机构:
Zenodo创建时间:
2026-05-13



