遇见数据集

AutoReLUNMF — PubMed corpora and topic factors

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

SQLite databases and pre-computed sentence-transformer embeddings for the five corpora used in A Fast Factorization Pipeline for Topic Modeling on Sentence-Transformer Embeddings (Berber, Karayağız, Siyah). Includes the four PubMed slices (nutri ~530k, oncology ~50k, cardio ~100k, precision ~400k) and the 20 Newsgroups embedding cache. Title and abstract text are public-domain PubMed records fetched through the NCBI E-utilities and stored in the schema produced by experiments/fetch_pubmed.py in the companion code repository.

提供机构:
Zenodo
创建时间:
2026-05-13
二维码
社区交流群
二维码
科研交流群
商业服务