遇见数据集

Word Embeddings Learnt On Medline Abstracts

收藏
Zenodo2020-09-18 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Accompanying a preprint manuscript and code repository, this folder contains both raw text data and learnt word embeddings. The data source is the set of MEDLINE articles published on or after 2000. Preprocessing consists of extraction of each article's title and abstract and some minor text processing. The result is a corpus of 10.5 million documents in a single 14 GB file. word2vec and fastText are used to learn word embeddings on this corpus and three sets of word embeddings are shared here: 1) word2vec skip-gram, 2) word2vec CBOW, and 3) fastText skip-gram. All three sets use the default parameters of the software (e.g. context=5) with the exception of hierarchical softmax optimization and dimension=200. Preprint manuscript: https://arxiv.org/abs/1705.06262<br> GitHub repository: https://github.com/vincentmajor/ctsa_prediction

提供机构:
Zenodo
创建时间:
2017-06-05
二维码
社区交流群
二维码
科研交流群
商业服务