遇见数据集

RedMed: Extending drug lexicons for social media applications

收藏
Zenodo2019-06-05 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

Data associated with the RedMed project. Details for the process behind the data creation can be found in the associated paper: <strong>Lavertu, A. &amp; Altman, R. B. </strong>"RedMed: Extending drug lexicons for social media applications"<br> Journal of Biomedical Informatics, (2019) https://doi.org/10.1016/j.jbi.2019.103307 <strong>RedMed embedding model:</strong> Word vectors trained on comments from health related subreddits and optimized for drug synonym retrieval. The Redmed model was train using only social media data from Reddit and achieves comparable performance on the UMNSRS and MayoSRS similarity tasks. Vectors are 64 dimensional. <strong><strong>redmed_model_vectors.tsv.gz - </strong></strong>Tab-separated word vectors (token\tdim1\tdim2\t...dim64) <strong><strong>redmed_model.bin - </strong></strong>Binary word2vec file saved using gensim, can be loaded into python gensim Other Files: supp_file_1_sidebar_subreddits.txt - List of health-related subreddits based on "r/Health" and "r/Drugs" sidebars<br> supp_file_2_enrichment_based_subreddits.txt - List of health-related subreddits based on amount of health-related content<br> supp_file_3_custom_stopword_list.txt - List of stopwords based on counts derived from Reddit comments

提供机构:
Stanford University
创建时间:
2019-06-05
二维码
社区交流群
二维码
科研交流群
商业服务