遇见数据集

Word Embedding Data Sets Learned From Tweets And General Data

收藏
Zenodo2020-09-18 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This includes 10 word embedding data sets learned from about 400 million tweets and 7 billion words from general data. They can be used in tasks involving social media data, especially tweets, and other types of textual data. Users can choose different embedding sets based on their use cases; they can also easily try all of them to see which one provides the best performance for their application. More details about the training data collection, word embedding generation, preprocessing steps, and how to use them can be found from the following paper: Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh, Data Set: Word Embeddings Learned from Tweets and General Data, The 11th International AAAI Conference on Web and Social Media (ICWSM-17). Montreal, Canada. May 16-18, 2017

本数据集包含10个词嵌入(word embedding)数据集,均基于约4亿条推文(tweets)与通用语料中的70亿词汇训练得到。该系列数据集可应用于涉及社交媒体数据(尤其是推文)及其他类型文本数据的各类任务。用户可根据自身应用场景选择适配的嵌入集,也可便捷地逐一测试全部嵌入集,以筛选出在其特定应用中表现最优的方案。 有关训练数据采集流程、词嵌入生成方法、预处理步骤以及数据集使用方式的更多细节,可参阅以下论文: Quanzhi Li、Sameena Shah、Xiaomo Liu、Armineh Nourbakhsh:《数据集:从推文与通用数据中学习得到的词嵌入》,第11届国际AAAI网络与社交媒体会议(ICWSM-17),加拿大蒙特利尔,2017年5月16日至18日。

提供机构:
Zenodo
创建时间:
2017-05-22
二维码
社区交流群
二维码
科研交流群
商业服务