Emoji Prediction Datasets
收藏资源简介:
本数据集名为Emoji Prediction Datasets,由达特茅斯学院的研究团队创建,主要用于表情预测任务。数据集包含来自Twitter的1,480,685条推文,每条推文平均包含1.89个表情符号。创建过程中,研究团队首先从Twitter收集数据,然后通过手动设计的启发式方法进行标注。该数据集主要应用于自然语言处理领域,旨在通过预测文本中适当的表情符号,帮助模型学习文本的交流意图,特别是在情感预测、情感分析和讽刺检测等任务中。
This dataset, named Emoji Prediction Datasets, was created by a research team from Dartmouth College and is primarily used for emoji prediction tasks. It contains 1,480,685 tweets sourced from Twitter, with an average of 1.89 emojis per tweet. During its development, the research team first collected data from Twitter, then annotated it using manually designed heuristic methods. Primarily applied in the field of natural language processing, this dataset aims to assist models in learning the communicative intent of text by predicting appropriate emojis within the text, particularly in tasks such as emotion prediction, sentiment analysis, and sarcasm detection.




