iSarcasm
收藏资源简介:
iSarcasm数据集由爱丁堡大学信息学院的Silviu Vlad Oprea和Walid Magdy创建,包含4484条由作者直接标记为讽刺的英语推文。数据集旨在解决现有讽刺检测数据集可能存在的偏差问题,鼓励未来NLP研究开发更准确地捕捉文本作者意图的讽刺检测方法。数据集中的每条讽刺推文都附有作者提供的讽刺解释和非讽刺表达方式,适用于研究讽刺的编码和解码,以及讽刺类别预测等任务。
The iSarcasm dataset was created by Silviu Vlad Oprea and Walid Magdy from the School of Informatics, University of Edinburgh. It contains 4,484 English tweets directly labeled as sarcasm by their respective authors. This dataset aims to address the potential biases inherent in existing sarcasm detection datasets, and encourages future NLP research to develop sarcasm detection methods that more accurately capture the intended meanings of text authors. Each sarcastic tweet in the dataset is accompanied by a sarcasm explanation and a non-sarcastic rephrasing provided by the original author, making it suitable for research tasks including sarcasm encoding and decoding, as well as sarcasm category prediction.




