CoAID dataset texts with OCR degradations
收藏资源简介:
This is the text of the CoAID dataset dedicated to fake news detection that has been updated to be used in event detection. Cui, Limeng, et Dongwon Lee. 2020. « CoAID: COVID-19 Healthcare Misinformation Dataset ». <em>ArXiv:2006.00885 [Cs]</em>, novembre. http://arxiv.org/abs/2006.00885. Guillaume Bernard. (2022). CoAID dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630405 Some degradations are applied using the DocCreator [1] tool in order to degrade the text of the tweets and to reproduce some common errors found in OCRised documents [2]. [1]: Journet, Nicholas, Muriel Visani, Boris Mansencal, Kieu Van-Cuong, et Antoine Billy. 2017. « DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images ». <em>Journal of Imaging</em> 3 (4): 62. https://doi.org/10.3390/jimaging3040062. [2]: Linhares Pontes, Elvys, Ahmed Hamdi, Nicolas Sidere, et Antoine Doucet. 2019. « Impact of OCR Quality on Named Entity Linking ». In <em>Digital Libraries at the Crossroads of Digital Information for the Future</em>, 11853:102‑15. Lecture Notes in Computer Science. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-34058-2_11. The results of the OCR degradations are as follow: CoAID CER/WER Without Character degradation Phantom degradation Bleed Blur All CoAID CER 2.105 6.358 2.105 2.122 2.616 7.898 CoAID WER 2.494 20.230 2.496 2.580 3.726 20.230
本数据集为专为虚假新闻检测设计的CoAID数据集(CoAID dataset),经更新后可用于事件检测。 Cui, Limeng 与 Dongwon Lee. 2020年. 《CoAID: COVID-19 Healthcare Misinformation Dataset》,发表于*ArXiv:2006.00885 [Cs]*,2020年11月。访问链接:http://arxiv.org/abs/2006.00885。 Guillaume Bernard.(2022). 含多提取特征(稀疏特征与稠密特征)的CoAID数据集(版本1.0)[数据集]. Zenodo. https://doi.org/10.5281/zenodo.6630405。 为还原经光学字符识别(Optical Character Recognition, OCR)处理的文档中常见的各类典型错误,本研究使用DocCreator工具[1]对推文文本进行多种退化处理。 [1] Journet, Nicholas, Muriel Visani, Boris Mansencal, Kieu Van-Cuong 与 Antoine Billy. 2017年. 《DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images》,发表于*Journal of Imaging* 3(4): 62。https://doi.org/10.3390/jimaging3040062。 [2] Linhares Pontes, Elvys, Ahmed Hamdi, Nicolas Sidere 与 Antoine Doucet. 2019年. 《Impact of OCR Quality on Named Entity Linking》,收录于《Digital Libraries at the Crossroads of Digital Information for the Future》,属于《计算机科学讲义(Lecture Notes in Computer Science)》丛书第11853卷,Cham:施普林格国际出版公司,第102-115页。https://doi.org/10.1007/978-3-030-34058-2_11。 OCR退化处理的效果如下: | 数据集 | 指标 | 无字符退化 | 伪影退化 | 透印退化 | 模糊退化 | 全部退化 | | ------ | ---- | ---------- | -------- | -------- | -------- | -------- | | CoAID | 字符错误率(Character Error Rate, CER) | 2.105 | 6.358 | 2.105 | 2.122 | 2.616 | 7.898 | | CoAID | 词错误率(Word Error Rate, WER) | 2.494 | 20.230 | 2.496 | 2.580 | 3.726 | 20.230 |



