CMU-MisCOV19
收藏资源简介:
CMU-MisCOV19数据集由卡内基梅隆大学创建,专注于COVID-19相关的错误信息分析。该数据集包含4573条经过标注的推文,覆盖了多种信息和错误信息类型。数据收集过程中使用了特定的关键词和时间点,确保了数据的相关性和时效性。创建过程中,数据经过多轮标注和分类,最终形成了包含17个分类的详细代码本。该数据集主要用于研究COVID-19错误信息社区的网络结构、语言模式及其在其他错误信息子社区中的成员身份,旨在通过分析和模型开发来揭露和纠正网络上的错误信息。
CMU-MisCOV19 dataset was developed by Carnegie Mellon University, focusing on COVID-19-related misinformation analysis. This dataset comprises 4,573 annotated tweets covering diverse factual and misinformation types. Specific keywords and time points were employed during data collection to ensure the relevance and timeliness of the dataset. The data underwent multiple rounds of annotation and classification during its creation, ultimately yielding a detailed coding manual with 17 categories. This dataset is primarily utilized to study the network structures and linguistic patterns of COVID-19 misinformation communities, as well as their memberships in other misinformation sub-communities, with the aim of exposing and correcting online misinformation through analysis and model development.



