遇见数据集

Datasets from Costa Rican news sources for fake news detection

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Today, technology has changed the way information is propagated and how the message is received. The interpretation of the news may have different angles depending on the source of origin. Because of this, there has been an increase in misinformation, in the way of influencing public opinion and in how we perceive or estimate reality.<br> The objective of this beta dataset is to be used for the evaluation of data mining models that allow the classification of true or potentially fake news that are generated by Costa Rican news sites only. This is intended to assess the level of reliability of the models and extend the scope of this research in future work. The dataset has been pre-processed (standarized using lower cases, lemmatized and removed any possible noise from it) and analyzed using LIWC dictionaries. One version has the news text in Spanish and was processed using LIWC2007 dictionary in Spanish. The second version was processed using LIWC2015 English dictionary and has the news text in English. The reason to having two versions is to be able to test using the newer LIWC dictionary which includes more Summary Language Variables that the Spanish version doesn't have and analyze how this and other variables can contribute to different results when creating models. The file "DescripcionVariables" provides a description of all variables used.

提供机构:
Zenodo
创建时间:
2019-10-26
二维码
社区交流群
二维码
科研交流群
商业服务