A Multidimensional Dataset for Analyzing and Detecting News Bias based on Crowdsourcing
收藏资源简介:
We provide a large data set consisting of <strong>2,057 sentences</strong> from 90 news articles and annotations of crowdworkers with respect to <strong>bias itself</strong> and the following <strong>bias dimensions</strong>: <strong>hidden assumptions</strong> <strong>subjectivity</strong> <strong>representation tendencies</strong> Our data set contains <strong>44,547 labels in total</strong> (43,197 sentence labels and 1,350 article labels). The news articles deal with the <strong>Ukraine crisis</strong>. They were published in 33 countries in total and were selected based on the data set of Cremisini et al. (Cremisini, A., Aguilar, D., & Finlayson, M. A. <em>A Challenging Dataset for Bias Detection: The Case of the Crisis in the Ukraine</em>, Proc. of SBP-BRiMS'19, pp. 173-183, 2019). Each sentence was annotated by 5 crowdworkers. In total, we spent $ 3,335 for the crowdworkers annotations. More information can be found in our GitHub repository. A description of the used file format is given in the codebook attached to the dataset. Please cite our data set as follows: <pre><code>@unpublished{Faerber2020Bias, author = {Michael F{\"{a}}rber and Victoria Burkard and Adam Jatowt and Sora Lim}, title = {{A Multidimensional Dataset for Analyzing and Detecting News Bias based on Crowdsourcing}}, year = {2020} }</code></pre>



