遇见数据集

Crowdsourced Temporal Data

收藏
DataONE2021-12-21 更新2024-06-08 收录
官方服务:

资源简介:

Data attained through crowdsourcing have an essential role in the development of computer vision algorithms. Crowdsourced data might include reporting biases, since crowdworkers usually describe what is “worth saying\" in addition to images’ content. We explore how the unprecedented events of 2020, including the unrest surrounding racial discrimination, and the COVID-19 pandemic, might be refected in responses to an open-ended annotation task on people images, originally executed in 2018 and replicated in 2020. Analyzing themes of Identity and Health conveyed in workers’ tags, we found evidence that supports the potential for temporal sensitivity in crowdsourced data. The 2020 data exhibit more race-marking of images depicting non-Whites, as well as an increase in tags describing Weight. We relate our findings to the emerging research on crowdworkers’ moods. This dataset includes all the tags, provided by crowdworkers, relevant to the topics of Health and Identity, providing aggregated counts of the occurrences of each tag in 2018 and 2020. Additionally, separate counts of the occurrences of each tag in 2018 and 2020 are provided for each depicted race (a.k.a., White, Latino, Black and Asian).

众包获取的数据在计算机视觉(computer vision)算法的开发中发挥着至关重要的作用。众包数据往往存在报告偏差,这是因为众包标注人员除了描述图像本身的内容外,还会补充标注他们认为“值得提及”的信息。我们探究了2020年的一系列重大事件——包括围绕种族歧视的社会动荡以及新型冠状病毒肺炎(COVID-19)大流行——如何体现在针对人物图像的开放式标注任务的应答结果中。该任务最初于2018年启动,并于2020年完成重复实验。通过分析标注人员所提供标签中蕴含的身份(Identity)与健康(Health)主题,我们找到了能够证明众包数据具有时间敏感性的证据。2020年的数据集针对非白人(non-White)群体的图像标注了更多种族相关信息,同时描述体重(Weight)的标签数量也有所增加。我们将本次研究发现与当前关于众包标注人员情绪的新兴研究相结合。本数据集包含了众包标注人员提供的所有与身份和健康主题相关的标签,并给出了2018年与2020年各标签出现次数的汇总统计值。此外,本数据集还针对每一类被标注的种族(即白人、拉丁裔、黑人和亚裔),分别给出了2018年与2020年各标签的出现次数统计。

创建时间:
2023-11-12
二维码
社区交流群
二维码
科研交流群
商业服务