CODA-19
收藏资源简介:
CODA-19是由宾夕法尼亚州立大学创建的一个包含10,966篇英文摘要的人工标注数据集,用于标注COVID-19开放研究数据集中的背景、目的、方法、发现/贡献和其他部分。该数据集由248名亚马逊Mechanical Turk的众包工作者在10天内完成,其标注质量与专家相当。每个摘要由九名不同的工作者标注,最终标签通过多数投票确定。CODA-19的标签在与生物医学专家标签比较时准确率达到82.2%,表明非专家众包可以大规模快速参与COVID-19的研究。该数据集有助于科学家访问和整合快速增长的冠状病毒文献,并作为AI/NLP研究的基石,解决获取专家标注速度慢的问题。
CODA-19 is a manually annotated dataset containing 10,966 English abstracts developed by The Pennsylvania State University. It is intended to annotate standard sections including Background, Objective, Methodology, Finding/Contribution, and Miscellaneous in the COVID-19 Open Research Dataset. This dataset was completed by 248 crowdworkers recruited from Amazon Mechanical Turk over a 10-day period, with its annotation quality comparable to that of domain experts. Each abstract was annotated by nine unique workers, and the final consensus label is determined via majority voting. When compared with annotations provided by biomedical experts, the labels from CODA-19 achieve an accuracy of 82.2%, demonstrating that non-expert crowdsourcing can rapidly engage in COVID-19 research at scale. This dataset enables scientists to access and integrate the rapidly growing body of coronavirus literature, and serves as a foundational resource for AI/NLP research to address the long-standing challenge of slow expert annotation acquisition.

- 1CODA-19: Using a Non-Expert Crowd to Annotate Research Aspects on 10,000+ Abstracts in the COVID-19 Open Research Dataset宾夕法尼亚州立大学 · 2020年



