遇见数据集

Citation Reason Dataset

收藏
DataCite Commons2020-08-27 更新2024-07-27 收录
官方服务:

资源简介:

<b>A dataset of ~4K statements from English Wikipedia annotated with the reason why they need a citation.</b><b><br></b>Each line of this tab-separated file contains:* <i>entity_id: </i> the Wikidata ID corresponding to the page* <i>revision_id</i>: the revision of the corresponding Wikipedia article <i>* t</i><i>imestamp: </i>the timestamp of the revision* <i>entity_title: </i>the page/Wikidata ID title <i>* section_id: </i>the section ID where the statement is <i>* section: </i>the section title * <i>prg_idx: </i>the index of the paragraph in the page* <i>sentence_idx: </i> the index of the statement in the paragraph* <i>statement: </i>the statement text<i>* citations</i>: the source cited in the statement:<i>* vote1: </i>first Mechanical Turk judgment<i>* vote2</i>: second Mechanical Turk judgment<i>* vote3: </i>third Mechanical Turk judgment<br>The numbers in the last 3 fields correspond to the following citation reasons:1='direct quotation'2='statistics'3='controversial'4='opinion'5='life'6='scientific'7='historical'8='other'

本数据集包含来自英文维基百科的约4000条语句,每条语句均标注了其需要引用的原因。 该制表符分隔文件的每一行均包含以下字段: * 实体ID(entity_id):对应维基百科词条的维基数据标识符(Wikidata ID) * 修订版本ID(revision_id):对应维基百科词条的修订版本号 * 时间戳(timestamp):该修订版本的时间戳 * 实体标题(entity_title):该词条页面/维基数据标识符对应的标题 * 段落ID(section_id):该语句所在的段落ID * 段落标题(section):该段落的标题 * 段落索引(prg_idx):该词条页面内的段落序号 * 语句索引(sentence_idx):该段落内的语句序号 * 语句内容(statement):该语句的文本 * 引用来源(citations):该语句所引用的来源 * 投票1(vote1):第一条亚马逊机械土耳其人(Mechanical Turk)标注结果 * 投票2(vote2):第二条亚马逊机械土耳其人(Mechanical Turk)标注结果 * 投票3(vote3):第三条亚马逊机械土耳其人(Mechanical Turk)标注结果 最后三个字段的数字对应以下引用原因:1='直接引语',2='统计数据',3='有争议内容',4='个人观点',5='生活相关内容',6='科学内容',7='历史内容',8='其他'

提供机构:
figshare
创建时间:
2019-02-22
搜集汇总
数据集介绍
Citation Reason Dataset 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务