供应商动态新闻事件抽取语料集
收藏资源简介:
该语料集主要给出关于核电制造企业供应链节点供应商企业的经营状态新闻事件,为事件抽取方法研究提供数据支持,总共有1250个句子。每个句子都是单独进行标注的。文件标注的格式是bio格式,其将每个元素标注为“B-X”、“I-X”或者“O”。其中,“B-X”表示此元素所在的片段属于X类型并且此元素在此片段的开头,“I-X”表示此元素所在的片段属于X类型并且此元素在此片段的中间位置,“O”表示不属于任何类型。比如,我们将 X 表示为名词短语(Noun Phrase, NP),则BIO的三个标记为:(1)B-NP:名词短语的开头(2)I-NP:名词短语的中间(3)O:不是名词短语。数据文件为txt文件、png图片,总计3个数据文件,数据量共1.24MB。
This corpus mainly provides news events concerning the operational status of supplier enterprises at the supply chain nodes of nuclear power manufacturing enterprises, to support research on event extraction methods. There are 1,250 sentences in total, each individually annotated. The dataset uses the BIO annotation format, where each element is labeled as "B-X", "I-X", or "O". Specifically, "B-X" indicates that the segment containing this element belongs to type X and this element is at the start of the segment; "I-X" indicates that the segment belongs to type X and this element is in the middle of the segment; "O" indicates that the element does not belong to any type. For example, if we take X as Noun Phrase (NP), the three BIO tags are: (1) B-NP: the start of a noun phrase; (2) I-NP: the middle of a noun phrase; (3) O: not a noun phrase. The dataset includes 3 files in total, namely TXT documents and PNG images, with an overall data size of 1.24 MB.




