中文突发事件语料库(Chinese Emergency Corpus)
收藏资源简介:
中文突发事件语料库是由上海大学(语义智能实验室)所构建,根据国务院颁布的《国家突发公共事件总体应急预案》的分类体系,从互联网上收集了5类(地震、火灾、交通事故、恐怖袭击和食物中毒)突发事件的新闻报道作为生语料,然后再对生语料进行文本预处理、文本分析、事件标注以及一致性检查等处理,最后将标注结果保存到语料库中,CEC合计332篇。CEC采用了XML语言作为标注格式,其中包含了六个最重要的数据结构(标记):Event、Denoter、Time、Location、Participant和Object。Event用于描述事件;Denoter、Time、Location、Participant和Object用于描述事件的指示词和要素。此外,我们还为每一个标记定义了与之相关的属性。与ACE和TimeBank语料库相比,CEC语料库的规模虽然偏小,但是对事件和事件要素的标注却最为全面。
The Chinese Emergency Corpus (CEC) was constructed by the Semantic Intelligence Laboratory at Shanghai University. It collects news reports on five types of emergencies (earthquakes, fires, traffic accidents, terrorist attacks, and food poisoning) from the internet, based on the classification system of the 'National Emergency Response Plan for Public Emergencies' issued by the State Council. The raw corpus undergoes text preprocessing, text analysis, event annotation, and consistency checking before the annotated results are saved into the corpus, totaling 332 articles. The CEC uses XML as the annotation format, which includes six key data structures (tags): Event, Denoter, Time, Location, Participant, and Object. The Event tag describes the event, while Denoter, Time, Location, Participant, and Object describe the indicators and elements of the event. Additionally, we have defined attributes related to each tag. Compared to the ACE and TimeBank corpora, the CEC corpus, although smaller in scale, provides the most comprehensive annotation of events and their elements.
中文突发事件语料库(CEC)概述
数据集构建
- 构建机构:上海大学语义智能实验室
- 数据来源:互联网上的新闻报道
- 事件分类:地震、火灾、交通事故、恐怖袭击、食物中毒,共5类
- 文本数量:332篇
数据处理
- 预处理步骤:文本预处理、文本分析、事件标注、一致性检查
- 标注格式:XML语言
- 主要数据结构:Event、Denoter、Time、Location、Participant、Object
- 属性定义:每个标记都有相关属性
研究支持
- 资助项目:国家自然科学基金项目“基于描述逻辑的事件推理关键问题研究(编号:61305053)”和“事件本体模型与应用技术”(编号:60975033)
学术成果
- 研究论文:包括事件本体的文本事件要素抽取方法、事件因果关系抽取等
- 学位论文:涉及面向事件的知识处理、事件本体构建等关键问题的研究
数据集特点
- 规模对比:与ACE和TimeBank语料库相比,CEC规模较小
- 标注全面性:对事件和事件要素的标注最为全面




