遇见数据集

TAC KBP Comprehensive English Source Corpora 2009-2014

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>TAC KBP Comprehensive English Source Corpora 2009-2014 was developed by the Linguistic Data Consortium (LDC) and contains the 3,877,207 English source documents used in support of the TAC KBP tasks from 2009-2014.</p><br> <p>Text Analysis Conference (<a href="https://tac.nist.gov/">TAC</a>) is a series of workshops organized by the National Institute of Standards and Technology (<a href="https://www.nist.gov/">NIST</a>). TAC was developed to encourage research in natural language processing and related applications by providing a large test collection, common evaluation procedures, and a forum for researchers to share their results. Through its various evaluations, the Knowledge Base Population (KBP) track of TAC encourages the development of systems that can match entities mentioned in natural texts with those appearing in a knowledge base and extract novel information about entities from a document collection and add it to a new or existing knowledge base.</p><br> <h3>Data</h3><br> <p>The source data consists of newswire, broadcast material, and web text collected by LDC. Documents are released as a collection of zip files for overall compactness, and ease and efficiency of use. When unpacked the documents are all UTF-8 text files with a basic markup structure. Also provided are a series of lists and tables to aid in specific zip file to doc mappings and the recreation of specific test sets. See the included documentation for more information.</p><br> <h3>Acknowledgement</h3><br> <p>This material is based on research sponsored by Air Force Research Laboratory and Defense Advance Research Projects Agency under agreement number FA8750-13-2-0045. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of Air Force Research Laboratory and Defense Advanced Research Projects Agency or the U.S. Government.</p><br> <h3>Samples</h3><br> <p>Please view this <a href="desc/addenda/LDC2018T03.txt">sample</a>.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 1994-1997, 2001-2010 Agence France Presse, © 2005 Aljazeera, © 1996-1997 American Broadcasting Company, © 1994-2010 The Associated Press, © 1994-1997 Cable News Network, LP, LLLP, © 1997-1999, 2001, 2003-2010 Central News Agency (Taiwan), © 2005 Dubai TV, © 2005-2006 National Broadcasting Company, Inc., © 1996-1997 National Cable Satellite Corporation, © 1994-1998, 2003-2009 Los Angeles Times - Washington Post News Service, Inc., © 1994-2010 New York Times, © 1996-1997 Public Radio Iternational,© 1994-1995 Reuters America, Inc., © 1996 The University of California, USC Radio and Marketplace, © 2010 The Washington Post Service with Bloomberg News, © 1995-2010 Xinhua News Agency, © 1996-1998, 2007, 2011, 2018 Trustees of the University of Pennsylvania

<h3>引言</h3> <p>2009-2014年TAC KBP综合英语源语料库由语言数据联盟(Linguistic Data Consortium, LDC)开发,收录了2009至2014年间支撑TAC KBP任务的3,877,207份英语源文档。</p> <p>文本分析会议(Text Analysis Conference, TAC)是由美国国家标准与技术研究院(National Institute of Standards and Technology, NIST)主办的系列研讨会。TAC旨在通过提供大规模测试集、统一评估流程以及供研究者分享研究成果的交流平台,推动自然语言处理及其相关应用领域的研究。通过其各类评估活动,TAC下设的知识基础人口普查(Knowledge Base Population, KBP)赛道致力于推动相关系统的研发,此类系统可将自然文本中提及的实体与知识库中的实体进行匹配,并从文档集合中提取实体的新信息,将其添加至新建或已有知识库中。</p> <h3>数据集</h3> <p>本数据集的源数据由LDC采集的新闻专线稿件、广播素材与网络文本组成。为实现整体存储紧凑性并提升使用的便捷性与效率,所有文档以压缩包集合的形式发布。解压后,所有文档均为带有基础标记结构的UTF-8编码文本文件。此外,数据集还附带一系列列表与表格,可辅助完成特定压缩包与文档的映射工作,以及复现特定测试集。更多详情请参阅随附的文档说明。</p> <h3>致谢</h3> <p>本材料基于美国空军研究实验室(Air Force Research Laboratory)与美国国防高级研究计划局(Defense Advanced Research Projects Agency)根据编号FA8750-13-2-0045的协议资助的研究成果。尽管本材料带有任何版权标注,美国政府仍有权出于政府用途复制和分发其重印本。本文所载观点与结论仅代表作者本人,不应被视为必然反映美国空军研究实验室、美国国防高级研究计划局或美国政府的官方政策或认可(无论明示或默示)。</p> <h3>示例</h3> <p>请查看此<a href="desc/addenda/LDC2018T03.txt">示例文件</a>。</p> <h3>更新情况</h3> <p>暂无更新。</p> <p>本数据集部分内容 © 1994-1997、2001-2010 法新社(Agence France Presse),© 2005 半岛电视台(Aljazeera),© 1996-1997 美国广播公司(American Broadcasting Company),© 1994-2010 美联社(The Associated Press),© 1994-1997 美国有线电视新闻网有限责任合伙(Cable News Network, LP, LLLP),© 1997-1999、2001、2003-2010 中央通讯社(中国台湾)(Central News Agency (Taiwan)),© 2005 迪拜电视台(Dubai TV),© 2005-2006 美国国家广播公司(National Broadcasting Company, Inc.),© 1996-1997 国家有线卫星公司(National Cable Satellite Corporation),© 1994-1998、2003-2009 洛杉矶时报-华盛顿邮报新闻服务公司(Los Angeles Times - Washington Post News Service, Inc.),© 1994-2010 《纽约时报》(New York Times),© 1996-1997 国际公共广播(Public Radio International,原文笔误为Public Radio Iternational),© 1994-1995 路透美国公司(Reuters America, Inc.),© 1996 加利福尼亚大学、南加州大学广播与市场频道(The University of California, USC Radio and Marketplace),© 2010 华盛顿邮报通讯社与彭博新闻合办机构(The Washington Post Service with Bloomberg News),© 1995-2010 新华通讯社(Xinhua News Agency),© 1996-1998、2007、2011、2018 宾夕法尼亚大学托管委员会(Trustees of the University of Pennsylvania)</p>

创建时间:
2020-11-30
搜集汇总
数据集介绍
TAC KBP Comprehensive English Source Corpora 2009-2014 数据集图片
背景与挑战
背景概述
该数据集是一个大规模的英语语料库,包含约387.7万份文档,来源于新闻专线、广播、论坛、新闻组和博客等多种渠道,覆盖2009年至2014年。它专门用于支持TAC KBP(知识库填充)任务,旨在促进信息提取和知识表示的研究,帮助系统从自然文本中匹配实体并丰富知识库。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务