遇见数据集

Portuguese Newswire Text

收藏
DataCite Commons2021-07-01 更新2024-07-13 收录
官方服务:

资源简介:

<h3>Introduction</h3> <p>This corpus builds on the Portuguese data published previously in the <a href="http://catalog.ldc.upenn.edu/LDC95T11" rel="nofollow">European Language Newswire Text Corpus</a> and contains the previously published material, as well as more recent material. </p><h3>Data</h3> <p>The data in this corpus comes from Agence France Presse from May 13, 1994 through December 31, 1998 (June 27, 1996 - December 31, 1998 was previously unpublished by the LDC). The data has been tagged using SGML to identify article boundaries. </p><h3>Updates</h3> There are no updates at this time. </br>

<h3>引言</h3> <p>本语料库基于此前发布于<a href="http://catalog.ldc.upenn.edu/LDC95T11" rel="nofollow">欧洲语言新闻专线文本语料库(European Language Newswire Text Corpus)</a>的葡萄牙语数据构建,既涵盖已公开的历史资料,也补充了最新的内容。</p><h3>数据</h3> <p>本语料库的数据源自法新社(Agence France Presse)1994年5月13日至1998年12月31日的新闻稿件,其中1996年6月27日至1998年12月31日的内容此前未由语言数据联盟(Linguistic Data Consortium,LDC)发布。该数据集已使用标准通用标记语言(Standard Generalized Markup Language,SGML)进行标注,以识别单篇稿件的分界。</p><h3>更新情况</h3> 目前暂无更新计划。</br>

创建时间:
2020-11-30
搜集汇总
数据集介绍
Portuguese Newswire Text 数据集图片
背景与挑战
背景概述
该数据集是葡萄牙语新闻电讯文本集合,包含来自Agence France Presse的1994年5月至1998年12月的数据,其中部分内容为首次发布。它基于早期欧洲语言新闻语料库扩展而来,采用SGML标记,适用于信息检索和语言建模任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务