遇见数据集

The JRC-Acquis Corpus, version 3.0

收藏
SSH Open MarketPlace2024-09-30 更新2024-10-05 收录
官方服务:

资源简介:

This is a parallel corpus of Acquis Communautaire, which is the total body of European Union law applicable in European member states. Most texts have been manually classified according to the EUROVOC subject domains so that the collection can also be used to train and test multi-label classification algorithms and keyword-assignment software. The corpus is encoded in XML, according to the Text Encoding Initiative Guidelines. Due to the large number of parallel texts in many languages, the JRC-Acquis is particularly suitable to carry out all types of cross-language research, as well as to test and benchmark text analysis software across different languages (for instance for alignment, sentence splitting and term extraction). The sentence-level alignment was done using the [hunalign](https://github.com/danielvarga/hunalign) tool. The corpus is available for download from the CLARIN:EL repository.

本数据集为欧盟既有法律体系(Acquis Communautaire)平行语料库,该语料库涵盖适用于欧盟成员国的全部欧盟法律文本。绝大多数文本已依据欧盟叙词表(EUROVOC)的主题领域完成人工分类,因此该语料库可用于训练和测试多标签分类算法与关键词分配软件。本语料库依据文本编码倡议(Text Encoding Initiative)指南采用XML格式编码。由于该语料库包含多语言平行文本且规模庞大,JRC-Acquis语料库尤其适用于各类跨语言研究,同时可用于跨语言文本分析软件的测试与基准评测(例如文本对齐、分句及术语抽取等任务)。句级对齐工作借助hunalign工具完成,该工具的开源仓库地址为https://github.com/danielvarga/hunalign。该语料库可从CLARIN:EL仓库下载获取。

创建时间:
2024-09-30
搜集汇总
数据集介绍
The JRC-Acquis Corpus, version 3.0 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务