遇见数据集

DEplain-APA

收藏
Zenodo2024-10-04 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

<strong>DEplain: A corpus for German Text Simplification</strong>This repository contains the corpus called DEplain-APA for German text simplification (document and sentence simplification). The corpus contains Austrian nexts text provided by the APA - Austria Presse Agentur eG. All of the sentence-wise aligned pairs (complex-simple) are manually aligned. The following table summarizes the most important meta data of the corpus. <strong>meta data</strong> <strong>value</strong> language DE-AT (Austrian German) domain news source language level B1 target language level A2 # document pairs (total, train/dev/test) 483 (387/48/48) # sentence pairs (total, train/dev/test) 13,122 (10,660/1,231/1,231) # complex sentences 25,607 # simple sentences 26,471 <strong>Updates:</strong> Version 1.1: Alignment Labels in Simplification Plans are repaired. For more info see https://github.com/rstodden/DEPlain/issues/2#issue-1875006089 For more information, please have a look at our paper. If you use this corpus, please also cite our paper and name APA - Austria Presse Agentur eG as data provider: Regina Stodden, Omar Momen, and Laura Kallmeyer. 2023. DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification. In <em>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, pages 16441–16463, Toronto, Canada. Association for Computational Linguistics.

提供机构:
Zenodo
创建时间:
2023-08-31
二维码
社区交流群
二维码
科研交流群
商业服务