遇见数据集

EPIC-EuroParl-UdS: A GPT-2 and MT Surprisal-Annotated Corpus for Translation and Interpreting

收藏
Zenodo2026-03-13 更新2026-05-26 收录
官方服务:

资源简介:

General description EPIC-EuroParl-UdS is a bidirectional document- and sentence-aligned English–German corpus of European Parliament debates (up to mid-July 2018). It includes the official written versions of speeches and their translations, as well as manual transcriptions produced from audio recordings of the original spoken speeches and their simultaneous interpreting. Spoken subcorpora (=epic) contain filler particles ('euh', 'hum', 'hm') in up to 44 % of segments and sentence alignment artefacts, reflecting the specificity of the spoken mediation mode. Namely, up to 8% of aligned segments have empty strings on either side to signal missing content identified during manual alignment. A version of the data without filler particles and empty segments is marked as epic_clean. Multisentence segments (one-to-many or many-to-one) make up to 20% of all segments in some spoken subcorpora; in written data, this type of sentence alignment affects around 8% of all segments in any given subcorpus. The written data are over 40 times larger than the spoken data and are imbalanced across translation directions. To ensure comparability and reliable evaluation, the written data were split into training and test sets, with the test set matched to the spoken data in document size. Document integrity and parallelism were preserved throughout corpus construction. Due to the size, the training data are stored by language pair. This combined resource supports the study of translation and interpreting using information-theoretic indices derived from monolingual GPT-2 language models and direction-specific MT models. The corpus is further enriched with CoNLL-U linguistic annotation, document-, segment-, and word-level alignment, enabling analysis across multiple granularities. For all items where the value is not available (empty strings manually aligned to source or target content, absent translator names) or where surprisal can not be computed (filler particles, multiword expansions, errors) were set to N/A (represented as pandas NA). Corpus structure and data formats The corpus is organised by mode, version, and data format. Modes (subcorpora) Spoken (epic) Written (europ) Versions for each mode Spoken: epic-clean, epic-fp-nones Written: europ-train (by translation direction due to size), europ-test Data formats vrt: word-level representation with CoNLL-U fields, per-token GPT-2 and MT surprisal values, and word alignment long: segment-level representation (independent of text type: source or target) with average GPT-2 surprisal wide: segment-pair-level representation with average MT alignment scores and segment-level BLEU Data representation and identifiers All formats are provided as compressed tab-separated tables (.tsv.gz). A consistent system of unique identifiers (doc_id, seg_id, and word_id) follows a rigid naming convention, allowing unambiguous mapping across formats. Item identifiers follow the pattern: <ttype>_<mode>_<src_lang>_<tgt_lang>_<doc_id>-<seg_id>:<word_id> For example, ORG_SP_DE_EN_131-02:001 refers to the first word of the second segment in the original German document 131, interpreted into English in the spoken subcorpus. Annotation and metadata Depending on the format, tables include categorical variables such as doc_id, seg_id, word_id, mode, ttype, lpair, lang, and speaker_id. These facilitate data slicing and can be used as fixed effects in statistical modelling. Document-level metadata are provided in: epic_org_doc_meta.tsv.gz (spoken): date, nationality, title, topic europ_org_doc_meta.tsv.gz (written): date, nationality, birth_date, birth_place, n_party, p_group Descriptive statistics and mean index values by language pair and mode are distributed in stats.zip (e.g. wide_basics.tsv, wide_indices.tsv). Further documentation and references A slide deck (EPIC–EuroParl–UdS_2026_surprisal_indexing_slides.pdf) presents corpus statistics and key methodological aspects of data preprocessing and annotation, including evaluation of various surprisal indexing options (aggregating LLM subword surprisals, fine-tuned models, sliding window approach). If you use EPIC–EuroParl–UdS in your research, please cite the corresponding LREC-2026 publication. @inproceedings{kunilovskaya2026perspectives, author = {Maria Kunilovskaya and Christina Pollkläsener}, title = {{EPIC–EuroParl–UdS: Information-Theoretic Perspectives on Translation and Interpreting}}, booktitle = {Proceedings of the 15th Conference on Language Resources and Evaluation (LREC 2026)}, publisher = {The European Language Resources Association (ELRA)}, editor = {}, year = {2026}, location = {Palma de Mallorca, Spain}, month = {11--16 May}, subject = {Computational Linguistics}, pages = {in print} }

提供机构:
Zenodo
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务