ETR-fr
收藏资源简介:
ETR-fr is the first high-quality French dataset specifically designed for Easy-to-Read (ETR) text generation. The corpus consists of 523 paragraph-aligned pairs of original and Easy-to-Read texts, created according to the official European Easy-to-Read guidelines. It was developed to support research on automatic text adaptation for people with cognitive disabilities and intellectual impairments. Unlike traditional text simplification datasets, ETR-fr focuses on paragraph-level rewriting and combines lexical, syntactic, and structural transformations to produce texts that are easier to understand while preserving the essential meaning of the original content. The dataset was constructed from the François Baudez Publishing collection, whose books are specifically written for readers with cognitive impairments. Original and Easy-to-Read paragraphs were manually aligned to ensure high-quality supervision for machine learning applications. Dataset characteristics Language: French Task: Easy-to-Read (ETR) text generation Alignment: Paragraph-level Number of aligned pairs: 523 Average source length: 102.8 words Average target length: 46.2 words Average compression ratio: 50.1% Average lexical novelty: 53.8% The dataset is distributed with predefined training, validation, and test splits to ensure reproducible evaluation. In addition, the associated benchmark introduces ETR-fr-politic, an out-of-domain evaluation set composed of 33 paragraph pairs derived from French political programs written in Easy-to-Read format, enabling the assessment of model generalization. Intended uses ETR-fr is intended for research on: Easy-to-Read text generation Text simplification Accessible Natural Language Processing (NLP) Cognitive accessibility Large Language Model (LLM) fine-tuning Retrieval-Augmented Generation (RAG) Multi-task learning Evaluation of automatic text adaptation systems License and citation If you use this dataset in your research, please cite the corresponding publication: Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation (Ledoyen et al., 2025).



