Propicto/propicto-polylexical
收藏资源简介:
--- language: - fr license: mit task_categories: - translation tags: - pictograms - AAC pretty_name: Propicto-polylexical --- # Propicto-polylexical ## 📝 Dataset Description Propicto-polylexical is a dataset of aligned text and pictograms (the pictograms correspond to the identifiers associated with ARASAAC pictograms) in French. This dataset was manually created to specifically provide a resource containing texts with polylexical expressions translated into pictograms. The dataset contains is a single file of 1,462 utterances. ## ⚒️ Dataset Structure The dataset is structured as follows: ```csv id : the unique identifier of the utterance text : the sentence in French pictos : the sequence of pictogram IDs from ARASAAC tokens : the sequence of tokens, each of which is a keyword associated with an ARASAAC pictogram ID ``` ## 💡 Dataset example For the given sample : ```csv id : 43 text : le collier du chien est trop serré il faut l'ajuster pictos : [8476, 6987, 36480, 25708, 5380, 15523, 8476, 8516] tokens : le collier_du_chien être trop serrer devoir le adapter ``` - `pictos` is the sequence of pictogram IDs, each of them can be retrieved from here : 15523 = https://static.arasaac.org/pictograms/15523/15523_2500.png<br /> - `tokens` are retrieved from a specific lexicon and can be used to train translation models.  ## 💻 Uses Propicto-polylexical is intended to be used to train Text-to-Pictograms translation models. This dataset can also be used to fine-tune large language models to perform translation into pictograms. ## ⚙️ Dataset Creation The dataset is created by applying a specific formalism that converts french transcriptions into a corresponding sequence of pictograms.<br /> The formalism includes a set of grammatical rules to handle specific phenomenon (negation, name entities, pronominal form, plural, ...) to the French language, as well as a dictionary which associates each ARASAAC ID pictogram with a set of keywords (tokens).<br /> This formalism was presented at [LREC](https://aclanthology.org/2024.lrec-main.76/). ## ⁉️ Limitations The translation can be partially incorrect, due to incorrect or missing words translated into pictograms. ## 💡 Information - **Curated by:** Cécile MACAIRE - **Funded by :** [PROPICTO ANR-20-CE93-0005](https://anr.fr/Projet-ANR-20-CE93-0005) - **Language(s) (NLP):** French - **License:** CC-BY-NC-SA-4.0 ## 📌 Citation ```bibtex @inproceedings{macaire-etal-2024-multimodal, title = "A Multimodal {F}rench Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation", author = "Macaire, C{\'e}cile and Dion, Chlo{\'e} and Arrigo, Jordan and Lemaire, Claire and Esperan{\c{c}}a-Rodier, Emmanuelle and Lecouteux, Benjamin and Schwab, Didier", booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)", year = "2024", publisher = "ELRA and ICCL", url = "https://aclanthology.org/2024.lrec-main.76", pages = "839--849", } ``` ## 👩🏫 Dataset Card Authors **Cécile MACAIRE, Chloé DION, Emmanuelle ESPÉRANÇA-RODIER, Benjamin LECOUTEUX, Didier SCHWAB**
--- 语言: - 法语 许可证: MIT许可证 任务类别: - 机器翻译 标签: - 象形图(pictograms) - 辅助与替代沟通(Augmentative and Alternative Communication, AAC) 展示名称: Propicto-polylexical --- # Propicto-polylexical ## 📝 数据集描述 Propicto-polylexical 是一套法语对齐文本与象形图(pictograms)的数据集,其中象形图对应ARASAAC象形图库的关联标识符。本数据集为手工构建,旨在专门提供包含多词表达式并被转换为象形图的文本资源。 本数据集包含1462条话语的单个数据文件。 ## ⚒️ 数据集结构 数据集结构如下: csv id : 话语的唯一标识符 text : 法语句子 pictos : ARASAAC象形图ID序列 tokens : 词元(Token)序列,每个词元均为与ARASAAC象形图ID相关联的关键词 ## 💡 数据集示例 以下为数据集示例: csv id : 43 text : le collier du chien est trop serré il faut l'ajuster pictos : [8476, 6987, 36480, 25708, 5380, 15523, 8476, 8516] tokens : le collier_du_chien être trop serrer devoir le adapter - `pictos` 为象形图ID序列,每个ID均可通过以下链接获取:15523 = https://static.arasaac.org/pictograms/15523/15523_2500.png - `tokens` 源自特定词典,可用于训练翻译模型。  ## 💻 应用场景 Propicto-polylexical 旨在用于训练文本到象形图翻译模型。本数据集还可用于微调大语言模型(Large Language Model, LLM)以实现文本到象形图的翻译任务。 ## ⚙️ 数据集构建 本数据集通过特定形式化方法构建,该方法可将法语转录内容转换为对应的象形图序列。该形式化方法包含一套针对法语语言的特定语法规则,用于处理否定、命名实体、代词形式、复数等语言现象,同时附带词典,将每个ARASAAC象形图ID与一组关键词(词元)相关联。该形式化方法已在[LREC](https://aclanthology.org/2024.lrec-main.76/)会议上发表。 ## ⁉️ 局限性 由于部分单词的象形图翻译存在错误或缺失,本数据集的翻译结果可能存在部分不准确之处。 ## 💡 相关信息 - **整理者:** 塞西尔·马克尔(Cécile MACAIRE) - **资助方:** [PROPICTO ANR-20-CE93-0005](https://anr.fr/Projet-ANR-20-CE93-0005) - **NLP使用语言:** 法语 - **许可证:** CC-BY-NC-SA-4.0 ## 📌 引用 bibtex @inproceedings{macaire-etal-2024-multimodal, title = "A Multimodal {F}rench Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation", author = "Macaire, C{é}cile and Dion, Chlo{é} and Arrigo, Jordan and Lemaire, Claire and Esperan{ç}a-Rodier, Emmanuelle and Lecouteux, Benjamin and Schwab, Didier", booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)", year = "2024", publisher = "ELRA and ICCL", url = "https://aclanthology.org/2024.lrec-main.76", pages = "839--849", } ## 👩🏫 数据集卡片作者 **塞西尔·马克尔(Cécile MACAIRE)、克洛埃·迪翁(Chloé DION)、埃马纽埃尔·埃斯佩兰萨-罗迪耶(Emmanuelle ESPÉRANÇA-RODIER)、本杰明·勒库特(Benjamin LECOUTEUX)、迪迪埃·施瓦布(Didier SCHWAB)**



