JudSacr/DFP
收藏资源简介:
这是一个法语提示数据集(DFP),包含113,129,978行数据(但由于许可原因,仅共享107,796,041行,其中训练集102,720,891个样本,验证集2,584,400个样本,测试集2,490,750个样本)。它涵盖了30个不同的自然语言处理(NLP)任务。数据集基于724个编写的提示生成,这些提示以命令式、非正式(tutoiement)和正式(vouvoiement)形式表达,旨在尽可能覆盖预训练数据的范围,以适配使用这些数据的模型(模型的具体细节未知)。数据集包含四列:inputs(字符串)、targets(字符串)、dataset(字符串)和task(字符串)。其中,inputs和targets列遵循Muennighoff等人提出的xP3数据集的格式;dataset列允许用户筛选所需的数据集;task列允许用户筛选所需的任务。该数据集源自34个其他数据集(每个都有其自己的许可证),而724个提示则使用cc-by-4.0许可证,因此可以自由应用于用户自己的数据集。数据集是74个提示数据集的串联,用户可以在指定链接中找到这些数据集。命名规则为“原始数据集名称”+“_fr_prompt_”+“任务名称”。
This dataset of prompts in French (DFP) contains 113,129,978 rows but for licensing reasons we can only share 107,796,041 rows (train: 102,720,891 samples, validation: 2,584,400 samples, test: 2,490,750 samples). It presents data for 30 different NLP tasks. 724 prompts were written, including requests in imperative, tutoiement and vouvoiement form in an attempt to have as much coverage as possible of the pre-training data used by the model that will use these data and which are unknown to us. The dataset contains four columns: inputs (string), targets (string), dataset (string), and task (string). The inputs and targets columns follow the same format as the xP3 dataset by Muennighoff et al. The dataset column allows the user to filter the datasets he wants to keep for his work, and the task column allows the user to filter the tasks he wants to keep for his work. The dataset was created from 34 other datasets each with its own license, and the 724 prompts are licensed under the cc-by-4.0 license, so youre free to apply them to your own datasets. The dataset is the concatenation of 74 prompts datasets, and the nomenclature adopted for these datasets is original dataset name + _fr_prompt_ + task name.



