alexandrainst/ragtruth-translated-hallucinations
收藏资源简介:
该数据集是一个多语言文本数据集,包含丹麦语(da)、德语(de)、英语(en)、西班牙语(es)、冰岛语(is)、斯洛文尼亚语(sl)、瑞典语(sv)和乌克兰语(uk)八个语言版本。每个样本由提示文本(prompt)、答案文本(answer)和标签列表(labels)组成,其中标签包括起始位置(start)、结束位置(end)和标签类型(label)。此外,每个样本还包含分割信息(split)、任务类型(task_type)、数据集来源(dataset)和语言代码(language)。数据集仅提供训练集,数据量从988个示例(斯洛文尼亚语)到17790个示例(丹麦语和英语)不等,总大小从约4.3MB(斯洛文尼亚语)到74.3MB(乌克兰语)。该数据集可能用于多语言自然语言处理任务,如文本标注、问答或序列标注。
This dataset is a multilingual text dataset comprising eight language versions: Danish (da), German (de), English (en), Spanish (es), Icelandic (is), Slovenian (sl), Swedish (sv), and Ukrainian (uk). Each sample consists of a prompt text, an answer text, and a list of labels, where labels include start position, end position, and label type. Additionally, each sample contains split information, task type, dataset source, and language code. The dataset only provides training splits, with the number of examples ranging from 988 (Slovenian) to 17,790 (Danish and English), and total sizes from approximately 4.3MB (Slovenian) to 74.3MB (Ukrainian). It is likely designed for multilingual natural language processing tasks such as text annotation, question answering, or sequence labeling.




