il_calderone
收藏资源简介:
该数据集是[The Cauldron](https://huggingface.co/datasets/HuggingFaceM4/the_cauldron)的机器翻译版本,专门为意大利语设计。原始数据集包含50个任务,但只有15个任务在机器翻译后仍保持其意义,因此被保留。在这15个任务中,选择了前10,000行进行机器翻译,未正确翻译的问答对被丢弃。图像路径的格式化策略如下:{task-name}/images/{row_number}_{image_number},其中{task-name}是原始数据集中的任务名称,{row_number}是原始数据集中的行号,{image_number}是图像的索引(在有多个图像作为输入的任务中)。
This dataset is a machine-translated version of [The Cauldron](https://huggingface.co/datasets/HuggingFaceM4/the_cauldron), specifically designed for the Italian language. The original dataset contains 50 tasks, but only 15 tasks retained their meaning after machine translation and were thus preserved. For these 15 tasks, the first 10,000 rows were selected for machine translation, and incorrectly translated question-answer pairs were discarded. The formatting strategy for image paths is as follows: {task-name}/images/{row_number}_{image_number}, where {task-name} refers to the task name in the original dataset, {row_number} is the row number in the original dataset, and {image_number} is the index of the image (used in tasks with multiple input images).




