nlu-question_answering
收藏资源简介:
SEA Question Answering数据集用于评估模型在给定段落中回答问题的能力。它包含印度尼西亚语、泰米尔语、泰语和越南语的样本,每个语言部分都有100个示例,并且有少样本示例的额外分割。数据集的特征包括ID、标签、提示(包括问题和文本)、提示模板和元数据(包括语言信息)。数据集的统计信息包括每个分割的示例数量、GPT-4o、Gemma 2和Llama 3的标记数量。数据集的来源包括TyDi QA-GoldP、IndicQA和XQuaD,每个来源都有其特定的许可证。
The SEA Question Answering dataset is designed to evaluate a model's ability to answer questions using provided paragraphs. It contains samples in Indonesian, Tamil, Thai, and Vietnamese, with 100 examples per language subset and an additional few-shot split. The dataset's features include ID, label, prompt (comprising the question and supporting text), prompt template, and metadata (including language information). Its statistical metrics cover the number of examples per split, as well as the token counts generated by GPT-4o, Gemma 2, and Llama 3. The dataset is sourced from TyDi QA-GoldP, IndicQA, and XQuaD, each with its own specific license.




