遇见数据集

digitalfabrik/integreat-qa

收藏
Hugging Face2024-09-27 更新2025-04-26 收录
官方服务:

资源简介:

--- license: cc-by-4.0 task_categories: - question-answering task_ids: - extractive-qa language: - de - en tags: - migration - refugees - extractive-qa size_categories: - n<1K annotations_creators: - crowdsourced source_datasets: - original pretty_name: Integreat QA --- # Dataset Our dataset consists of 906 diverse QA pairs in German and English. The dataset is extractive, i.e., answers are given as sentence indices (breaking at the newline character `\n`). Questions are automatically generated using an LLM. The answers are manually annotated using voluntary crowdsourcing. **Repository:** [More Information Needed](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) **Paper:** - [https://arxiv.org/abs/1806.03822](https://arxiv.org/abs/1806.03822) - [https://aclanthology.org/2024.konvens-main.25/](https://aclanthology.org/2024.konvens-main.25/) Our dataset is licensed under [cc-by-4.0](https://choosealicense.com/licenses/cc-by-4.0). ## Properties A QA pair consists of - `question` (string): Question - `context` (string): Full text from the Integreat-App - `answers` (number[]): Indices of answer sentences Furthermore, the following properties are present: - `id` (number): A unique id for the QA pair - `language` (string): The language of question and context. - `sourceLanguage` (string | null): If question and context are machine translated, the source language. - `city` (string): The city the page in the Integreat-App belongs to. - `pageId` (number): The page id of the page in the Integreat-App. - `jaccard` (number): The sentence-level inter-annotator agreement of manual answer annotation.

许可证:cc-by-4.0 任务类别: - 问答 任务子类别: - 抽取式问答(Extractive QA) 语言: - 德语 - 英语 标签: - 移民 - 难民 - 抽取式问答(Extractive QA) 规模类别: - 样本量小于1000 注释生成者: - 众包 源数据集: - 原创 展示名称:Integreat QA # 数据集 本数据集包含906条覆盖多元领域的德英双语问答对。 本数据集属于抽取式问答(Extractive QA)数据集,即答案以句子索引的形式给出(以换行符` `作为句子拆分边界)。 问题通过大语言模型(LLM)自动生成,答案则通过自愿众包方式完成人工标注。 **仓库:** [更多信息待补充](https://github.com/huggingface/datasets/blob/master/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards) **论文:** - [https://arxiv.org/abs/1806.03822](https://arxiv.org/abs/1806.03822) - [https://aclanthology.org/2024.konvens-main.25/](https://aclanthology.org/2024.konvens-main.25/) 本数据集采用[CC-BY-4.0(知识共享署名4.0许可协议)](https://choosealicense.com/licenses/cc-by-4.0)进行授权。 ## 数据集属性 每条问答样本包含以下字段: - `question`(字符串类型):问题文本 - `context`(字符串类型):来自Integreat应用的完整文本 - `answers`(数字数组):答案句子的索引值 此外,样本还包含以下附加字段: - `id`(数字类型):问答样本的唯一标识符 - `language`(字符串类型):问题与上下文的语言 - `sourceLanguage`(字符串类型或空值):若问题与上下文为机器翻译所得,则标注其源语言 - `city`(字符串类型):Integreat应用中对应页面所属的城市 - `pageId`(数字类型):Integreat应用中对应页面的页面ID - `jaccard`(数字类型):人工答案标注的句子级注释者间一致性(雅卡尔系数)

提供机构:
digitalfabrik
二维码
社区交流群
二维码
科研交流群
商业服务