遇见数据集

Computer Science Contextual Question Answering [CSCQA]

收藏
Mendeley Data2026-05-21 收录
官方服务:

资源简介:

The dataset is a context-based question answering (QA) resource for Computer Science, where each instance is structured into six components: Topic, Context, Sentence, Keyword, Question, and Answer. Topic - defines the domain or concept. Context - provides the source paragraph. Sentence - represents a processed or simplified form of the context. Keyword - highlights the key term used to guide question generation. Question - is automatically or manually generated from the context. Answer - contains the corresponding response derived from the context. This structured format supports tasks such as automatic question generation, answer extraction, and contextual understanding in NLP systems.

本数据集为面向计算机科学领域的基于上下文的问答(QA, Question Answering)资源,每个样本均包含六大组成部分:主题(Topic)、上下文(Context)、语句(Sentence)、关键词(Keyword)、问题(Question)与答案(Answer)。 主题(Topic):用于界定所属领域或核心概念。 上下文(Context):提供源段落文本。 语句(Sentence):为上下文经过处理或简化后的形式。 关键词(Keyword):用于指引问题生成的核心高亮术语。 问题(Question):由上下文自动或手动生成的查询内容。 答案(Answer):包含从上下文中提取得到的对应响应内容。 该结构化格式可支持自然语言处理(NLP, Natural Language Processing)系统中的自动问题生成、答案抽取以及上下文理解等多项任务。

创建时间:
2026-04-28
二维码
社区交流群
二维码
科研交流群
商业服务