遇见数据集

Task2Model

收藏
Zenodo2026-02-07 更新2026-05-26 收录
官方服务:

资源简介:

We release EarlyBird and Task2Model, two datasets derived from scientific publications that document how NLP models are mentioned and linked to scientific tasks in the research literature. EarlyBird (EarlyBird Model Mention Dataset) is an intermediate dataset that captures all detected mentions of NLP models within individual scientific papers. Each entry is accompanied by rich metadata, including paper identifiers, section indicators, and model mention counts, enabling fine-grained analyses of model usage, reporting practices, and mention patterns across the scientific literature. Building on this resource, we further release Task2Model, a large-scale dataset of scientific problem descriptions or tasks linked to NLP models.Task2Model emphasizes hierarchical and family-level evaluation, systematic quality audits, and explicit leakage-prevention measures, providing a reliable framework for benchmarking task-to-model mapping approaches. The dataset comprises more than 99,000 scientific problem descriptions associated with over 500 unique NLP models. Each row in the Task2Model dataset contains a cleaned and filtered text snippet with the model name masked, capturing a scientific task or problem context, along with associated metadata. Columns: Scientific_problem_descriptions: Cleaned and filtered text snippets in which the model name is masked, capturing a scientific task or problem context. Full_ID: Canonical full identifier of the NLP model associated with the task (e.g., openai-community/gpt2). Short_ID: Normalized short identifier of the model, used as a collapsed label for analysis (e.g., gpt2). paper_id: Identifier linking the task description to the original scientific publication. family_label: Higher-level model family corresponding to the model (e.g., BERT, GPT, LLaMA), enabling hierarchical and family-level evaluation. Purpose and Use Cases: These datasets are designed to support research on how NLP models are mentioned, reported, and applied in scientific literature. They can be used for: Analyzing trends and common usage patterns of NLP models in research. Studying reporting practices and model adoption across scientific domains. Grouping models by families for aggregated and hierarchical analyses. Training and evaluating systems for model mention recognition and task-to-model mapping. Linking model mentions to their scientific task contexts for downstream empirical studies. Overall, these resources provide a foundation for systematic analysis of NLP model usage in scientific publications and for the development and evaluation of task-to-model recommendation and benchmarking methods.The full codebase for dataset construction, preprocessing, and split generation is publicly available athttps://github.com/khushbakhtahmed/Task2Model.

本研究发布EarlyBird与Task2Model两款数据集,二者均源自学术文献,旨在记录自然语言处理(Natural Language Processing,NLP)模型在科研文献中的提及情况,以及其与科学任务的关联方式。 EarlyBird(EarlyBird模型提及数据集)是一款中间级数据集,用于捕获单篇学术论文中所有被识别出的NLP模型提及内容。每条数据均附带丰富的元数据,包括论文标识符、章节标识与模型提及次数,支持对学术文献中模型的使用情况、报告规范及提及模式进行细粒度分析。 基于上述数据集,本研究进一步发布Task2Model,这是一款关联NLP模型与科学问题描述或任务的大规模数据集。Task2Model聚焦层级化与家族级评估、系统性质量审核,以及明确的防数据泄露措施,可为任务-模型映射方法的基准测试提供可靠框架。该数据集包含超过9.9万条科学问题描述,关联了500余种独特的NLP模型。 Task2Model数据集中的每一行数据均包含一段经清洗与过滤后的文本片段,其中模型名称已被掩码处理,用于捕获科学任务或问题语境,并附带相关元数据。 数据集字段说明: Scientific_problem_descriptions:经清洗与过滤后的文本片段,其中模型名称已被掩码处理,用于捕获科学任务或问题语境。 Full_ID:与该任务关联的NLP模型的标准完整标识符(例如:openai-community/gpt2)。 Short_ID:经过归一化处理的模型短标识符,用作分析时的聚合标签(例如:gpt2)。 paper_id:用于将任务描述关联至原始学术文献的标识符。 family_label:与该模型对应的高层级模型家族(例如:BERT、GPT、LLaMA),支持层级化与家族级分析。 数据集用途与应用场景: 本数据集旨在支持关于NLP模型在学术文献中如何被提及、报告与应用的相关研究,可应用于以下场景: - 分析科研领域中NLP模型的使用趋势与通用模式; - 研究不同科学领域中的报告规范与模型采用情况; - 按模型家族对模型进行分组,以开展聚合分析与层级化分析; - 训练与评估模型提及识别及任务-模型映射相关的系统; - 将模型提及内容与其科学任务语境相关联,以支持下游实证研究。 总体而言,上述两款数据集为学术文献中NLP模型使用情况的系统性分析,以及任务-模型推荐与基准测试方法的开发与评估提供了基础。数据集构建、预处理与数据集划分的完整代码库已在https://github.com/khushbakhtahmed/Task2Model公开。

提供机构:
Zenodo
创建时间:
2026-02-07
二维码
社区交流群
二维码
科研交流群
商业服务