BIP! DB: A Dataset of Impact Measures for Research Products
收藏资源简介:
Overview This dataset contains citation-based impact indicators (also referred as measures) for ~321M distinct persistent identifiers (PIDs) that correspond to various types of research products (publications, datasets, software, and other products). The calculated indicators are organized into categories based on the aspect of impact they capture. Influence indicators Reflect the "total" impact of a research product; how established it is in general. Citation Count: The total number of citations of the product, the most well-known influence indicator. PageRank score: An influence indicator based on the PageRank (Page et al., 1999), a popular network analysis method. PageRank estimates the influence of each product based on its centrality in the whole citation network. It alleviates some issues of the Citation Count indicator (e.g., two products with the same number of citations can have significantly different PageRank scores if the aggregated influence of the products citing them is very different - the product receiving citations from more influential products will get a larger score). Popularity indicators Capture the "current" impact of a research product; how popular it currently is. RAM score: A popularity indicator based on the RAM (Ghosh et al., 2011) method. It is essentially a Citation Count where recent citations are considered as more important. This type of "time awareness" alleviates problems of methods like PageRank, which are biased against recently published products (new products need time to receive a number of citations that can be indicative for their impact). AttRank score: A popularity indicator based on the AttRank (Kanellos et al., 2020) method. AttRank alleviates PageRank's bias against recently published products by incorporating an attention-based mechanism, akin to a time-restricted version of preferential attachment, to explicitly capture a researcher's preference to examine products which received a lot of attention recently. Impulse indicators Measure the initial momentum that a research product received right after its publication. Incubation Citation Count (3-year CC): This impulse indicator is a time-restricted version of the Citation Count, where the time window length is fixed for all products and the time window depends on the publication date of the product, i.e., only citations 3 years after each product's publication are counted. FIeld-weighted indicators Capture the impact of a research product relative to the average performance in its field, accounting for differences in citation practices across disciplines. Field-Weighted Citation Impact (FWCI): A field-weighted indicator that measures how a research product performs compared to the global average in its research field. An FWCI of 1.0 indicates that the product is cited exactly as expected for similar publications in the same field; values above 1.0 indicate above-average impact, while values below 1.0 indicate below-average impact. 3-year FWCI: A time-restricted version of the FWCI that considers citations received within the first three years after publication. By limiting the citation window, this indicator captures the early relative impact of a research product, providing insight into how quickly it gains influence in its field. In our analysis, the expected number of citations for each research product is computed by grouping them by concept, publication year, and product type and then averaging the citations within each group. More details about the aforementioned impact indicators, the way they are calculated and their interpretation can be found here and in the respective references (Kanellos et al., 2019). Indicator calculation levels The impact indicators are calculated in two levels: PID level: assuming that each PID corresponds to a distinct research product. Currently PIDs are DOIs, PMCIDs, and PMIDs. OpenAIRE-id level: leveraging PID synonyms based on OpenAIRE's deduplication algorithm (Manghi et al., 2020) - each distinct article has its own OpenAIRE id. Impact classes Each researcj product is also assigned an impact class, reflecting its percentile rank among all products in the dataset: Class Percentile Description C1 Top 0.01% Exceptional impact C2 Top 0.1% Very high impact C3 Top 1% High impact C4 Top 10% Good impact C5 Rest 90% Remaining products File structure For each calculation level (PID / OpenAIRE-id) we provide five (5) compressed CSV files (one for each measure/score provided). The structure of the files differs slightly depending on the level: PID-level files: Each line follows the format:identifier <tab> identifier_type <tab> score <tab> class OpenAIRE-id-level files: These files contain the keyword "openaire_ids" in the filename. Each line follows the format:identifier <tab> score <tab> class The parameter setting of each measure is encoded in the corresponding filename. For more details on the different measures/scores see our extensive experimental study (Kanellos et al., 2019) and the configuration of AttRank in the original paper (Kanellos et al., 2020). Topic-related files In addition to the main indicator files, the dataset also includes topic-level outputs, providing field-weighted impact indicators as well percentile classes within the associated 2nd-level concepts from OpenAlex. Specifically, we associated all research products with their 2nd level concepts from OpenAlex (using only their DOIs); we kept only the three most dominant concepts for each product, based on their confidence score, and only if this score was greater than 0.3. Since currently only the DOIs are used to associate concepts from OpenAlex to research products, all identifiers in these files refer to DOIs. Topic-specific impact classes file: Fore each concept and indicator, precentile classes are computed and provided in topic_based_impact_classes.txt in the following format: identifier <tab> concept <tab> pagerank_class <tab> attrank_class <tab> 3-cc_class <tab> cc_class Field-weighted indicator files: Each line follows the format:identifier <tab> concept <tab> score Note that to prevent division by zero, the score column is left empty whenever the average score for a specific combination of concept, publication year, and product type equals zero. Data sources The data used to produce the citation network on which we calculated the provided measures have been gathered from the OpenAIRE Graph v10.8.1, including data from (a) OpenCitations' COCI & POCI dataset, (b) MAG (Sinha et al, 2015; Wang et al., 2019), and (c) Crossref. The union of all distinct citations that could be found in these sources have been considered. Additionally, all topic-related computations are derived from OpenAlex concepts. Access and Use Find our Academic Search Engine built on top of these data here. Further note, that we also provide all calculated scores through BIP! Finder's API. Terms: These data are provided "as is", without any warranties of any kind. The data are provided under the CC0 license. Changelog v19.1 [major update] Added field-weighted indicators: FWCI and 3-year FWCI. v19.0 Added PMCID as an additional type of PID. v15.1 Fixed missing records that were unintentionally omitted in v15.0 Ensures all popularity indicators correctly use current_year = 2025 v12.0 Added PMIDs as an additional type of PID. v10.0 [Major update] Introduced deduplication of research products using the latest OpenAIRE article deduplication algorithm. Each node in the citation network is now a deduplicated product having a distinct OpenAIRE id. Corrected overcounting of citations caused by multiple versions of the same product. PID-level scores are now derived from deduplicated OpenAIRE nodes. Added filtering rules described here to remove from dataset PIDs with problematic metadata. v9.0 [Major update] Introduced topic-specific impact classes for PID-identified products based on OpenAlex 2nd-level concepts. v7.0 [Major update] Added impact class labels (C1-C5) for each procuct, indicating the percentile-bsaed impact levels. Classes reflect relative position within the global score distribution. v5.1 [Major update] Introduced dual-level score computation: PID level and OpenAIRE ID level.
### 概览 本数据集包含约3.21亿个不同持久标识符(Persistent Identifier, PID)对应的基于引用的影响指标(亦称为测度),这些标识符对应各类研究成果(包括学术论文、数据集、软件及其他类型成果)。 所计算的影响指标根据其捕捉的影响维度被划分为不同类别。 #### 影响力指标 用于反映研究成果的“整体”影响力,即其在学术领域的总体认可度。 **引用计数(Citation Count)**:指研究成果的总被引次数,是最广为人知的影响力指标。 **PageRank得分**:基于PageRank算法(Page等人,1999)的影响力指标,该算法是一种主流的网络分析方法。PageRank通过研究成果在整个引用网络中的中心性来评估其影响力,可缓解引用计数指标的部分缺陷(例如:若两篇成果的被引次数相同,但引用它们的成果的整体影响力差异显著,则二者的PageRank得分会有明显不同——从影响力更高的成果获得引用的成果将获得更高得分)。 #### 活跃度指标 用于捕捉研究成果的“当下”影响力,即其当前的受关注程度。 **RAM得分**:基于RAM方法(Ghosh等人,2011)的活跃度指标,本质上是对近期引用赋予更高权重的引用计数。这种“时间敏感性”设计可缓解PageRank等方法对新近发表成果的偏向性问题(新成果需要时间积累足够的被引次数,才能体现其影响力)。 **AttRank得分**:基于AttRank方法(Kanellos等人,2020)的活跃度指标。AttRank通过引入基于注意力的机制(类似受时间约束的优先连接模型),显式捕捉研究者对近期受关注度较高成果的关注偏好,从而缓解PageRank对新近发表成果的偏向性问题。 #### 脉冲影响力指标 用于衡量研究成果在发表后即刻获得的初始发展势头。 **孵化期引用计数(Incubation Citation Count,简称3年CC)**:该脉冲影响力指标是受时间窗口约束的引用计数变体,所有成果的时间窗口长度固定,且窗口起始时间基于成果的发表日期——即仅统计成果发表后3年内的被引次数。 #### 领域加权指标 用于衡量研究成果相对于其所在领域平均表现的影响力,可修正不同学科间引用习惯差异带来的偏差。 **领域加权引用影响(Field-Weighted Citation Impact,FWCI)**:该领域加权指标用于评估研究成果相对于其所在研究领域的全球平均表现的情况。FWCI值为1.0时,表示该成果的被引情况与同领域同类成果的预期一致;值高于1.0则代表影响力高于领域平均水平,低于1.0则代表影响力低于领域平均水平。 **3年FWCI**:FWCI的时间约束变体,仅统计成果发表后前3年内的被引次数。通过限定引用窗口,该指标可捕捉研究成果的早期相对影响力,帮助分析其在领域内的影响力积累速度。 在本研究的分析中,每篇研究成果的预期被引次数通过以下方式计算:按研究主题、发表年份及成果类型对成果进行分组,再计算每组内的平均被引次数。 关于上述影响指标的更多细节、计算方法及解读方式,可参考本文及相关参考文献(Kanellos等人,2019)。 ### 指标计算层级 本数据集的影响指标通过两种层级计算: **PID层级**:假设每个PID对应唯一的研究成果。当前支持的PID类型包括数字对象标识符(Digital Object Identifier, DOI)、PubMed Central标识符(PMCID)及PubMed标识符(PMID)。 **OpenAIRE ID层级**:基于OpenAIRE的去重算法(Manghi等人,2020),利用PID的同义词映射,每个独立的学术文章对应唯一的OpenAIRE ID。 ### 影响力类别 每篇研究成果还会被分配一个影响力类别,反映其在数据集中所有成果中的百分位排名: | 类别 | 百分位范围 | 描述 | |------|------------|------| | C1 | 前0.01% | 影响力卓越 | | C2 | 前0.1% | 影响力极高 | | C3 | 前1% | 影响力较高 | | C4 | 前10% | 影响力良好 | | C5 | 其余90% | 普通影响力 | ### 文件结构 针对两种计算层级(PID层级/OpenAIRE ID层级),本数据集分别提供5个压缩后的CSV文件(每个文件对应一种指标/得分)。不同层级的文件结构略有差异: **PID层级文件**:每行遵循以下格式:标识符 <制表符> 标识符类型 <制表符> 得分 <制表符> 类别 **OpenAIRE ID层级文件**:此类文件的文件名中包含关键词“openaire_ids”,每行遵循以下格式:标识符 <制表符> 得分 <制表符> 类别 各项指标的参数设置已编码至对应文件名中。关于各类指标/得分的更多细节,可参考本团队的大规模实验研究(Kanellos等人,2019)及AttRank方法的原始论文(Kanellos等人,2020)。 ### 主题相关文件 除主指标文件外,本数据集还包含主题层级的输出结果,提供领域加权影响指标,以及基于OpenAlex二级主题的百分位类别。 具体而言,本数据集仅通过数字对象标识符(DOI)将所有研究成果与OpenAlex的二级主题进行关联;仅保留每个成果置信度高于0.3的前三个主导主题。 由于当前仅通过DOI将研究成果与OpenAlex主题进行关联,因此此类文件中的所有标识符均为DOI。 **主题专属影响力类别文件**:针对每个主题与指标,本数据集会计算百分位类别,并保存至`topic_based_impact_classes.txt`文件中,格式如下: 标识符 <制表符> 主题 <制表符> PageRank类别 <制表符> AttRank类别 <制表符> 3年CC类别 <制表符> 引用计数类别 **领域加权指标文件**:每行遵循以下格式:标识符 <制表符> 主题 <制表符> 得分 请注意:若某一主题、发表年份及成果类型组合的平均得分为零,则为避免除以零的错误,对应行的得分列将留空。 ### 数据来源 本数据集用于构建引用网络的数据源自OpenAIRE Graph v10.8.1,具体包括:(a) OpenCitations的COCI与POCI数据集;(b) MAG学术图谱(Sinha等人,2015;Wang等人,2019);(c) Crossref元数据平台。本研究整合了所有可从上述来源获取的唯一引用记录。 此外,所有主题相关的计算均基于OpenAlex的主题体系。 ### 获取与使用 可通过本团队基于该数据集构建的学术搜索引擎访问数据。此外,本数据集的所有计算得分还可通过BIP! Finder的应用程序编程接口(API)获取。 ### 使用条款 本数据集按“现状”提供,不附带任何形式的担保。数据集采用CC0许可协议发布。 ### 更新日志 #### v19.1 **重大更新**:新增领域加权指标FWCI与3年FWCI。 #### v19.0 新增PMCID作为支持的PID类型之一。 #### v15.1 修复了v15.0版本中因疏忽遗漏的部分记录;确保所有活跃度指标均正确使用当前年份2025。 #### v12.0 新增PMIDs作为支持的PID类型之一。 #### v10.0 **重大更新**:采用最新的OpenAIRE学术文章去重算法对研究成果进行去重处理。如今引用网络中的每个节点均为去重后的成果,对应唯一的OpenAIRE ID。修复了同一成果的多个版本导致的引用重复计数问题;PID层级的得分现在基于去重后的OpenAIRE节点计算;新增过滤规则,用于从数据集中移除元数据存在问题的PID。 #### v9.0 **重大更新**:基于OpenAlex二级主题,为PID标识的成果新增主题专属影响力类别。 #### v7.0 **重大更新**:为每篇研究成果新增影响力类别标签(C1-C5),用于标识基于百分位的影响力等级;该类别反映成果在全球得分分布中的相对位置。 #### v5.1 **重大更新**:引入双层级得分计算:PID层级与OpenAIRE ID层级。



