遇见数据集

Citation network dataset covering the work of RP Millar and its citing literature

收藏
Zenodo2024-06-09 更新2026-05-26 收录
官方服务:

资源简介:

The following describes the citation network datasets that underpins the manuscript “A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research” [1]. Data collection We retrieved data from the Web of Science Core Collection under the University of Edinburgh’s subscription in January 2024. We sought to retrieve all indexed papers of Professor Robert P. Millar (RPM). We searched the following AU = (Millar, R), and then retrieved records that corresponded to his WoS profile (n=428 records) and an additional 49 paper that were authored by Robert but had not been included in his WoS record – validating the records against a CV of his published works. We retrieved the full citation history as record by Web of Science to these papers from other indexed records. The 477 RPM papers had been cited 21,677 times by 11,138 documents by date of retrieval, and removing self-citations left 19,256 citations by 10,719 documents. We then retrieved all metadata from WoS concerning the 477 RPM papers and the 10,719 citation papers, resulting in a dataset covering 11,196 documents. Citation network dataset We constructed a citation network dataset by parsing data from each paper’s full bibliography consisting of: i. ‘Edge-list’ that records citation links from a citing to a cited document. This is constructed by assigning unique IDs to each retrieved paper and to every unique reference string contained in their bibliographies. The edge list is composed of a ‘Source’ column that contains the ID of the citing document and a ‘Target’ column containing the IDs of its citations, with one record per row. Given that we were only interested in citations between the WoS retrieved documents, we discarded any reference string that represented a document outwith our search. ii. ‘Node-attribute list’ that contains the ID, with relevant metadata contained in adjacent columns to identify documents, including authors, title of publication, journal, year of publication. We also parsed into this dataset the WoS full citation count for each paper and the total number of references in the bibliographies of each paper. This results in a dataset containing 11,196 nodes and 115,834 edges between nodes. We removed a total of 67 papers for which metadata was incomplete and/or corrupted. We further focussed on the largest interconnected component, removing nodes with no connections (isolates) or smaller components that were detached from the main network. We excluded papers <10 references to remove meeting abstracts and other minor journal items, and papers not published in English. This resulted in a final dataset containing 10,901 nodes and 113,742 edges, and it is this dataset that we share as it is the basis for the analyses within the paper. Description of dataset variables ‘RPM_Edgelist.csv’ is a comma-separate values file that consists of all 113,742 citations between the 10,901 documents of the citation network analysed in the manuscript. The columns refer to: ‘Source’, the unique identifier for the citing document ‘Target’, the unique identifier for the cited document ‘Syr’, the year of publication of the citing document ‘Tyr’, the year of publication of the cited document ‘SC’, the cluster ID of the citing document ‘TC’, the cluster ID of the cited document ‘RPM_Nodelist.csv’ is a comma-separate values file that consists of the 10,901 documents of the citation network analysed in the manuscript. The columns refer to: ‘Id’, the unique ID assigned to a document that corresponds with the edgelist ‘Reference string’, the reference string of the document ‘WoS ID’, the unique accession number assigned to a document by the Web of Science. These can be used to query WoS to find further data on all papers via the ‘UT= ’ field tag. ‘Authors’, all authors formatted by full last name and initials ‘# of authors’, number of authors ‘Title’, title of document ‘Publication year’, publication year of document ‘Document type’, document type defined by WoS (e.g. article, review, etc.) ‘Total references’, total number of references within a documents bibliography as recorded by WoS ‘Total WoS citations’, total number of citations recorded to a document from other documents indexed in the Web of Science ‘Indegree’, total number of within network citations (i.e. counting only citations from other papers retrieved by our query) ‘Outdegree’, total number of within network references (i.e. counting only reference to other papers retrieved by our query) ‘Degree’, total number of node connections (i.e. indegree + outdegree) ‘Class’, variable used to distinguish between RPM’s publications (‘RPM’) and the citing documents (‘CITE’) ‘Cluster’, provides the cluster membership number as discussed within the manuscript. This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.67 | 25 clusters). References [1] Leng, R. I., Leng. G. (Under review). A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research. J. Neuroendocrinol All bibliographic data included in this study are derived originally from Clarivate™ (Web of Science™) and downloaded in January 2024. © Clarivate 2024. All rights reserved.

以下为本手稿《A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH研究》[1]所依托的引文网络数据集说明。 数据采集 本研究于2024年1月通过爱丁堡大学订阅权限,从Web of Science核心合集(Web of Science Core Collection)中获取数据。我们旨在获取罗伯特·P·米勒(Robert P. Millar,简称RPM)教授所有已被Web of Science收录的论文。首先通过作者字段AU = (Millar, R)进行检索,获取了匹配其Web of Science个人主页的记录共428条;此外还补充检索到49篇由罗伯特·P·米勒撰写但未被纳入其Web of Science个人主页的论文,并通过其已发表成果的简历对所有记录进行了核验。 随后,我们通过Web of Science获取了上述477篇RPM论文被其他已收录文献引用的完整引文历史。截至数据检索当日,这477篇RPM论文累计被11138篇文献引用共计21677次;去除自引后,剩余19256次引用,涉及10719篇文献。接着,我们从Web of Science中获取了这477篇RPM论文以及10719篇引用文献的全部元数据,最终得到覆盖11196篇文献的数据集。 引文网络数据集 我们通过解析每篇文献的完整参考文献列表构建引文网络数据集,具体包含以下两部分: i. 边列表(Edge-list):用于记录施引文献至被引文献的引文关联。构建方式为:为每一篇已检索获取的文献,以及其参考文献列表中每一条唯一的参考文献字符串分配唯一标识符。边列表包含两列:"Source"列记录施引文献的唯一ID,"Target"列记录其引用的被引文献ID,每行对应一条记录。鉴于本研究仅关注Web of Science检索范围内文献间的引用关系,我们剔除了所有不属于本次检索范围的参考文献字符串。 ii. 节点属性列表(Node-attribute list):包含文献的唯一标识符,相邻列附带用于标识文献的相关元数据,包括作者、文献标题、期刊名称、出版年份。此外,我们还将每篇文献的Web of Science总被引频次,以及其参考文献列表中的总参考文献数纳入该数据集。 由此得到的初始数据集包含11196个节点与115834条节点间边。我们剔除了共67条元数据不完整或已损坏的文献记录。随后,我们仅保留最大连通分量,移除了无连接的孤立节点以及与主网络脱离的小型连通分量。此外,我们剔除了参考文献数少于10的文献(以去除会议摘要及其他小型期刊文献),以及非英文出版的文献。最终得到的数据集包含10901个节点与113742条边,本研究共享的正是该数据集,其为论文内所有分析的基础。 数据集变量说明 "RPM_Edgelist.csv"为逗号分隔值文件,包含本论文分析的引文网络中10901篇文献间的全部113742条引用关系,各字段含义如下: - "Source":施引文献的唯一标识符 - "Target":被引文献的唯一标识符 - "Syr":施引文献的出版年份 - "Tyr":被引文献的出版年份 - "SC":施引文献的聚类ID - "TC":被引文献的聚类ID "RPM_Nodelist.csv"为逗号分隔值文件,包含本论文分析的引文网络中的10901篇文献,各字段含义如下: - "Id":分配给文献的唯一标识符,与边列表中的ID对应 - "Reference string":文献的参考文献字符串 - "WoS ID":Web of Science为文献分配的唯一收录编号,可通过"UT="字段标识在Web of Science中检索该文献以获取更多相关数据 - "Authors":所有作者,格式为完整姓氏加名字首字母缩写 - "# of authors":作者总数 - "Title":文献标题 - "Publication year":文献出版年份 - "Document type":Web of Science定义的文献类型(如期刊论文、综述等) - "Total references":Web of Science记录的该文献参考文献列表中的总参考文献数 - "Total WoS citations":Web of Science记录的该文献被其他已收录文献引用的总次数 - "Indegree":网络内总被引次数(即仅统计本研究检索范围内其他文献对该文献的引用) - "Outdegree":网络内总施引次数(即仅统计该文献对本研究检索范围内其他文献的引用) - "Degree":节点总连接数(即Indegree与Outdegree之和) - "Class":用于区分RPM发表的文献(标注为"RPM")与施引文献(标注为"CITE")的分类变量 - "Cluster":聚类成员编号,与论文中讨论的聚类结果一致。该聚类通过莱顿算法(Leiden algorithm)最大化模块度得到(结果1;Q=0.67,共25个聚类)。 参考文献 [1] Leng, R. I., Leng, G.(已投稿待审). 以数为笔:RP米勒的学术生涯——引文网络分析其研究工作及对促性腺激素释放激素(Gonadotropin-Releasing Hormone, GnRH)研究的贡献. J. Neuroendocrinol 本研究涉及的所有文献元数据最初均来自科睿唯安(Clarivate™)旗下的Web of Science™,并于2024年1月下载获取。© 科睿唯安2024年版权所有,保留一切权利。

提供机构:
Zenodo
创建时间:
2024-06-09
二维码
社区交流群
二维码
科研交流群
商业服务