遇见数据集

Mine, Yours, Ours? Sharing Data on Human Genetic Variation

收藏
Figshare2016-10-31 更新2026-04-29 收录
官方服务:

资源简介:

The achievement of a robust, effective and responsible form of data sharing is currently regarded as a priority for biological and bio-medical research. Empirical evaluations of data sharing may be regarded as an indispensable first step in the identification of critical aspects and the development of strategies aimed at increasing availability of research data for the scientific community as a whole. Research concerning human genetic variation represents a potential forerunner in the establishment of widespread sharing of primary datasets. However, no specific analysis has been conducted to date in order to ascertain whether the sharing of primary datasets is common-practice in this research field. To this aim, we analyzed a total of 543 mitochondrial and Y chromosomal datasets reported in 508 papers indexed in the Pubmed database from 2008 to 2011. A substantial portion of datasets (21.9%) was found to have been withheld, while neither strong editorial policies nor high impact factor proved to be effective in increasing the sharing rate beyond the current figure of 80.5%. Disaggregating datasets for research fields, we could observe a substantially lower sharing in medical than evolutionary and forensic genetics, more evident for whole mtDNA sequences (15.0% vs 99.6%). The low rate of positive responses to e-mail requests sent to corresponding authors of withheld datasets (28.6%) suggests that sharing should be regarded as a prerequisite for final paper acceptance, while making authors deposit their results in open online databases which provide data quality control seems to provide the best-practice standard. Finally, we estimated that 29.8% to 32.9% of total resources are used to generate withheld datasets, implying that an important portion of research funding does not produce shared knowledge. By making the scientific community and the public aware of this important aspect, we may help popularize a more effective culture of data sharing.

构建稳健、高效且负责任的数据共享模式,目前已被视为生物学与生物医学研究的核心要务。对数据共享开展实证评估,是明确关键问题、制定策略以提升科研数据面向全体科学界可获取性的不可或缺的首要环节。人类遗传变异相关研究,有望成为推动原始数据集广泛共享的先行领域。然而截至目前,尚无针对性分析用以明确原始数据集共享是否已成为该研究领域的通行惯例。为此,我们分析了2008年至2011年间收录于PubMed数据库的508篇论文中报道的共计543套线粒体与Y染色体数据集。研究发现,有相当比例(21.9%)的数据集未被共享;且无论是严格的期刊投稿政策,还是高影响因子期刊,均未能将共享率提升至当前80.5%的水平之上。按研究领域对数据集进行细分后可见,医学遗传学领域的数据共享率显著低于进化遗传学与法医学遗传学领域,这一差异在全线粒体DNA(mtDNA)序列数据集上表现得尤为显著(二者共享率分别为15.0%与99.6%)。向未共享数据集的通讯作者发送邮件索取数据的响应率仅为28.6%,这一结果表明,应将数据共享作为论文最终录用的前置条件;同时,要求研究者将研究结果上传至具备数据质量管控功能的开放在线数据库,似乎是最佳实践标准。最后,我们估算得出,未共享数据集的生成成本占总科研资源的29.8%至32.9%,这意味着有相当一部分科研经费并未转化为可共享的科研知识。通过向科学界与公众普及这一关键问题,我们或可助力构建更为高效的数据共享文化。

创建时间:
2016-10-31
二维码
社区交流群
二维码
科研交流群
商业服务