Evidence that identifiers are a source of problems for data integrators.
收藏资源简介:
Advances in computing power and expansion of the Internet have led to increasing optimism that big data will lead to new insights. However, in the life sciences, relevant data is not only "big"; it is also highly decentralized across thousands of online databases. Wringing value from it depends on the discipline of data science and on the humble bricks and mortar that make it possible -- identifiers. However, our collective handling of identifiers has lagged behind these advances. Diverse identifier problems (for instance broken links and ‘content drift’) make it difficult to integrate data and derive new knowledge from it. This is a snapshot of a living document intended to show real-world examples of identifier problems representative of those encountered by data integrators. It is not meant to be exhaustive.
随着计算能力的提升与互联网的普及,人们对于大数据将催生全新科学洞见的乐观情绪日益高涨。然而在生命科学领域,相关数据不仅体量庞大,还高度分散于数千个在线数据库之中。从这类数据中挖掘价值,既依赖数据科学的学科规范,也离不开支撑其实现的基础要素——标识符(identifiers)。然而,我们在标识符处理方面的整体进展却落后于上述技术革新。各类标识符相关问题(例如链接失效与内容漂移(content drift))使得数据整合与从中获取新知识的过程困难重重。本文件为动态更新的文档,此快照旨在展示数据整合工作者实际遭遇的各类典型标识符问题的真实案例。本快照并非旨在穷尽所有相关问题。



