The Materials Data Facility
收藏资源简介:
MDF streamlines and automates data sharing, discovery, access and analysis by: 1) enabling data publication, regardless of data size, type, and location; 2) automating metadata extraction from submitted data into MDF metadata records (i.e., JSON formatted documents following the MDF schema) using open-source materials-aware extraction pipelines and ingest pipelines; and 3) unifying search across many materials data sources, including both MDF and other repositories with potentially different vocabularies and schemas. Currently, MDF stores 60 TB of data from simulation and experiment, and also indexes hundreds of datasets contained in external repositories, with millions of individual MDF metadata records created from these datasets to aid fine-grained discovery.
MDF简化并自动化了数据共享、发现、获取与分析全流程,具体实现方式如下:1) 支持数据发布,不受数据规模、类型与存储位置的限制;2) 依托开源的材料感知提取流水线与数据摄入流水线,自动从提交的数据中提取元数据 (metadata),生成遵循MDF元数据模式 (schema)的JSON格式元数据记录;3) 统一跨多类材料数据源的检索功能,涵盖MDF自身以及其他存在词汇表与架构差异的外部存储库。目前,MDF已存储60 TB来自模拟与实验的数据集,并为外部存储库中的数百个数据集建立了索引,同时基于这些数据集生成了数百万条独立的MDF元数据记录,以助力精细化的数据发现。




