遇见数据集

The entire OpenCitations Corpus triplestore data dump, archived on 2017-10-25

收藏
DataCite Commons2020-09-01 更新2024-07-25 收录
官方服务:

资源简介:

This archive contains the dump of the OpenCitations Corpus (OCC, http://opencitations.net) triplestore (Blazegraph, https://www.blazegraph.com/, licensed in GPLv2) containing all the data of the corpus, and created regularly every month.<br><br>After unzipping the archive, Disk ARchive (DAR, http://dar.linux.free.fr/, a multi-platform archive tool for managing huge amount of data) is needed for recreating the whole structure. For extracting the DAR archive, please run the command<br><br>dar -x [archive-name]<br><br>Where "[archive-name"] is the name of the DAR file without final package number and extension. E.g.:<br><br>dar -x 2016-09-23-triplestore<br><br>Please execute "run.sh" for running the triplestore, and use "stop.sh" for stopping it.<br><br>For further questions, comments, and suggestions please don't hesitate to contact Silvio Peroni at essepuntato@opencitations.net.

本归档文件包含开放引用语料库(OpenCitations Corpus,简称OCC,http://opencitations.net)的三元组存储库(triplestore)转储文件。该存储库基于Blazegraph(https://www.blazegraph.com/,采用GPLv2许可证)构建,涵盖该语料库的全部数据,且每月定期生成。<br><br>解压该归档文件后,需使用磁盘归档工具(Disk ARchive,简称DAR,http://dar.linux.free.fr/,一款用于管理海量数据的多平台归档工具)来还原完整的数据结构。若要解压DAR归档文件,请执行以下命令:<br><br>dar -x [归档文件名]<br><br>其中「[归档文件名]」指DAR文件的名称,需省略末尾的包编号与扩展名。例如:<br><br>dar -x 2016-09-23-triplestore<br><br>如需启动该三元组存储库,请执行`run.sh`脚本;如需停止,则执行`stop.sh`脚本。<br><br>如有任何疑问、意见或建议,请随时联系Silvio Peroni,邮箱地址为essepuntato@opencitations.net。

提供机构:
figshare
创建时间:
2017-11-06
二维码
社区交流群
二维码
科研交流群
商业服务