遇见数据集

AYNEC-Datasets

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

These datasets are presented in the article "AYNEC: All You Need for Evaluating Completion Techniques in Knowledge Graphs", sent for the ESWC19. Please, cite it in your work if you make use of them. The following datasets are included: WN18-AF, generated from WN18. WN18-AR, generated from WN18, removing inverses. WN11-AF, generated from WN11. WN11-AR, generated from WN11, removing inverses. FB13-A, generated from FB13. FB15K-AF, generated from FB15K. FB15K-AR, generated from FB15K, keeping relations that cover 95% of the graph and removing inverses. NELL-AF, generated from NELL. NELL-AR, generated from NELL, keeping relations that cover 95% of the graph and removing inverses. In all datasets, we removed relations with only one instance, used 20% of each relation in the graph for test, generated one negative for each positive in both training and testing by replacing the target of the positive with a random entity. In WN11 and WN18 all entities are potential candidates. In the rest of datasets, only entities that have appeared as targets of the relation are candidates.<br> <br> Two relations were considered inverses when there was a 90% overlap between them. That is, relationc A and B are inverses if for 90% of instances of A there is an instance of B with inversed source and target, and vice-versa. When removing inverses, the smallest of each pair of inverses was removed. Each zip file contains the following files about a dataset: train.txt - triples used for training. Each line contains the source, the relation, the target, and the label (1 for positives and -1 for negatives). test.txt - triples used for testing, following the same format. relations.txt - a list of the relations in the dataset, each with its frequency. entities.txt - a list of the entities in the dataset, eac with its total degree, inwards degree, and output degree. inverses.txt - a list of the inverses in the original graph, whether or not they were removed. Each inverse relationship is represented by a pair of relations. summary.html - the visual summary of the relation frequencies and entity degrees (without removed inverses). dataset.gexf - the entire dataset in the open graph format "gexf", which can be opened by applications such as Gephi.

本数据集配套发表于投稿至ESWC 2019的论文《AYNEC:评估知识图谱补全技术所需的全部资源》(AYNEC: All You Need for Evaluating Completion Techniques in Knowledge Graphs)。若您在研究中使用本数据集,请引用该论文。 本次提供的数据集如下:从WN18衍生的WN18-AF;从WN18移除逆关系后得到的WN18-AR;从WN11衍生的WN11-AF;从WN11移除逆关系后得到的WN11-AR;从FB13衍生的FB13-A;从FB15K衍生的FB15K-AF;从FB15K中保留覆盖图谱95%的关系并移除逆关系后得到的FB15K-AR;从NELL衍生的NELL-AF;从NELL中保留覆盖图谱95%的关系并移除逆关系后得到的NELL-AR。 所有数据集均经过如下预处理:移除仅包含单个实例的关系;将图谱中每个关系的20%划分为测试集;为训练集与测试集中的每个正样本生成一个负样本,具体方式为将正样本的尾实体替换为随机选取的实体。其中,WN11与WN18的候选实体集合包含全部实体;其余数据集的候选实体仅限定为该关系中出现过的尾实体。 当两个关系的实例重叠率达到90%时,即判定二者为逆关系。具体而言,若关系A的90%实例均存在对应的关系B实例(二者的头实体与尾实体互换),且反之亦然,则关系A与B互为逆关系。在移除逆关系时,我们将每对逆关系中出现频次更低的一方移除。 每个数据集对应的压缩包均包含以下文件: - train.txt:训练集三元组文件,每行格式为头实体、关系、尾实体、标签(1代表正样本,-1代表负样本); - test.txt:测试集三元组文件,格式与训练集一致; - relations.txt:数据集包含的关系列表,每行记录一个关系及其出现频次; - entities.txt:数据集包含的实体列表,每行记录一个实体及其总度数、入度数、出度数; - inverses.txt:原始图谱中的逆关系列表(无论该逆关系是否已被移除),每个逆关系对以两个关系的形式表示; - summary.html:关系频次与实体度数的可视化统计页面(未包含已移除的逆关系); - dataset.gexf:以开放图谱格式GEXF存储的完整数据集,可通过Gephi等工具打开。

提供机构:
Zenodo
创建时间:
2019-02-14
二维码
社区交流群
二维码
科研交流群
商业服务