Synthetic datasets for end-to-end Relation Extraction of relationships between Organisms and Natural-Products
收藏资源简介:
Synthetic datasets (training/validation) for end-to-end Relation Extraction of relationships between Organisms and Natural-Products. The datasets are provided for reproducibility purposes, but, can also be used to train new models. As in the corresponding article, 3 subtypes of synthetic datasets are provided: Diversity-synt: The seed literature references used in the generation process correspond to the top-500 extracted items per biological kingdoms using the GME-sampler. Random-synt: 5 datasets of equivalent sizes as Diversity-synt, but using randomly sampled seed literature references. Extended-synt: A merge of Diversity-synt and the 5 Random-synt datasets. All datasets were produced with Vicuna-13b-v1.3. Like the model, the produced synthetic data are also submitted to the License of the model used for generation, see the original LLaMA model card. LLaMA is licensed under the LLaMA License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.
用于生物与天然产物间关系端到端关系抽取任务的合成数据集(训练集/验证集)。本数据集旨在支持研究可复现,同时亦可用于训练新型模型。 如对应研究论文所述,本次发布共包含三类合成数据集亚型: 多样性合成数据集(Diversity-synt):生成过程所使用的种子文献参考文献,为通过GME采样器(GME-sampler)按各生物界提取的前500条条目。 随机合成数据集(Random-synt):共包含5个规模与多样性合成数据集一致的子集,但其种子文献参考文献均采用随机采样方式获取。 扩展合成数据集(Extended-synt):为多样性合成数据集与5个随机合成数据集的合并集合。 所有数据集均基于Vicuna-13b-v1.3生成。与该模型一致,生成的合成数据同样需遵守生成所用模型的许可协议,详见原始LLaMA模型卡片(LLaMA model card)。 LLaMA模型采用LLaMA许可协议进行授权,版权所有 © Meta Platforms公司,保留所有权利。



