遇见数据集

INDRA assembly Benchmark Corpus

收藏
Zenodo2023-01-22 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This data set accompanies the manuscript "Automated assembly of molecular mechanisms at scale from text mining and curated databases" which describes assembly methodology implemented in the INDRA system (https://github.com/sorgerlab/indra). The manuscript uses an example assembly pipeline on ~570k publications as input to create the INDRA Benchmark Corpus. This dataset provides INDRA Statements constituting the INDRA Benchmark Corpus as well as a set of curations on the corpus: - indra_benchmark_corpus.pkl: A Python pickle file of INDRA Statement objects. It requires INDRA to be installed to load in a Python environment. - indra_benchmark_corpus.json.gz: A gzipped JSON export of INDRA Statements. - indra_benchmark_corpus_curations.pkl: Aggregated curations on the Benchmark Corpus as a Python pickle file. The content is a list, where each element of the list is a dictionary providing a stmt_hash corresponding to the hash of a Statement in the Benchmark Corpus, and giving a correct flag (0 for overall incorrect, 1 for overall correct) to represent the overall result of curation, along with some other metadata. - indra_benchmark_corpus_curations.json: Aggregated curations on the Benchmark Corpus as a JSON file.

提供机构:
Zenodo
创建时间:
2022-11-03
二维码
社区交流群
二维码
科研交流群
商业服务