遇见数据集

Hackathon - TF-TG literature triage unlabelled data

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Once literature triage system is ready it is time to actually try to apply if to records that do not have any label in order to find the subset that does describe TF-TG interactions (are relevant). This is the corpus that has to be labeled by the systems created (hopefully) during the hackathon. To make the results more useful we have pre-selected records that do mention TFs by exploiting either automatic human TF mention recognition or external references from databases that have manually curated information on transcription factors (from GeneRif or UniProt). This means that these abstracts should be enriched with TF relevant records. This record has the same format as the training data except that the last column with the class label is missing. It contains PMIDs and Abstracts. Name: greekc_triage_unlabelled_v01.tsv Example: Format: tsv-separated columns (PMID, PubAnnotation JSON formated results of Pubtator for this record together with the automatically detected gene mentions using GnormPlus providing the Entrez Gene Identifiers together with the mention offsets, i.e. start and end character positions PubAnnotation format description: http://www.pubannotation.org/docs/annotation-format/ PubTator record retrieval description: https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/tmTools/curl.html <strong>Warning:</strong> This file is quite big!

提供机构:
Zenodo
创建时间:
2019-02-12
二维码
社区交流群
二维码
科研交流群
商业服务