The BioC-BioGRID corpus contains human annotations on 120 full text biomedical literature articles for genetic and protein interaction data. The annotated corpus contains 6409 mentions of genes and th
Samples of disease-gene association data set used in Master thesis project of Nuttapong Mekvipad. The data sets were automatically generated using distant supervision as described in BERT-based contex
The numbers of sentence pairs, genes, chemicals, and relationships for four diseases by DigChem, and the numbers of triple relationships from CTD, DrugBank, and WDD for four diseases.