Reference Dataset for Text Mining Type 2 Diabetes Candidate Genes
收藏资源简介:
The present disease-gene association data contains evidence or reference sentences which contain this disease-gene association information, which is further classified into 4 classes: Yes, No, Ambiguous and X each pertaining to Positive, Negative, Ambiguous and Not related disease-gene associations respectively. This data serves as reference data for the training text mining-based biological literature classifiers which can be used to predict classes of published literature, not just for Type 2 diabetes, but can also be expanded beyond to encompass a wide range of disease and their complications. The compilation of positively associated genes derived from these predictions can then be utilized for in-depth system-level analysis of T2D.
本疾病-基因关联(disease-gene association)数据集包含携带有该类疾病-基因关联信息的证据或参考文献语句,并进一步划分为4个类别:Yes、No、Ambiguous和X,分别对应正相关、负相关、模棱两可以及无关联的疾病-基因关联关系。该数据集可作为基于文本挖掘的生物文献分类器的训练参考数据,此类分类器不仅可用于预测2型糖尿病(Type 2 diabetes, T2D)相关的已发表文献类别,还可扩展覆盖多种疾病及其并发症。通过上述预测得到的阳性关联基因集合,可进一步用于2型糖尿病(T2D)的系统级深度分析。




