QTL
收藏资源简介:
QTL数据集是由爱荷华州立大学创建的一个实际应用数据集,专注于动物科学领域的命名实体识别(NER)任务。该数据集包含18706条从PubMed精心挑选的与六种物种相关的定量性状位点(QTL)研究摘要,总计有18706个句子,514176个Tokens。数据集的创建过程中,收集了来自四个领域本体专业的3884个性状名称字典,用于远距离标注过程。QTL数据集主要用于解决动物基因组学研究和育种方法中的重要任务,即识别描述性表达的性状实体。
The QTL dataset is a practical application dataset developed by Iowa State University, focusing on the named entity recognition (NER) task within the field of animal science. This dataset includes 18,706 abstracts of quantitative trait locus (QTL) studies related to six species, meticulously selected from PubMed, with a total of 18,706 sentences and 514,176 Tokens. During the dataset construction process, a dictionary containing 3,884 trait terms from four domain-specific ontologies was collected for use in the distant annotation process. The QTL dataset is primarily utilized to address a critical task in animal genomics research and breeding approaches: identifying trait entities in descriptive expressions.

- 1Re-Examine Distantly Supervised NER: A New Benchmark and a Simple Approach爱荷华州立大学 · 2024年



