Mendeley Data
收藏资源简介:
We developed a two-stream, Apache Solr-based information retrieval system in response to the bioCADDIE 2016 Dataset Retrieval Challenge. One stream was based on the principle of word embeddings, the other was rooted in ontology based indexing. Despite encountering several issues in the data, the evaluation procedure and the technologies used, the system performed quite well. We provide some pointers towards future work: in particular, we suggest that more work in query expansion could benefit future biomedical search engines.
为响应bioCADDIE 2016数据集检索挑战赛(bioCADDIE 2016 Dataset Retrieval Challenge),我们研发了一套基于Apache Solr的双流信息检索系统。该系统包含两个分支:其一基于词嵌入(word embeddings)原理,其二则依托基于本体的索引技术。尽管在数据、评估流程及所用技术环节均遭遇了若干问题,本系统仍展现出了优异的性能。我们针对后续研究提出了若干指引:具体而言,进一步开展查询扩展(query expansion)相关研究,将对未来的生物医学搜索引擎大有裨益。




