遇见数据集

Mendeley Data

收藏
知名数据库2026-06-11 收录
官方服务:

资源简介:

We developed a two-stream, Apache Solr-based information retrieval system in response to the bioCADDIE 2016 Dataset Retrieval Challenge. One stream was based on the principle of word embeddings, the other was rooted in ontology based indexing. Despite encountering several issues in the data, the evaluation procedure and the technologies used, the system performed quite well. We provide some pointers towards future work: in particular, we suggest that more work in query expansion could benefit future biomedical search engines.

为响应bioCADDIE 2016数据集检索挑战赛(bioCADDIE 2016 Dataset Retrieval Challenge),我们研发了一套基于Apache Solr的双流信息检索系统。该系统包含两个分支:其一基于词嵌入(word embeddings)原理,其二则依托基于本体的索引技术。尽管在数据、评估流程及所用技术环节均遭遇了若干问题,本系统仍展现出了优异的性能。我们针对后续研究提出了若干指引:具体而言,进一步开展查询扩展(query expansion)相关研究,将对未来的生物医学搜索引擎大有裨益。

提供机构:
Elsevier
搜集汇总
数据集介绍
Mendeley Data 数据集图片
背景与挑战
背景概述
该数据集包含Elsevier团队为bioCADDIE 2016数据集检索挑战赛开发的信息检索系统相关代码与数据。系统采用基于Apache Solr的双流架构,整合了词嵌入和本体索引技术,并提供了多个处理模块,如Aspire内容处理、Solr查询生成工具以及字典转换脚本。数据集还包括挑战赛的提交文件和详细的技术实现说明。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务