基于提取实体的生物医学搜索结果随机重排序

Stochastic reranking of biomedical search results based on extracted entities

Journal of the Association for Information Science and Technology (JASIST) · 2017
被引 6
ABS 3

中文导读

提出一种利用生物医学实体对搜索结果进行重排序的方法,通过实体提取构建文档与实体图,并用随机游走分析来提升低排名但相关的结果,实验表明能显著改善经典检索模型的效果。

Abstract

Health‐related information is nowadays accessible from many sources and is one of the most searched‐for topics on the Internet. However, existing search systems often fail to provide users with a good list of medical search results, especially for classic (keyword‐based) queries. In this article we elaborate on whether and how we can exploit biomedicine‐related entities from the emerging Web of Data for improving (through reranking) the results returned by a search system. The aim is to promote relevant but low‐ranked hits containing entities that are important to the current search context. We introduce an approach that is based on entity extraction applied on the retrieved documents, yielding a graph of documents along with entities, which in turn is analyzed probabilistically using a Random Walk‐based method. The proposed approach is independent of the submitted query and the underlying retrieval models, and thus can be applied over any ranked list of medical search results. Evaluation results using the data set of TREC Clinical Decision Support track demonstrate that the proposed approach can significantly improve the results returned by classic and widely applicable retrieval models. The results also enabled us to identify cases where the proposed reranking method fails to improve the ranking.

信息检索生物医学数据挖掘搜索引擎