Development and evaluation of a biomedical search engine using a predicate-based vector space model

Development and evaluation of a biomedical search engine using a predicate-based vector space model
复制标题

DOI:
10.1016/j.jbi.2013.07.006
复制
发表时间:
2013-10-01
影响因子:
4.5
通讯作者:
Harwell, Jeffrey
Harwell, Jeffrey
中科院分区:
医学3区
文献类型:
--
作者:
Kwak, Myungjae;Leroy, Gondy;Harwell, Jeffrey

文献摘要

被引文献

相似文献

尽管文章和专利中的生物医学信息呈指数级增长,但我们仍然依赖相同的信息检索方法,使用很少的关键字来搜索数百万文档。我们正在开发一种从根本上不同的方法,通过使用谓词而不是关键字进行查询和文档表示的单个查询来查找更精确和完整的信息。谓词是三元组,是比关键字更复杂的数据结构,包含更多的结构化信息。为了更好地利用它们,我们提出了一个新的基于谓词的向量空间模型和查询文档相似度函数,并调整了tf-idf和boost函数。使用107,367 PubMed摘要的测试床,我们评估了第一个基本功能:检索信息。癌症研究人员提供了20个真实的查询,其中前15个摘要使用基于谓词(新)和基于关键词(基线)的方法进行检索。癌症研究人员对每份摘要进行双盲评估,以0-5分的量表计算精确度(0分与更高分)和相关性(0-5分)。基于谓词的方法(80%)的精确度显著高于基于关键词的方法(71%)(p <0.001)。使用基于谓词的方法,相关性几乎增加了一倍--在不调整排名顺序的情况下,相关性为2.1比1.6(p < .001),在调整排名顺序的情况下,相关性为1.34比0.98(p < .001)。谓词可以支持比关键字更精确的搜索,为丰富和复杂的信息搜索奠定了基础。(C)2013 Elsevier Inc. All rights reserved.
Although biomedical information available in articles and patents is increasing exponentially, we continue to rely on the same information retrieval methods and use very few keywords to search millions of documents. We are developing a fundamentally different approach for finding much more precise and complete information with a single query using predicates instead of keywords for both query and document representation. Predicates are triples that are more complex datastructures than keywords and contain more structured information. To make optimal use of them, we developed a new predicate-based vector space model and query-document similarity function with adjusted tf-idf and boost function. Using a test bed of 107,367 PubMed abstracts, we evaluated the first essential function: retrieving information. Cancer researchers provided 20 realistic queries, for which the top 15 abstracts were retrieved using a predicate-based (new) and keyword-based (baseline) approach. Each abstract was evaluated, double-blind, by cancer researchers on a 0-5 point scale to calculate precision (0 versus higher) and relevance (0-5 score). Precision was significantly higher (p < .001) for the predicate-based (80%) than for the keyword-based (71%) approach. Relevance was almost doubled with the predicate-based approach-2.1 versus 1.6 without rank order adjustment (p < .001) and 1.34 versus 0.98 with rank order adjustment (p < .001) for predicate versus keyword-based approach respectively. Predicates can support more precise searching than keywords, laying the foundation for rich and sophisticated information search. (C) 2013 Elsevier Inc. All rights reserved.