PubMed Query Log Analysis and Use in Access Enhancement
PubMed Query Log Analysis and Use in Access Enhancement
批准号:
8177730
负责人:
Willy Wilbur
金额:
$48.97万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
BooksCategoriesComputational TechniqueDataDatabasesDiseaseDrug FormulationsEventGene ProteinsGenesGoalsHealthHereditary DiseaseHumanInformation ServicesInternationalInvestigationLearningLinkLiteratureMapsMethodsNamesPeer ReviewPharmaceutical PreparationsPopulationPubMedRecordsResearchResourcesServicesTechniquesTextabstractingbasefallsimprovedmeetingsoperationresponsesensorsuccesstext searching
中文摘要
生物医学文献检索是获取不断增加的信息的主要入口。PubMed/MEDLINE是用于此目的的最广泛的服务。然而,在PubMed中找到与用户信息需求相关的引文并不总是容易的。提高我们对PubMed用户不断增长的人口,他们的信息需求以及他们满足这些需求的方式的理解,为改善PubMed提供的信息服务和信息访问提供了机会。用于理解和表征搜索引擎的顾客的一个资源是事务日志。我们以前的用户查询日志的调查,使我们开发和部署一个有用的应用程序,在帮助用户查询制定在PubMed,即相关检索(RQ)。受其成功的启发,我们继续使用日志分析来确定与PubMed操作密切相关的研究问题。
例如,通过我们对PubMed日志的分析,我们了解到人们搜索某些生物医学概念的频率高于其他概念,并且不同概念之间存在很强的关联。例如,疾病名称通常与基因/蛋白质和药物名称共同出现。为此,我们组织了一个国际挑战赛,以自动识别基因/蛋白质全文。挑战中提出的成功技术可能会被NCBI用来增强其更好地将基因记录与文献联系起来的能力。我们还在PubMed中开发并部署了一种自动方法,以识别用户查询中的疾病概念(称为PubMeds疾病传感器)。这种传感器为PubMed用户提供了PubMed中文章之外的其他相关信息。例如,这种特殊的疾病传感器将用户链接到GeneReviews中的相关章节,这是一本由专家撰写和同行评审的关于各种遗传疾病的书。最后,我们通过使用标准的本体映射,集成了来自各种权威资源(例如DailyMed的药物适应症字段)的疾病-药物关系的自动提取结果。这些结果将用于丰富国家协调机构卫生相关数据库中不同记录之间的联系。
此外,我们还继续开发计算技术,以应对PubMed中返回零结果的查询。正如我们的日志分析所示,大约15%的PubMed搜索属于这一类。在某些情况下,确实没有文档或摘要可以满足特定的查询。 然而,在分析提交给PubMed的一个月的查询时,我们发现,通常情况下,没有检索到结果的查询如果构造不同,则会检索到相关的内容。在学习人类如何修改不成功的查询的基础上,我们成功地开发了一种自动方法,通过删除查询词将失败的查询转化为成功的查询,同时最大限度地保留原始用户搜索意图。
英文摘要
Biomedical literature search is the main entry point for an ever-increasing range of information. PubMed/MEDLINE is the most widely used service for this purpose. However, finding citations relevant to a users information need is not always easy in PubMed. Improving our understanding of the growing population of PubMed users, their information needs and the way in which they meet these needs opens opportunities to improve information services and information access provided by PubMed. One resource for understanding and characterizing patrons of search engines is the transaction logs. Our previous investigation of user query logs has led us to develop and deploy a useful application in assisting user query formulation in PubMed, namely Related Queries (RQ). Inspired by its success, we have continued using log analysis to identify research problems which are closely related to PubMed operations.
For instance, through our analysis of PubMed logs, we learn that people search certain biomedical concepts more often than others and that there exist strong associations between different concepts. For example, a disease name often co-occurs with gene/protein and drug names. To this end, we have organized an international challenge event for automatically identifying gene/proteins in full text. Successful techniques presented in the challenge may be used by NCBI to enhance its ability to better link gene records to literature. We have also developed and deployed an automatic method in PubMed to recognize disease concepts in user queries (known as PubMeds disease sensor). Such a sensor provides PubMed users with additional relevant information beyond articles in PubMed. For instance, this particular disease sensor links users to related chapters in GeneReviews, an expert-authored and peer-reviewed book of various genetic diseases. Finally, we have integrated automatic extraction results of disease-drug relations from various authoritative resources (e.g. drug indication field from DailyMed) through the use of standard ontological mappings. Such results will be used to enrich links among different records in NCBIs health related databases.
In addition, we have continued developing computational techniques in response to queries that return zero results in PubMed. As shown in our log analysis, about 15% of PubMed searches fall into this category. In some cases there really is no document or abstract that will satisfy a particular query. However, in analyzing one month of queries submitted to PubMed, we find that more often than not, queries that retrieved no results are queries that would retrieve something relevant if they were constructed differently. Based on learning how humans modify unsuccessful queries, we have successfully developed an automatic approach to turning failed queries into successful ones by removing query terms, while maximally preserving the original user search intent.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Document Processing System
-
批准号:8344939
-
项目类别:
-
资助金额:$7.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8344960
-
项目类别:
-
资助金额:$23.98万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8558105
-
项目类别:
-
资助金额:$56.4万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Natural Language Processing Techniques To Enhance Information Access.
-
批准号:8943224
-
项目类别:
-
资助金额:$56.15万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
-
批准号:7969244
-
项目类别:
-
资助金额:$77.4万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:8149591
-
项目类别:
-
资助金额:$13.71万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8149592
-
项目类别:
-
资助金额:$17.63万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8149602
-
项目类别:
-
资助金额:$47.01万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:9160906
-
项目类别:
-
资助金额:$43.97万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:7969199
-
项目类别:
-
资助金额:$20.27万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8344948
-
项目类别:
-
资助金额:$59.96万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:8344950
-
项目类别:
-
资助金额:$17.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8558117
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8943215
-
项目类别:
-
资助金额:$18.72万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:7969197
-
项目类别:
-
资助金额:$12.9万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:8344938
-
项目类别:
-
资助金额:$7.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:9160916
-
项目类别:
-
资助金额:$14.32万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:9160928
-
项目类别:
-
资助金额:$12.27万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
-
批准号:7735088
-
项目类别:
-
资助金额:$29.31万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:8149604
-
项目类别:
-
资助金额:$19.59万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
海外基金