课题基金 / 基金详情

Query Log Analysis for Improving User Access to NCBI Web Services

Query Log Analysis for Improving User Access to NCBI Web Services
用于改善用户对 NCBI Web 服务的访问的查询日志分析
批准号:
8344934
负责人:
Zhiyong Lu
金额:
$49.97万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年里,对生物信息的在线搜索发展迅速,已成为任何科学发现过程中不可或缺的一部分。今天,如果不依赖NCBI开发和维护的那种网络资源,几乎不可能进行生物医学的研发。事实上,每天都有数百万用户通过NCBIS在线Entrez系统搜索生物信息。然而,在Entrez中查找与用户信息需求相关的数据并不总是很容易。提高我们对Entrez用户日益增长的人口、他们的信息需求以及他们满足这些需求的方式的了解,为改进NCBI提供的信息服务和信息获取提供了机会。了解和描述搜索引擎用户特征的一个资源是交易日志。我们之前对PubMed查询日志的研究使我们开发和部署了几个有用的应用程序来帮助用户进行搜索和检索,例如PubMed中的查询公式,即相关查询和查询自动补全。受其成功的启发,我们继续使用日志分析来确定与NCBI操作密切相关的研究问题。 在所有Entrez数据库中,PubMed是使用最多的数据库,经常作为人们访问其他Entrez数据库中相关数据的入口点。在最近的一次调查中,我们将PubMed与其他研究人员开发的其他类似文献搜索工具进行了比较和对比。根据我们的调查,我们发现PubMed可以在某些领域向他人学习,以提高检索和用户搜索体验。例如,有几个工具与PubMed的不同之处在于它们允许相关性搜索,这是一个重要的功能,可以帮助一些PubMed搜索。关于用户界面,其他工具已经尝试使用诸如集群、词云或网络之类的新颖方案来可视化搜索结果。尽管这些方法没有在大规模的用户研究中得到正式验证,但更好地可视化搜索结果的概念可能仍然有助于考虑改进PubMeds当前基于列表的呈现方式。 2011年,我们还研究了PubMed以外的查询日志。一个具体的项目涉及分析NCBIS Global Search的用户日志,其中根据所有Entrez数据库搜索用户查询,并显示结果,但没有说明不同数据库与用户查询的相关性。因此,我们的任务是根据用户的输入查询,预测哪个Entrez数据库(S)最有可能包含用户的相关信息。在我们当前的方法中,我们首先从日志中收集数据语料库,其中每个数据点包含一个用户查询,然后是用户对特定数据库的点击。接下来,我们应用机器学习算法来学习用户查询中的特征,以区分用户寻找不同生物数据的意图。基于学习到的特征,我们对新输入的查询进行分类,并将用户引导到相关数据库(S)中的结果以满足他们的搜索需求。 查询日志的另一个用途是我们为PubMed Health所做的工作:这是一项新推出的NCBI服务,为健康消费者和医疗保健专业人员提供关于疾病、状况、药物、治疗方案和健康生活的最新信息。根据我们在PubMed Health中对疾病和药物主题的实际使用情况的日志分析,我们发现大约80%的使用情况落在数据库内容的20%上。也就是说,用户访问模式满足帕累托原则(也称为80-20规则),这可能会对进一步改进我们的Web服务产生很多影响。例如,该原则建议我们将资源优先放在那些访问频繁的内容上。此外,当我们在PubMed Health中部署我们的研究,以构建链接以丰富相关药物和疾病页面之间的用户访问时,还使用了查询日志。具体地说,我们开发了文本挖掘方法来自动识别药物及其密切相关的疾病(例如立普妥和心脏病)。特别是,我们利用用户查询中药物和疾病提及的共现信息来帮助确定它们在用户需求中的相关度和受欢迎程度。结果,我们计算了数千对药物和疾病,这些药物和疾病不仅密切相关,而且经常被用户要求。
英文摘要
Over the last decade, the online search for biological information has progressed rapidly and has become an integral part of any scientific discovery process. Today, it is virtually impossible to conduct R&D in biomedicine without relying on the kind of Web resources developed and maintained by the NCBI. Indeed, each day millions of users search for biological information via NCBIs online Entrez system. However, finding data relevant to a users information need is not always easy in Entrez. Improving our understanding of the growing population of Entrez users, their information needs and the way in which they meet these needs opens opportunities to improve information services and information access provided by NCBI. One resource for understanding and characterizing patrons of search engines is the transaction logs. Our previous investigation of PubMed query logs has led us to develop and deploy several useful applications in assisting user searches and retrieval such as the query formulation in PubMed, namely Related Queries and Query Autocomplete. Inspired by its success, we have continued using log analysis to identify research problems which are closely related to NCBI operations. Among all Entrez databases, PubMed is the most used one and often serves as an entry point for people to access related data in other Entrez databases. In a recent survey, we compared and contrasted PubMed with other similar literature search tools developed by other researchers. Based on our investigation, we found that there are areas where PubMed may learn from others for self-improvement with respect to better retrieval and user search experience. For instance, several tools differ from PubMed in that they allow relevance search, an important feature that can be helpful for some PubMed searches. With respect to user interface, other tools have attempted to visualize search results using novel schemes such as clusters, word clouds or networks. Though these methods are not formally validated in large-scale user studies, the concept of better visualization of search results might still be useful for consideration towards improving PubMeds current list-based presentation. In 2011, we have also studied query logs beyond PubMed. One specific project involves the analysis of user logs of NCBIs Global Search where user queries are searched against all Entrez databases and results are presented without indicating the relevancy of different databases to the user queries. Hence our task is to predict which Entrez database(s) is mostly likely to contain the relevant information to the users based on their input queries. In our current approach we first collect a data corpus from logs where each data point contains a user query followed by a user click to a specific database. Next, we apply machine-learning algorithms to learn the characteristics in user queries that distinguish users intention for seeking different biological data. Based on the learned features, we classify new input queries and direct users to results in relevant database(s) for their search needs. Another use of query logs lies in our work for PubMed Health: a newly launched NCBI service offering up-to-date information on diseases, conditions, drugs, treatment options, and healthy living for both health consumers and healthcare professionals. Based on our log analysis of actual usage on disease and drug topics in PubMed Health, we discovered that approximately 80% of the usage falls on 20% of the database content. That is, the user access pattern satisfies the Pareto principle (aka the 80-20 rule), which can have many implications for further improving our Web service. For instance, the principle suggests that we prioritize our resources on those heavily accessed content. In addition, query logs were used when we deployed our research on building links for enriching user access between related drug and disease pages as portlets in PubMed Health. Specifically, we developed text-mining methods for automatically identifying drugs and its closely related diseases (e.g. Lipitor and heart disease). In particular, we took advantage of the co-occurrence information of drug and disease mentions in user queries to help determine their strength in relatedness and popularity in user needs. As a result, we computed several thousand pairs of drug and diseases that are not only closely related but also frequently requested by our users.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    9362446
  • 项目类别:
  • 资助金额:
    $140.39万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
  • 批准号:
    9564626
  • 项目类别:
  • 资助金额:
    $160.63万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
  • 批准号:
    10927050
  • 项目类别:
  • 资助金额:
    $387.34万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    10007525
  • 项目类别:
  • 资助金额:
    $190.14万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: