课题基金 / 基金详情

Query Log Analysis for Improving User Access to NCBI Web Services

Query Log Analysis for Improving User Access to NCBI Web Services
用于改善用户对 NCBI Web 服务的访问的查询日志分析
批准号:
10007518
负责人:
Zhiyong Lu
金额:
$213.91万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Zhiyong Lu的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年里,对生物信息的在线搜索发展迅速,已成为任何科学发现过程中不可或缺的一部分。今天,如果不依赖NCBI开发和维护的那种网络资源,几乎不可能进行生物医学的研发。事实上,每天都有数百万用户通过NCBIS在线Entrez系统搜索生物信息。然而,在Entrez中查找与用户信息需求相关的数据并不总是很容易。提高我们对Entrez用户日益增长的人口、他们的信息需求以及他们满足这些需求的方式的了解,为改进NCBI提供的信息服务和信息获取提供了机会。 在所有Entrez数据库中,PubMed是使用最多的数据库,经常作为人们访问其他Entrez数据库中相关数据的入口点。帮助搜索PubMed的工具有查询建议、扩展和拼写更正。专用的最佳匹配算法通过忽略小错误来帮助导航查询,并使用机器学习来帮助信息搜索,以组合相关信号,如文章流行度、发布日期和类型以及查询-文档相关性分数。其他有价值的帮助包括识别相关文章和消除作者姓名的歧义。 PubMed Labs提供了一个试验和改进新搜索功能的地方。它的特点是专为小屏幕设备量身定做的干净和移动友好的设计,以及一个供用户提供反馈指导未来工作的平台。 虽然搜索通常集中在完整的文档或参考文献上,但句子搜索的价值正在上升。它可以确定具体的陈述,而不是关于一般主题的整篇文章。我们的新工具LitSense提供句子级别的搜索,在句子级别理解生物医学文献。 句子相似性的一种特定用途是帮助保守域数据库(CDD)中的精选工作。为此,LitSense已用于在已用于创建CDD摘要的PubMed文章中查找句子,并识别与现有CDD摘要密切相关的新句子。 对于深度学习任务中的句子使用,BioSentVec是第一个专门的句子编码器 在生物医学领域。它比一般的领域编码器更好地捕获生物医学语义。 当然,单词嵌入仍然是在NLP任务中使用深度学习的主要方法。BioWordVec使用子单词信息和网格生成生物医学单词嵌入,可以显著提高性能。生物医学术语通常包含重要的子词信息。在诸如网格之类的本体中可用的语义信息是有意义的。一般的单词嵌入不能利用这些有价值的补充信息。 这些机器学习方法受益于拥有大量可用文本。为此,Web API提供PMC开放获取子集的Bioc版本和作者手稿。这是对我们现有的FTP服务的一个不断更新的补充。这些文档以JSON或XML格式提供,ASCII和Unicode编码均可用。
英文摘要
Over the last decade, the online search for biological information has progressed rapidly and has become an integral part of any scientific discovery process. Today, it is virtually impossible to conduct R&D in biomedicine without relying on the kind of Web resources developed and maintained by the NCBI. Indeed, each day millions of users search for biological information via NCBIs online Entrez system. However, finding data relevant to a users information need is not always easy in Entrez. Improving our understanding of the growing population of Entrez users, their information needs and the way in which they meet these needs opens opportunities to improve information services and information access provided by NCBI. Among all Entrez databases, PubMed is the most used one and often serves as an entry point for people to access related data in other Entrez databases. Tools to aid searching PubMed are query suggestion, expansion, and spelling correction. Dedicated best match algorithms aid navigational queries by ignoring minor errors and aid informational searches using machine learning to combine relevant signals such as article popularity, publication date and type, and query-document relevance score. Additional valuable aids including identifying related articles and author name disambiguation. PubMed Labs provides a place to trial and improve new search features. It features a clean and mobile-friendly design tailored specifically towards small screen devices and a platform for users to provide feedback guiding future work. While search has usually focused on full documents or references, the value of sentence search is rising. It can identify specific statements rather than whole articles on a general topic. Our new tool, LitSense, provides sentence level search, making sense of biomedical literature at sentence level. A specific use of sentence similarity is to aid the curation efforts in the Conserved Domain Database (CDD). To this end, LitSense has been used to both finds sentences in PubMed articles already used to create CDD summaries and identify new sentences closely related to existing CDD summaries. For using sentences in Deep Learning tasks, BioSentVec is the first sentence encoder specifically for the biomedical domain. It better captures biomedical semantics than general domain encoders. Of course, word embeddings remain the primary method of using Deep Learning in NLP tasks. BioWordVec uses subword information and MeSH to generate biomedical word embeddings that can significantly improve performance. Biomedical terminology often includes important subword information. The semantic information available in ontologies such as MeSH is meaningful. A generic word embedding cannot take advantage of this valuable supplemental information. These machine learning methods benefit from having a large amount of text available. To that end a Web API serves BioC versions of the PMC Open Access Subset and Author Manuscripts. This is a continuously updated complement to our existing FTP service. The documents are available in either JSON or XML and both ASCII and Unicode encodings are available.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    9362446
  • 项目类别:
  • 资助金额:
    $140.39万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
  • 批准号:
    9564626
  • 项目类别:
  • 资助金额:
    $160.63万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
  • 批准号:
    10927050
  • 项目类别:
  • 资助金额:
    $387.34万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
  • 批准号:
    10007525
  • 项目类别:
  • 资助金额:
    $190.14万
  • 财政年份:
    --
  • 负责人:
    Zhiyong Lu
  • 依托单位:
海外基金