Query Log Analysis for Improving User Access to NCBI Web Services
Query Log Analysis for Improving User Access to NCBI Web Services
批准号:
10261212
负责人:
Zhiyong Lu
金额:
$170.11万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
2019-nCoVAlgorithmsAreaAssisted SuicideAutomationBiologicalBiomedical ResearchCOVID-19COVID-19 pandemicCollectionCommunitiesDataDatabasesDiagnosisDiseaseFloodsGoalsGrowthHumanInformation ServicesInternetLearningLinkMEDLINEMachine LearningManualsMeSH ThesaurusMeasuresMedical ResearchMethodsMilitary PersonnelMineralsMinorMolecular BiologyMorphologyNamesOrganPopulationPreventionProbabilityProcessPubMedPublicationsResearchRetrievalShockSignal TransductionSuggestionSuicideSuicide attemptSuicide preventionSymptomsSystemTextTimeUpdateVocabularyWeightWorkbasecomorbiditycostimprovednavigation aidonline resourcephrasesresearch and developmentspellingtoolvirtualweb services
中文摘要
在过去的十年中,生物信息的在线搜索发展迅速,已成为任何科学发现过程中不可或缺的一部分。今天,如果不依赖NCBI开发和维护的网络资源,几乎不可能进行生物医学领域的研发。事实上,每天都有数百万用户通过ncbi更新的在线PubMed系统搜索生物信息。然而,找到与用户信息需求相关的数据并不总是那么容易。提高我们对不断增长的Entrez用户群体、他们的信息需求以及他们满足这些需求的方式的理解,为改善NCBI提供的信息服务和信息获取提供了机会。不幸的是,SARS-CoV-2和COVID-19大流行的到来导致了前所未有的生物医学研究和传播所学信息的新机会。
英文摘要
Over the last decade, the online search for biological information has progressed rapidly and has become an integral part of any scientific discovery process. Today, it is virtually impossible to conduct R&D in biomedicine without relying on the kind of Web resources developed and maintained by the NCBI. Indeed, each day millions of users search for biological information via NCBIs updated online PubMed system. However, finding data relevant to a users information need is not always easy. Improving our understanding of the growing population of Entrez users, their information needs and the way in which they meet these needs opens opportunities to improve information services and information access provided by NCBI. The unfortunate arrival of SARS-CoV-2 and the COVID-19 pandemic has led to unprecedented focused biomedical research and new opportunities to distribute the information learned.
Tools to aid searching PubMed are query suggestion, expansion, and spelling correction. Dedicated best match algorithms aid navigational queries by ignoring minor errors and aid informational searches using machine learning to combine relevant signals such as article popularity, publication date and type, and query-document relevance score. Additional valuable aids including identifying related articles and author name disambiguation.
Historically, an important part of MedLINE search has been the MeSH terms assigned to each article. These were critical when only titles were available, or there were only a few articles on a given topic. But with the growth of MEDLINE, the cost of humans assigning MeSH terms has become increasingly prohibitive.
As part of a trans-NLM initiative called NLM Labs, we investigated the value of MeSH terms assigned to each article, in contrast to the value of the MeSH vocabulary for synonymy. We measured that a MeSH term assigned to an article only leads to a click for a small percentage of total queries. When a query leads to multiple articles, an article that did not need MeSH for retrieval is more likely to be clicked by a small margin.
Because manual MeSH assignment is costly, there are projects that investigate the use of automation for increasing its efficiency. Assignments using the FullMeSH tool benefit from using the full text of an article, not just the title and abstract. Not only does it use the full text, it uses Learning-to-Rank to weight the different sections of the article in the most valuable manner. It performs several percentage points better than other state of the art methods.
TermVariants is a collection of synonyms developed via morphology and token distribution. To improve these synonyms, we used similarity of word embeddings to produce a probability of synonymy. This allows us to identify pairs of words that while typographically similar, have very different meanings. Examples include mushrooming vs. mushrooms and mineralizer vs. minerals.
When queries in PubMed return a large, incomprehensible set of documents, PDC, a probabilistic distributional clustering algorithm, can group the articles into titled subtopics. For example, the articles returned by the query suicide are grouped into topics such as assisted suicide, attempted suicide, prevention and control of suicide, and military personnel. Related phrases can further clarify the subtopic.
Our work is directly visible in the new PubMed in several ways. One is that the default Best Match sort order used by the new PubMed is based on a Learning to Rank algorithm we recently developed. While author disambiguation has been available for some time, it now also uses ORCID which is available in a growing number of articles. It can distinguish even articles that do not themselves include the ORCID.
SARS-CoV-2 has been a shock to the entire world. The medical research community has responded with a flood of studies on COVID-19 covering prevention, diagnosis, treatment and other areas. LitCovid provides quick, direct access to these articles. The articles are separated by broad area and can be searched for more specific articles.
To understand this collection better, we applied NER and NLP tools, and other tools mentioned above, to provide an overview. We identified bioentities such as diseases, internal body organs, symptoms and co-morbidities. Their relationship to COVID-19 was determined via co-occurrence. We also automatically clustered articles by topic. We then recognized emerging topics and their growth. These tools could be used to more fully understand any collection of articles.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9362446
-
项目类别:
-
资助金额:$140.39万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:9564626
-
项目类别:
-
资助金额:$160.63万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
-
批准号:10927050
-
项目类别:
-
资助金额:$387.34万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10007525
-
项目类别:
-
资助金额:$190.14万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8149607
-
项目类别:
-
资助金额:$39.17万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8558092
-
项目类别:
-
资助金额:$97.61万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9796762
-
项目类别:
-
资助金额:$225.49万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8344934
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8943212
-
项目类别:
-
资助金额:$20.8万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8943240
-
项目类别:
-
资助金额:$83.19万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8558091
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9160930
-
项目类别:
-
资助金额:$40.9万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10261222
-
项目类别:
-
资助金额:$166.47万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine learning for medical imaging: automated disease diagnosis and prognosis
-
批准号:10927041
-
项目类别:
-
资助金额:$138.33万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10007518
-
项目类别:
-
资助金额:$213.91万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8344935
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
海外基金