Query Log Analysis for Improving User Access to NCBI Web Services
Query Log Analysis for Improving User Access to NCBI Web Services
批准号:
10261212
负责人:
Zhiyong Lu
金额:
$170.11万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
2019-nCoVAlgorithmsAreaAssisted SuicideAutomationBiologicalBiomedical ResearchCOVID-19COVID-19 pandemicCollectionCommunitiesDataDatabasesDiagnosisDiseaseFloodsGoalsGrowthHumanInformation ServicesInternetLearningLinkMEDLINEMachine LearningManualsMeSH ThesaurusMeasuresMedical ResearchMethodsMilitary PersonnelMineralsMinorMolecular BiologyMorphologyNamesOrganPopulationPreventionProbabilityProcessPubMedPublicationsResearchRetrievalShockSignal TransductionSuggestionSuicideSuicide attemptSuicide preventionSymptomsSystemTextTimeUpdateVocabularyWeightWorkbasecomorbiditycostimprovednavigation aidonline resourcephrasesresearch and developmentspellingtoolvirtualweb services
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Over the last decade, the online search for biological information has progressed rapidly and has become an integral part of any scientific discovery process. Today, it is virtually impossible to conduct R&D in biomedicine without relying on the kind of Web resources developed and maintained by the NCBI. Indeed, each day millions of users search for biological information via NCBIs updated online PubMed system. However, finding data relevant to a users information need is not always easy. Improving our understanding of the growing population of Entrez users, their information needs and the way in which they meet these needs opens opportunities to improve information services and information access provided by NCBI. The unfortunate arrival of SARS-CoV-2 and the COVID-19 pandemic has led to unprecedented focused biomedical research and new opportunities to distribute the information learned.
Tools to aid searching PubMed are query suggestion, expansion, and spelling correction. Dedicated best match algorithms aid navigational queries by ignoring minor errors and aid informational searches using machine learning to combine relevant signals such as article popularity, publication date and type, and query-document relevance score. Additional valuable aids including identifying related articles and author name disambiguation.
Historically, an important part of MedLINE search has been the MeSH terms assigned to each article. These were critical when only titles were available, or there were only a few articles on a given topic. But with the growth of MEDLINE, the cost of humans assigning MeSH terms has become increasingly prohibitive.
As part of a trans-NLM initiative called NLM Labs, we investigated the value of MeSH terms assigned to each article, in contrast to the value of the MeSH vocabulary for synonymy. We measured that a MeSH term assigned to an article only leads to a click for a small percentage of total queries. When a query leads to multiple articles, an article that did not need MeSH for retrieval is more likely to be clicked by a small margin.
Because manual MeSH assignment is costly, there are projects that investigate the use of automation for increasing its efficiency. Assignments using the FullMeSH tool benefit from using the full text of an article, not just the title and abstract. Not only does it use the full text, it uses Learning-to-Rank to weight the different sections of the article in the most valuable manner. It performs several percentage points better than other state of the art methods.
TermVariants is a collection of synonyms developed via morphology and token distribution. To improve these synonyms, we used similarity of word embeddings to produce a probability of synonymy. This allows us to identify pairs of words that while typographically similar, have very different meanings. Examples include mushrooming vs. mushrooms and mineralizer vs. minerals.
When queries in PubMed return a large, incomprehensible set of documents, PDC, a probabilistic distributional clustering algorithm, can group the articles into titled subtopics. For example, the articles returned by the query suicide are grouped into topics such as assisted suicide, attempted suicide, prevention and control of suicide, and military personnel. Related phrases can further clarify the subtopic.
Our work is directly visible in the new PubMed in several ways. One is that the default Best Match sort order used by the new PubMed is based on a Learning to Rank algorithm we recently developed. While author disambiguation has been available for some time, it now also uses ORCID which is available in a growing number of articles. It can distinguish even articles that do not themselves include the ORCID.
SARS-CoV-2 has been a shock to the entire world. The medical research community has responded with a flood of studies on COVID-19 covering prevention, diagnosis, treatment and other areas. LitCovid provides quick, direct access to these articles. The articles are separated by broad area and can be searched for more specific articles.
To understand this collection better, we applied NER and NLP tools, and other tools mentioned above, to provide an overview. We identified bioentities such as diseases, internal body organs, symptoms and co-morbidities. Their relationship to COVID-19 was determined via co-occurrence. We also automatically clustered articles by topic. We then recognized emerging topics and their growth. These tools could be used to more fully understand any collection of articles.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9362446
-
项目类别:
-
资助金额:$140.39万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:9564626
-
项目类别:
-
资助金额:$160.63万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
-
批准号:10927050
-
项目类别:
-
资助金额:$387.34万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10007525
-
项目类别:
-
资助金额:$190.14万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8149607
-
项目类别:
-
资助金额:$39.17万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9796762
-
项目类别:
-
资助金额:$225.49万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8558092
-
项目类别:
-
资助金额:$97.61万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8344934
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8943212
-
项目类别:
-
资助金额:$20.8万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8943240
-
项目类别:
-
资助金额:$83.19万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8558091
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10261222
-
项目类别:
-
资助金额:$166.47万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9160930
-
项目类别:
-
资助金额:$40.9万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10007518
-
项目类别:
-
资助金额:$213.91万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine learning for medical imaging: automated disease diagnosis and prognosis
-
批准号:10927041
-
项目类别:
-
资助金额:$138.33万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8344935
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
海外基金