Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
批准号:
8149607
负责人:
Zhiyong Lu
金额:
$39.17万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
中文摘要
作为一个文档检索系统,PubMed的目标是提供对数百万科学文档的高效访问。为此,它依赖于将PubMed文档的关键字和语义表示与用户查询进行匹配。MEDLINE引文中使用的一种语义表示被称为医学主题标题(MESH)索引词,由国家医学图书馆的专业人类索引员分配。或者,作者在提交文章时提供的作者关键字从作者的角度捕捉文档主题的本质。最后但并非最不重要的一点是,读者对什么词对一篇文章有自己的重要看法,这可能与网状术语或同一文章的作者关键字一致,也可能不一致。
PubMed依靠人工索引器为PubMed文章分配适当的网状索引术语,这是一个非常耗时和劳动的过程。因此,这些条款不能立即用于新文章。事实上,我们的分析表明,一篇PubMed引文平均需要90天以上的时间才能被手动标注为网状术语。作为回应,我们开发了一种机器学习算法,用于自动预测具有一组新特征的网格项。与其他最先进的方法相比,我们的方法取得了明显更好的性能。我们目前正在探索其在实践中协助人工网目精选过程的潜力。
由于网状术语需要人工整理,作者关键字可以从期刊文章中免费获取。我们首次对生物医学论文中的作者关键词进行了研究,描述了作者关键词在生物医学期刊论文中的增长情况,并对作者关键词和网状标引术语进行了比较研究。我们使用了过去研究中的相似性度量来自动评估作者关键词和网状索引术语对之间的相关性。此外,还手动审查了一组300对术语,以评估该指标并描述术语类型之间的关系。结果表明,作者关键词在生物医学论文中的可获得性越来越高,超过60%的作者关键词可以链接到密切相关的标引术语。这项工作的结果对MEDLINE文档索引和网格术语的开发都有意义。
最后,通过对比,我们发现网格词和作者关键词与用户角度的重要词都没有明显的重叠,这促使我们从集体用户的角度来了解哪些特征使文档词变得重要。具体地说,我们应用机器学习来识别用户查询中可能频繁使用的文档关键字。每个词由一组特征表示,这些特征包括不同类型的信息,如语义类型、词性标记、TF-IDF权重和摘要位置。我们既研究了TF-IDF等传统特征,也研究了以前从未在此上下文中探索过的新特征,如命名实体。我们确定了最重要的功能,并使用数月的真实PubMed日志数据对我们的模型进行了评估。我们的结果表明,除了承载较高的TF-IDF权重外,从用户角度来看,重要的词往往是生物医学实体,存在于文章标题中,并在文章摘要中重复出现。这项研究使我们能够自动预测可能出现在导致文档点击的用户查询中的单词。预测单词的相对重要性也可以在根据相关性对文档进行排名方面发挥作用。
英文摘要
As a document retrieval system, PubMed aims at providing efficient access to millions of scientific documents. For this purpose, it relies on matching keywords and semantic representations of PubMed documents to user queries. One type of semantic representation used in MEDLINE citations is known as Medical Subject Heading (MeSH) indexing terms, which are assigned by professional human indexers at the National Library of Medicine. Alternatively, author keywords, provided by authors when submitting an article, capture the essence of the topic of a document from the authors perspective. Last but not least, readers have their own opinions about what words are of importance to an article, which may or may not agree with either MeSH terms or author keywords of the same article.
PubMed relies on human indexers to assign the appropriate MeSH indexing terms to PubMed articles a very time and labor-intensive process. As a result, these terms are not immediately available for new articles. In fact, our analysis shows that on average it takes over 90 days for a PubMed citation to be manually annotated with MeSH terms. In response, we have developed a machine learning algorithm for automatically predicting MeSH terms with a set of novel features. When compared to other state-of-the-art methods, our approach achieved significantly better performance. We are currently exploring its potential for assisting the manual MeSH curation process in practice.
As MeSH terms require human curation, author keywords can be obtained freely from journal articles when they are available. We conducted a first study on author keywords in biomedical articles where we described the growth of author keywords in biomedical journal articles and presented a comparative study of author keywords and MeSH indexing terms. A similarity metric from our past study was used to automatically assess the relatedness between pairs of author keywords and MeSH indexing terms. Furthermore, a set of 300 pairs was manually reviewed to evaluate the metric and characterize the relationships between the term types. Results show that author keywords are increasingly available in biomedical articles and that over 60% of author keywords can be linked to a closely related indexing term. Results of this work have implications in both MEDLINE document indexing and MeSH terminology development.
Finally by comparison, we found neither MeSH terms nor author keywords overlap significantly with the important words from the users point of view, which motivated us to learn what characteristics make document words important from a collective user perspective. Specifically, we applied machine learning to identify document keywords which would likely be used frequently in user queries. Each word was represented by a set of features that included different types of information, such as semantic type, part of speech tag, TF-IDF weight and location in the abstract. We examined both traditional features such as TF-IDF, as well as novel ones such as named entity, which have not been explored before in this context. We identified the most important features and evaluated our model using months of real-world PubMed log data. Our results suggest that, in addition to carrying high TF-IDF weight, important words from the users perspective tend to be biomedical entities, to exist in article titles, and to occur repeatedly in article abstracts. This study enabled us to automatically predict words likely to appear in user queries that lead to document clicks. The relative importance of predicted words can also play a role in ranking documents by relevance.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9362446
-
项目类别:
-
资助金额:$140.39万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:9564626
-
项目类别:
-
资助金额:$160.63万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine Learning and Natural Language Processing for Biomedical Applications
-
批准号:10927050
-
项目类别:
-
资助金额:$387.34万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10007525
-
项目类别:
-
资助金额:$190.14万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8558092
-
项目类别:
-
资助金额:$97.61万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9796762
-
项目类别:
-
资助金额:$225.49万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8344934
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8943212
-
项目类别:
-
资助金额:$20.8万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8943240
-
项目类别:
-
资助金额:$83.19万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:8558091
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:10261222
-
项目类别:
-
资助金额:$166.47万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:9160930
-
项目类别:
-
资助金额:$40.9万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10007518
-
项目类别:
-
资助金额:$213.91万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Machine learning for medical imaging: automated disease diagnosis and prognosis
-
批准号:10927041
-
项目类别:
-
资助金额:$138.33万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Named Entity Recognition and Relationship Extraction in Biomedicine
-
批准号:8344935
-
项目类别:
-
资助金额:$49.97万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
Query Log Analysis for Improving User Access to NCBI Web Services
-
批准号:10261212
-
项目类别:
-
资助金额:$170.11万
-
财政年份:--
-
负责人:Zhiyong Lu
-
依托单位:
海外基金