课题基金 / 基金详情

Automatic Analysis and Annotation of Document Keywords in Biomedical Literature

Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
生物医学文献中文档关键词的自动分析与标注
批准号:
9160928
负责人:
Willy Wilbur
金额:
$12.27万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Willy Wilbur的其他基金

相似基金

相关文献

中文摘要
翻译
1)电子教科书和PubMed中央标引 目前对电子教科书材料的处理涉及若干步骤,这些步骤旨在产生文本中最有意义的短语,用作参考点。第一个任务是找出语法上合理的短语。我们使用基于Brill变换的标记器的一个版本,用C++重写,用于词性标记。这构成了确定语法合理短语的基础。有一个重要的后处理步骤,该步骤删除涉及对上下文的不适当引用的短语(例如,不同的细胞、最终突变)。在找到语法上合理的短语后,我们会尝试删除那些太常见或太普通而没有用处的短语(例如,重要的结果、短时间)。下一步是将短语与在项目生命周期中收集的先前评级的短语进行比较。最后一个阶段是评估课本中某一短语在文章中的重要性。这样的估计是基于短语出现的频率和段落的大小,与整本书中短语的出现频率和书的整体大小相比较。为了改进这样的估计,我们试图考虑代表相同概念的短语或任何短语。为此,我们使用UMLS Metathesaurus,并对这两种方法进行词干处理并将其组合成文本中出现的概念的一致图景。这一处理的结果是每本教科书的短语-书籍部分对的评分列表。这些都是用来指导在书籍中进行一般搜索的反应。当用户输入我们精选列表中的短语时,首先给出的结果是该短语的高评级图书部分。我们现在对PMCentral的文章正文采用类似的索引方案。这使我们能够为每一篇文章提供一个高评级短语的列表,作为搜索者的增强参考点。 2)PubMed中的很大一部分查询是多项查询,PubMed通常将它们作为项的布尔连接进行处理。然而,对PubMed中的查询的分析表明,许多这样的查询是有意义的短语,而不是简单的术语集合。我们已经检查了如果将这些查询解释为短语或查询词的连接词,在检索质量方面是否会有所不同。如果是这样的话,用这种查询进行搜索的最佳方式是什么?为了解决这个问题,我们开发了一种基于机器学习技术的自动检索评估方法,使我们能够评估和比较各种检索结果。我们表明,包含所有搜索词(但不包含短语)的记录类与包含短语的记录类在性质上不同。我们还表明,这种差异是系统性的,这取决于记录中查询词彼此之间的接近程度。基于这些结果,可以为记录建立最佳检索顺序。我们的发现与近距离搜索的研究是一致的。这里对索引的重要见解是,在某些情况下,如果短语的单词出现在文本中,而不是作为短语出现,则短语可能仍然是用于索引文本的适当概念。 3)我们研究了如何根据好的短语的特征来识别它们,例如频率、在文档中出现的倾向和其他数字特性。这些特征允许人们预测哪些短语是高质量的。我们发现这样的预测在研究文本中可能出现的不同类型的术语以及如何从文本中提取本体时很有用。 4)我们发现通过提前停止的正则化随机梯度下降(SGD)是一种非常有效的训练支持向量机(SVM)的方法。我们发现,早期停止可以实现为在恒定的迭代次数后停止,结果与基于保持数据的停止一样好,也与在大数据集上训练支持向量机的更传统的方法一样好。SGD方法要快得多,并且允许人们容易地为所有27,000个网目术语训练分类器。结果优于以前发表的方法。该方法可以作为为PubMed记录或类似于MESH的自动概念分配系统的索引建议的基础。
英文摘要
1) Electronic Textbook and PubMed Central Indexing Current processing of the electronic textbook material involves a number of steps designed to produce the most meaningful phrases in the text to be used as reference points. The first task is to identify grammatically reasonable phrases. We use a version of the Brill transformation based tagger, rewritten in C++, for part-of- speech tagging. This forms the basis for determining grammatically reasonable phrases. There is a significant post processing step that removes phrases that involve inappropriate references to context (e.g., different cells, final mutation). After finding grammatically reasonable phrases we attempt to eliminate those that are too common or generic to be useful (e.g., significant result, short time). The next step is to compare a phrase with previously rated phrases that have been collected over the life of the project. The final stage is to estimate the importance of a phrase in the passage where it is found in a textbook. Such an estimate is based on the frequency of the phrase and the size of the passage compared with the frequency of the phrase throughout the book and the overall size of the book. In order to improve such an estimate we attempt to take account of the phrase or any phrase that represents the same concept. For this purpose we use the UMLS Metathesaurus and also stemming and combine these two approaches into a consistent picture of the concept as it occurs in the text. The result of this processing is a scored list of phrase-book section pairs for each textbook. These are used to guide the response of general searching in the books. When a user types in a phrase that is on our curated list the first results given are the highly rated book sections for that phrase. We are now applying a similar indexing scheme to the text of articles in PMCentral. This allows us to give a list of highly rated phrases for each article as an enhanced reference point for searchers. 2) A significant fraction of queries in PubMed are multiterm queries and PubMed generally handles them as a Boolean conjunction of the terms. However, analysis of queries in PubMed indicates that many such queries are meaningful phrases, rather than simply collections of terms. We have examined whether or not it makes a difference, in terms of retrieval quality, if such queries are interpreted as a phrase or as a conjunction of query terms. And, if it does, what is the optimal way of searching with such queries. To address the question, we developed an automated retrieval evaluation method, based on machine learning techniques, that enables us to evaluate and compare various retrieval outcomes. We show that classes of records that contain all the search terms, but not the phrase, qualitatively differ from the class of records containing the phrase. We also show that the difference is systematic, depending on the proximity of query terms to each other within the record. Based on these results, one can establish the best retrieval order for the records. Our findings are consistent with studies in proximity searching. The important insight here for indexing is that in some cases where the words of a phrase occur in text, but not as the phrase, the phrase may still be an appropriate concept to use in indexing the text. 3) We have studied how good phrases can be recognized by their characteristics, such as frequency, tendency to be repeated in documents where they occur, and other numerical properties. These features allow one to predict which phrases are of high quality. We have found such predictions to be useful in studying different kinds of terms that may appear in text and how an ontoloogy might be extracted from text. 4) We have found stochastic gradient descent (SGD) with regularization by early stopping to be a very efficient method for training a Support Vector Machine (SVM) for MeSH term assignment. We have discovered that the early stopping can be implemented as stopping after a constant number of iterations and the results are as good as stopping based on held out data and also as good as more conventional methods of training an SVM on large data sets. The SGD approach is much faster and allows one to readily train classifiers for all 27,000 MeSH terms. Results are superior to previously published methods. The approach could be the basis of indexing suggestions for PubMed records or for an automatic concept assignment system similar to MeSH.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
General and Semi-supervised Machine Learning Applied to Bioinformatics
  • 批准号:
    8558105
  • 项目类别:
  • 资助金额:
    $56.4万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Natural Language Processing Techniques To Enhance Information Access.
  • 批准号:
    8943224
  • 项目类别:
  • 资助金额:
    $56.15万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
  • 批准号:
    8344960
  • 项目类别:
  • 资助金额:
    $23.98万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
A Document Processing System
  • 批准号:
    8344939
  • 项目类别:
  • 资助金额:
    $7.99万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
海外基金