课题基金 / 基金详情

Automatic Bayesian Methods In Text Retrieval

Automatic Bayesian Methods In Text Retrieval
文本检索中的自动贝叶斯方法
批准号:
8344938
负责人:
Willy Wilbur
金额:
$7.99万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Willy Wilbur的其他基金

相似基金

相关文献

中文摘要
翻译
该项目目前的工作重点是开发改进的贝叶斯分类模型,并开发使用贝叶斯模型进行主动学习的新方法。 1)我们发展了基于术语的主动学习方法,为主动学习提供了一种不同的方法,并表明在许多情况下,它们比简单的不确定性抽样或错误减少抽样更有效。 2)PubMed数据库是一个独特的挑战,因为它的记录非常大,超过1900万条。由于这种规模,很少有机器学习方法可以在合理的周转时间内应用。一种可以有效应用的方法是朴素贝叶斯,但当要区分的不同类别表现出明显的大小差异时,它的性能很差。但对于人们希望在PubMed研究的问题来说,这种不平衡是很常见的。在这种情况下,我们发现可以通过一种主动的学习启发方法来选择比整个训练集小得多的训练集。结果表明,朴素贝叶斯在网格术语分配的文档分类中的性能提高了近200%。这种方法的结果明显优于KNN方法,而且这样定义的最优训练集可以用作更复杂的机器学习方法的训练集,其结果甚至比用朴素贝叶斯方法得到的结果更好。 3)我们使用基于两个泊松分布的概率计算来计算与文档相关的文档,一个用于文档中对文档内容更重要的术语,另一个用于较边缘的术语。这些被组合成基于术语在文档中的相对频率的术语在文档中的重要性的概率估计。该概率估计与术语的全局IDF权重相结合,以说明该术语在计算两个文档之间的相似度时的重要性。从开发这种方法的时候起,我们就知道它工作得很好。在过去的几年里,TREC基因组学轨道上的数据已经变得可用,这使得我们能够通过将其与Robertson和他的同事开发的BM25公式的结果进行比较来测试这种方法。我们发现我们的概率方法有一个很小但在统计上显著的优势。 4)我们目前正在研究一个问题,当一个数据集中出现几个不同类型的文档时,有人想要计算每个文档的相邻文档。在这种情况下,小组内的分数可能会高于小组之间的分数,因此分数不能真实反映相关性。可能会发生这种情况,因为在组内使用的术语可能很常见,但在组外很少使用。我们已经开发了一种贝叶斯方法来测试并删除这种术语。删除过程需要一些小心,我们通过测试要删除的术语来进行测试,以查看它们对生物学有多具体,以及它们对PubMed的用户有多有用。这些测试一起使我们能够成功地删除文档组中的术语并提高评分,以便它可以成功地用于邻居。
英文摘要
Current work on the project is focusing on developing an improved Bayesian classification model and developing new approaches to active learning with a Bayesian model. 1) We have developed term based active learning methods which provide a different approach to active learning and have shown that they are in many cases more effective then simple uncertainty sampling or error reduction sampling. 2) The PubMed database presents a unique challenge because of its very large size of over 19 million records. Because of this size few machine learning methods can be applied with a reasonable turn-around time. One method that can be applied efficiently is Naive Bayes, but it performs poorly when the different classes to be distinguished exhibit a marked size discrepancy. But such an imbalance is common for the problems one wishes to study in PubMed. In such a situation we have discovered that a training set much smaller than the whole set can be selected by an active learning inspired method. The result yields an almost 200% improvement in the performance of Naive Bayes in classifying documents for MeSH term assignment. The results are significantly better than a KNN method and there is the added advantage that the optimal training sets defined in this way can be used as the training sets for more sophisticated machine learning methods with even better results than those obtained from Naive Bayes. 3) We compute the documents related to a document using a probability calculation based on two Poisson distributions, one for the terms in a document that are more central to the documents content and one for the terms that are more peripheral. These are combined into a probability estimate of the importance of a term in a document based on its relative frequency in the document. This probability estimate is combined with the global IDF weight of a term to account for that terms importance in computing the similarity between two documents. We have known from the time this approach was developed that it worked well. In the last several years data has become available in the TREC genomics track that has allowed us to test this approach by comparing it with theresults of the bm25 formula developed by Robertson and colleagues. We find a small but statistically significant advantage for our probabilistic approach. 4) We are currently working on a problem which arises when several different kinds of documents appear in a dataset and one wants to compute neighboring documents for each document. In this situation it is possible that scores will be higher within groups than between groups so that scores do not give a true picture of relatedness. This can happen because within a group terminology may be common that is used but rarely outside the group. We have developed a Bayesian method to test for this kind of terminology and remove it. The removal process requires some care which we exercise by testing terms to be removed by tests to see how specific they are to biology and how useful they are to the users of PubMed. These test together have allowed us to successfully remove terminology within document groups and improve scoring so that it can be used for neighboring successfully.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Document Processing System
  • 批准号:
    8344939
  • 项目类别:
  • 资助金额:
    $7.99万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
  • 批准号:
    8344960
  • 项目类别:
  • 资助金额:
    $23.98万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
  • 批准号:
    8558105
  • 项目类别:
  • 资助金额:
    $56.4万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Natural Language Processing Techniques To Enhance Information Access.
  • 批准号:
    8943224
  • 项目类别:
  • 资助金额:
    $56.15万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
海外基金