课题基金 / 基金详情

Free Text Gene Name Recognition

Free Text Gene Name Recognition
自由文本基因名称识别
批准号:
8344950
负责人:
Willy Wilbur
金额:
$17.99万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Willy Wilbur的其他基金

相似基金

相关文献

中文摘要
翻译
1)我们已经确信,可以使用MEDLINE中句子中可能出现的不同类型实体的更多信息来提高名称识别。这使我们设计了一组语义类别,并试图用可以从数据库和网站中获得的实际名称来填充这些类别。我们称之为SEMCAT。它目前可以识别75个类别,包含分布在这些类别中的大约400万个名称字符串。我们尝试了概率上下文无关文法和文本字符串的马尔可夫模型,试图学习如何识别不同类别的实体。 然而,我们发现区分基因/蛋白质类别和非基因/蛋白质类别的最佳方法是一种新的算法,我们称之为优先级模型。 在SEMCAT中,与任何名称相关联的每个标记都与两个概率相关联。第一概率是标记指示它是基因/蛋白质名称的一部分的概率,第二概率是标记作为指示符的可靠性的指示符。使用该模型,给定一个短语,可以计算该短语是基因/蛋白质名称的概率的估计。我们发现,与优先级模型,我们可以实现96%的F分数相比,我们最好的PCFG方法为95%。(with Lorrie田边)。 BioCreative II中基因提及识别的最佳性能是由IBM的Rie Ando提出的,他引入了一种称为交替结构优化的技术。这种方法处理了许多类似于命名实体标记的标记问题,但只是试图从周围的文本上下文预测名称或标记的出现。当已经学习了这些许多辅助问题的SVM解决方案权重向量时,执行奇异值分解并从每个向量中减去分解中的前h个分量。该减法仅用于减少成本函数的正则化项中的惩罚。然后重新学习权重向量,并重复该过程。这一直持续到收敛。最后的结果是一组h分量的分解的许多权重向量。使用这些组件来增强对实际命名实体识别任务的学习。这是一个有点复杂和难以使用。我们正在研究如何使用类似的方法,但使用一种更简单的方法来应用辅助学习来提高命名实体识别。一个问题是如何将这种辅助学习与SEMCAT数据联合收割机相结合。我们目前正在努力改进这个模型,找到一种方法,将其应用于两个以上的类。 2)我们最近共同主持了BioCreative III研讨会,其中主要的竞争任务是在全文文章中找到基因提及并将其映射到其GenBank标识符并对其可靠性进行评分,将PubMed记录分类为可能代表包含蛋白质-蛋白质相互作用信息的文章,并在全文中找到描述实验者用于实验验证蛋白质-蛋白质相互作用的方法的文本。我们组织了第一项任务,并参加了第二项任务。在第二个任务中,我们使用优先级模型来定位蛋白质提及,它被证明是非常成功的,与其他方法相比具有竞争力。 3)我们目前正在努力开发更通用的方法,根据他们的摘要为PPI寻找高价值的文章。这项工作不仅涉及更强大的排名方法,还涉及向用户显示证据以供用户快速评估的方法。
英文摘要
1) We have become convinced that more information about the different types of entities that can occur in sentences in MEDLINE can be used to improve name recognition. This has led us to design a set of semantic categories and to attempt to fill these categories with actual names that can be harvested from databases and from web sites. We call the result SEMCAT. It currently recognizes seventy-five categories and contains about four million name strings distributed over those categories. We have experimented with probabilistic context free grammars and Markov models of text strings in an attempt to learn how to recognize the entities in different categories. However, the best approach we have found for distinguishing the categories of gene/protein and not gene/protein is a new algorithm we term a priority model. Every token associated with any name in SEMCAT has associated with it two probabilities. The first probability is the probability that the token indicates that it is part of a gene/protein name and the second probability is an indicator of how reliable the token is as an indicator. With this model, given a phrase, one can compute an estimate of the probability that the phase is a gene/protein name. We find that with the priority model we can achieve an F score of 96% as compared with 95% for our best PCFG approach. (with Lorrie Tanabe). The top performance for gene mention recognition in BioCreative II was by Rie Ando from IBM who introduced a technique called alternating structural optimization. This approach takes many labeling problems similar to named entity tagging, but simply tries to predict the occurrence of the names or the tokens from the surrounding textual context. When the SVM solution weight vectors for these many auxiliary problems have been learned, one performs a singular value decomposition and subtracts from each vector its first h components in the decomposition. This subtraction is only used to decrease the penalty in the regularization term of the cost function. The weight vectors are then relearned and the process is repeated. This is continued until convergence. The final result is a set of h components of the decomposition of the many weight vectors. One uses these components to enhance the learning on the actual named entity recognition task. This is a bit complicated and difficult to use. We are studying how we may be able to use a similar approach, but with a simpler method of applying the auxiliary learning to improve named entity recognition. One problem is how to combine such auxiliary learning with the SEMCAT data. We are currently working to improve this model by finding a way to apply it to more than two classes at a time. 2)We recently co-chaired the BioCreative III Workshop in which the main competitive tasks were to find gene mentions in a full text article and map them to their GenBank identifiers and score them as to reliability, to classify PubMed records as likely to represent articles containing information on protein-protein interactions, and to find the text in full papers that describes the method used by an experimenter to experimentally verify a protein-protein interaction. We organized the first of these task and participated in the second. In the second task we used the priority model to locate protein mentions and it proved very successful and competitive with other approaches. 3) We are currently working to develop more general methods of finding high value articles for PPI based on their abstracts. This effort involves not only more powerful ranking methods, but also ways to display evidence to the user for a users quick evaluation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Document Processing System
  • 批准号:
    8344939
  • 项目类别:
  • 资助金额:
    $7.99万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
  • 批准号:
    8344960
  • 项目类别:
  • 资助金额:
    $23.98万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
  • 批准号:
    8558105
  • 项目类别:
  • 资助金额:
    $56.4万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Natural Language Processing Techniques To Enhance Information Access.
  • 批准号:
    8943224
  • 项目类别:
  • 资助金额:
    $56.15万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
国内基金
海外基金
Journal of Integrative Plant Biology
  • 批准号:
    31024801
  • 项目类别:
    专项基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2010
  • 负责人:
    贺萍
  • 依托单位: