Gene Name Extraction Using FlyBase Resources

Gene Name Extraction Using FlyBase Resources
复制标题

DOI:
10.3115/1118958.1118959
复制
发表时间:
2003-07
期刊:
--
影响因子:
--
通讯作者:
Alexander A. Morgan;L. Hirschman;A. Yeh;M. Colosimo
Alexander A. Morgan;L. Hirschman;A. Yeh;M. Colosimo
中科院分区:
其他
文献类型:
--
作者:
Alexander A. Morgan;L. Hirschman;A. Yeh;M. Colosimo

文献摘要

被引文献

相似文献

基于机器学习的实体提取需要大量的标注训练语料库才能达到可接受的结果。然而,相关数据的专家注释的成本,加上注释者之间的可变性问题,使得创建必要的语料库既昂贵又耗时。我们在这里报告了一个简单的方法,用于自动创建大量的生物实体(基因或蛋白质)提取系统的不完美的训练数据。我们使用了FlyBase模式生物数据库中的可用资源;这些资源包括一个精心策划的基因列表和从中提取条目的文章,以及一个同义词词典。我们应用简单的模式匹配来识别相关摘要中的基因名称,并使用文章的精选条目列表过滤这些实体。这个过程创建了一个数据集,可以用来训练一个简单的隐马尔可夫模型(HMM)实体标记器。HMM标记器的结果与其他组报告的结果相当(F测量值为0.75)。这种方法的优点是可以快速转移到具有类似现有资源的新域。
Machine-learning based entity extraction requires a large corpus of annotated training to achieve acceptable results. However, the cost of expert annotation of relevant data, coupled with issues of inter-annotator variability, makes it expensive and time-consuming to create the necessary corpora. We report here on a simple method for the automatic creation of large quantities of imperfect training data for a biological entity (gene or protein) extraction system. We used resources available in the FlyBase model organism database; these resources include a curated lists of genes and the articles from which the entries were drawn, together a synonym lexicon. We applied simple pattern matching to identify gene names in the associated abstracts and filtered these entities using the list of curated entries for the article. This process created a data set that could be used to train a simple Hidden Markov Model (HMM) entity tagger. The results from the HMM tagger were comparable to those reported by other groups (F-measure of 0.75). This method has the advantage of being rapidly transferable to new domains that have similar existing resources.