tmVar: a text mining approach for extracting sequence variants in biomedical literature

tmVar: a text mining approach for extracting sequence variants in biomedical literature
复制标题

DOI:
10.1093/bioinformatics/btt156
复制
发表时间:
2013-06-01
期刊:
影响因子:
5.8
通讯作者:
Lu, Zhiyong
Lu, Zhiyong
中科院分区:
生物学3区
文献类型:
--
作者:
Wei, Chih-Hsuan;Harris, Bethany R.;Lu, Zhiyong

文献摘要

被引文献

相似文献

动机:从文献中挖掘突变信息成为后基因组时代复杂疾病序列变异分析和解释的生物信息学方法的关键部分。它还被用于协助建立与疾病有关的突变数据库。大多数现有的方法是基于规则的,并且集中于有限类型的序列变异,例如蛋白质点突变。因此,扩展其提取范围需要大量的人工工作来检查新实例并开发相应的规则。因此,新的自动化方法是非常需要提取不同种类的突变与高accuracy.Results:在这里,我们报告tmVar,文本挖掘方法的基础上,条件随机场(CRF)提取广泛的序列变异描述在蛋白质,DNA和RNA水平根据人类基因组变异学会制定的标准命名法。通过这样做,我们涵盖了过去研究中未考虑的几种重要类型的突变。使用一种新的CRF标签模型和特征集,我们的方法在我们的语料库(91.4%对78.1%的F-测量)和他们自己的金标准(93.9%对89.4%的F-测量)上都取得了比最先进的方法更高的性能。这些结果表明,tmVar是一种高性能的方法,从生物医学文献中提取突变。
Motivation: Text-mining mutation information from the literature becomes a critical part of the bioinformatics approach for the analysis and interpretation of sequence variations in complex diseases in the post-genomic era. It has also been used for assisting the creation of disease-related mutation databases. Most of existing approaches are rule-based and focus on limited types of sequence variations, such as protein point mutations. Thus, extending their extraction scope requires significant manual efforts in examining new instances and developing corresponding rules. As such, new automatic approaches are greatly needed for extracting different kinds of mutations with high accuracy.Results: Here, we report tmVar, a text-mining approach based on conditional random field (CRF) for extracting a wide range of sequence variants described at protein, DNA and RNA levels according to a standard nomenclature developed by the Human Genome Variation Society. By doing so, we cover several important types of mutations that were not considered in past studies. Using a novel CRF label model and feature set, our method achieves higher performance than a state-of-the-art method on both our corpus (91.4 versus 78.1% in F-measure) and their own gold standard (93.9 versus 89.4% in F-measure). These results suggest that tmVar is a high-performance method for mutation extraction from biomedical literature.