Literature-based concept profiles for gene annotation: The issue of weighting

Literature-based concept profiles for gene annotation: The issue of weighting
复制标题

DOI:
10.1016/j.ijmedinf.2007.07.004
复制
发表时间:
2008-05-01
影响因子:
4.9
通讯作者:
Kors, Jan A.
Kors, Jan A.
中科院分区:
医学2区
文献类型:
--
作者:
Jelier, Rob;Schuemie, Martijn J.;Kors, Jan A.

文献摘要

被引文献

相似文献

背景资料:文本挖掘已被用于将生物医学概念(如基因或生物过程)相互联系起来,以用于注释目的或生成新的假设。为了将两个概念相互联系起来,一些作者使用了向量空间模型,因为向量可以有效和透明地进行比较。使用这个模型,一个概念的特点是一个列表的相关概念,以及权重,指示的关联强度。向量中的相关概念及其权重是从链接到感兴趣概念的一组文档中导出的。这种方法的一个重要问题是确定相关概念的权重。为确定这些权重提出了各种方案,但没有对不同方法进行比较研究。在这里,我们比较了几种加权方法在一个大规模的分类experiment.Methods:三种不同的技术进行了评估:(1)加权平均的基础上,一种经验的方法;(2)对数似然比,一种基于测试的措施;(3)不确定性系数,一种基于信息理论的措施。加权方案被应用在一个系统中,该系统用基因本体编码注释基因。作为我们研究的金标准,我们使用了基因本体注释项目提供的注释。分类性能进行了评估,通过使用曲线下面积(AUC)作为measure of performance.Results和讨论:所有方法的受试者工作特征(ROC)曲线,中位数AUC得分大于0.84,得分大大高于二进制方法没有任何加权。特别是对于更具体的基因本体代码,观察到优异的性能。当考虑整个实验时,方法之间的差异很小。然而,与一个概念相联系的文件数量被证明是一个重要的变量。当大量的文本可用于生成概念向量时,方法的性能差异很大,不确定性系数则优于其他两种方法。(c)2007爱思唯尔爱尔兰有限公司保留所有权利。
Background: Text-mining has been used to link biomedical concepts, such as genes or biological processes, to each other for annotation purposes or the generation of new hypotheses. To relate two concepts to each other several authors have used the vector space model, as vectors can be compared efficiently and transparently. Using this model, a concept is characterized by a list of associated concepts, together with weights that indicate the strength of the association. The associated concepts in the vectors and their weights are derived from a set of documents linked to the concept of interest. An important issue with this approach is the determination of the weights of the associated concepts. Various schemes have been proposed to determine these weights, but no comparative studies of the different approaches are available. Here we compare several weighting approaches in a large scale classification experiment.Methods: Three different techniques were evaluated: (1) weighting based on averaging, an empirical approach; (2) the log likelihood ratio, a test-based measure; (3) the uncertainty coefficient, an information-theory based measure. The weighting schemes were applied in a system that annotates genes with Gene Ontology codes. As the gold standard for our study we used the annotations provided by the Gene Ontology Annotation project. Classification performance was evaluated by means of the receiver operating characteristics (ROC) curve using the area under the curve (AUC) as the measure of performance.Results and discussion: All methods performed well with median AUC scores greater than 0.84, and scored considerably higher than a binary approach without any weighting. Especially for the more specific Gene Ontology codes excellent performance was observed. The differences between the methods were small when considering the whole experiment. However, the number of documents that were linked to a concept proved to be an important variable. When larger amounts of texts were available for the generation of the concepts' vectors, the performance of the methods diverged considerably, with the uncertainty coefficient then outperforming the two other methods. (c) 2007 Elsevier Ireland Ltd. All rights reserved.