Identification of transcription factor contexts in literature using machine learning approaches.

Identification of transcription factor contexts in literature using machine learning approaches.
复制标题

DOI:
10.1186/1471-2105-9-s3-s11
复制
发表时间:
2008-04-11
期刊:
影响因子:
3
通讯作者:
Keane JA
Keane JA
中科院分区:
生物学4区
文献类型:
--
作者:
Yang H;Nenadic G;Keane JA

文献摘要

被引文献

相似文献

关于转录因子(TF)的信息的可用性对于基因组生物学至关重要,因为TF在基因表达的调控中起着核心作用。虽然手动文献管理是昂贵的,劳动密集型的,半自动化的文本挖掘支持的发展受到训练数据不可用的阻碍。目前还没有关于如何使用现有数据源(例如来自MeSH词库和GO本体的TF相关数据)或潜在噪声示例数据(例如蛋白质-蛋白质相互作用,PPI)来提供用于识别文献中TF上下文的训练数据的研究。在本文中,我们描述了一个文本分类系统,旨在自动识别相关的转录因子在文献中的上下文。学习模型基于一组被认为与任务相关的生物学特征(例如蛋白质和基因名称、相互作用词、其他生物学术语)。我们利用现有生物资源(MeSH和GO)的背景知识来设计这些功能。从MeSH和GO中TF相关概念的描述,PPI数据和代表非蛋白质功能描述的数据中收集了弱和噪声训练数据集。研究了三种机器学习方法,沿着基于投票的各个方法和/或不同训练数据集的合并。该系统取得了非常令人鼓舞的结果,大多数分类器实现了90%以上的F-措施。实验结果表明,该模型可以用于识别TF相关的上下文(即句子)具有较高的准确性,与传统的词袋方法相比,具有显着减少的功能集。考虑现有PPI数据的结果表明,TF和PPI上下文之间的相似性并不像我们预期的那样高。我们还表明,现有的知识来源是有用的特征工程和获得嘈杂的积极的训练数据。
Availability of information about transcription factors (TFs) is crucial for genome biology, as TFs play a central role in the regulation of gene expression. While manual literature curation is expensive and labour intensive, the development of semi-automated text mining support is hindered by unavailability of training data. There have been no studies on how existing data sources (e.g. TF-related data from the MeSH thesaurus and GO ontology) or potentially noisy example data (e.g. protein-protein interaction, PPI) could be used to provide training data for identification of TF-contexts in literature. In this paper we describe a text-classification system designed to automatically recognise contexts related to transcription factors in literature. A learning model is based on a set of biological features (e.g. protein and gene names, interaction words, other biological terms) that are deemed relevant for the task. We have exploited background knowledge from existing biological resources (MeSH and GO) to engineer such features. Weak and noisy training datasets have been collected from descriptions of TF-related concepts in MeSH and GO, PPI data and data representing non-protein-function descriptions. Three machine-learning methods are investigated, along with a vote-based merging of individual approaches and/or different training datasets. The system achieved highly encouraging results, with most classifiers achieving an F-measure above 90%. The experimental results have shown that the proposed model can be used for identification of TF-related contexts (i.e. sentences) with high accuracy, with a significantly reduced set of features when compared to traditional bag-of-words approach. The results of considering existing PPI data suggest that there is not as high similarity between TF and PPI contexts as we have expected. We have also shown that existing knowledge sources are useful both for feature engineering and for obtaining noisy positive training data.