Mutation mining - A prospector's tale

Mutation mining - A prospector's tale
复制标题

DOI:
10.1007/s10796-006-6103-2
复制
发表时间:
2006-02-01
影响因子:
5.9
通讯作者:
Witte, R
Witte, R
中科院分区:
计算机科学3区
文献类型:
--
作者:
Baker, CJO;Witte, R

文献摘要

被引文献

相似文献

蛋白质结构可视化工具呈现图像,允许用户探索蛋白质的结构特征。然而,与特定蛋白质或蛋白质家族相关的背景特异性信息不容易整合,并且必须从数据库上传或通过输入文件的人工管理提供。蛋白质工程师花费相当多的时间迭代地审查文献和用突变残基手动注释的蛋白质结构可视化。与此同时,文本挖掘工具越来越多地用于从科学文献中提取特定的原始文本单元,并已证明有潜力支持蛋白质工程师的活动。将突变特定的原始文本注释转移到蛋白质结构需要集成的数据处理管道,可以协调信息检索,信息提取,蛋白质序列检索,序列比对和突变残基映射。我们描述了为此目的而设计的Mutation Miner管道,并介绍了该过程中关键步骤的案例研究评估。从蛋白质家族突变的文献开始,卤代烷脱卤酶,联苯双加氧酶和木聚糖酶,我们列举了可用于文本挖掘分析的相关文件,可用的电子格式,以及给定蛋白质家族的突变数量。我们回顾了NLP驱动的蛋白质序列检索的效率,并报告了Mutation Miner在将注释映射到蛋白质结构可视化中的有效性。我们强调的可行性和实用性的方法。
Protein structure visualization tools render images that allow the user to explore structural features of a protein. Context specific information relating to a particular protein or protein family is, however, not easily integrated and must be uploaded from databases or provided through manual curation of input files. Protein Engineers spend considerable time iteratively reviewing both literature and protein structure visualizations manually annotated with mutated residues. Meanwhile, text mining tools are increasingly used to extract specific units of raw text from scientific literature and have demonstrated the potential to support the activities of Protein Engineers.The transfer of mutation specific raw-text annotations to protein structures requires integrated data processing pipelines that can co-ordinate information retrieval, information extraction, protein sequence retrieval, sequence alignment and mutant residue mapping. We describe the Mutation Miner pipeline designed for this purpose and present case study evaluations of the key steps in the process. Starting with literature about mutations made to protein families; haloalkane dehalogenase, bi-phenyl dioxygenase, and xylanase we enumerate relevant documents available for text mining analysis, the available electronic formats, and the number of mutations made to a given protein family. We review the efficiency of NLP driven protein sequence retrieval from databases and report on the effectiveness of Mutation Miner in mapping annotations to protein structure visualizations. We highlight the feasibility and practicability of the approach.