Modeling and prediction of genome sequence information by using information representation models
Modeling and prediction of genome sequence information by using information representation models
批准号:
12208010
负责人:
YADA Tetsushi
金额:
$46.34万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research on Priority Areas
财政年份:
2000
资助国家:
日本
项目状态:
已结题
起止时间:
2000 至 2004
中文摘要
在这项研究中,我们的重点是能够从基因组序列中找到基因的基因模型。首先,我们开发了一种通用算法,通过组合多个现有的基因寻找器来寻找基因。该算法已实现到一个名为DIGIT的新型基因查找器中。算法的概要如下。首先,将现有的基因查找器应用于未表征的基因组序列(输入序列)。接下来,DIGIT根据基因寻找者的结果生成所有可能的外显子,并为它们分配外显子类型、阅读框和外显子得分。最后,DIGIT搜索一组外显子,这些外显子的附加分数在其阅读框约束下最大。使用贝叶斯方法推断外显子分数,使用隐马尔可夫模型(HMM)搜索外显子集。我们结合FGENESH、GENSCAN和HMMgene的结果设计了DIGIT,并使用最近编制的基准数据集对其预测精度进行了评估。对于所有数据集,…More DIGIT成功地丢弃了许多由单个基因查找器预测的假阳性外显子,并且与任何单个基因查找器获得的最佳基因水平准确性相比,在基因水平上取得了显著的灵敏度和特异性提高。其次,我们开发了一种新的索引,可以精确地从跨物种基因组比对中提取蛋白质编码区域。该索引与在编码序列比对中观察到的帧恢复密切相关,即如果核苷酸的插入或删除导致编码区域的帧移位,则在附近通常会观察到其他恢复阅读帧的in-del。相比之下,在其他保守区域中没有观察到这种帧恢复。建立了两种基因模型:一种是利用序列相似度和内在基因测度寻找基因的模型(基本模型),另一种是利用序列相似度和内在基因测度同时利用帧恢复指标寻找基因的模型(帧恢复模型)。我们对两种模型的预测精度进行了评估,我们的基准测试表明,帧恢复模型比基本模型显著提高了预测精度。第三,我们开发了GeneDecoder,这是一种基于hmm的真核生物基因发现技术。该算法采用动态规划方法和经注释基因组序列训练的统计模型,将输入的核酸序列划分为若干有意义的片段。此外,GeneDecoder还具有以下特点:(1)多流架构;(2)结合相似性搜索;(3)支持向量机驱动的假设剪接位点筛选。(1)除核酸序列外,GeneDecoder还允许添加任何其他数据流。通常,双齿子双元图值可以提前计算并在“直接”流上对齐,这使得状态转换网络简单得多。预先提取的任何其他有意义的特征都可以合并到。使用这个方案的基因寻找过程。(2)将编码势计算和相似性搜索与已知序列数据库相结合,实现更可靠的推定外显子。为此,GeneDecoder能够将已知的motif模型嵌入到外显子模型中,并使用BLAST搜索发现的与已知序列相似的片段。(3)支持向量机(Support Vector Machine, SVM)是目前已知分类能力较强的模式再识别技术之一,已成功应用于剪接位点预测。在GeneDecoder中,这个功能和基于pwm的剪接站点模型一样被实现。在解析过程中,排除了从基于pwm的模型中得到的假设剪接位点,但这些剪接位点被设计为剪接位点分类器的支持度较差的支持向量机。少
英文摘要
In this research, we have focused on gene models which are capable of finding genes from genome sequences.First, we have developed a general purpose algorithm which finds genes by combining plural existing gene-finders. The algorithm has been implemented into a novel gene-finder named DIGIT. An outline of the algorithm is as follows. First, existing gene-finders are applied to an uncharacterized genomic sequence (input sequence). Next, DIGIT produces all possible exons from the results of gene-finders, and assigns them their exon types, reading frames and exon scores. Finally, DIGIT searches a set of exons whose additive score is maximized under their reading frame constraints. Bayesian procedure and a hidden Markov model (HMM) are used to infer exon scores and search the exon set, respectively. We have designed DIGIT so as to combine the results of FGENESH, GENSCAN and HMMgene, and have assessed its prediction accuracy by using recently compiled benchmark data sets. For all data sets, … More DIGIT successfully discarded many false-positive exons predicted by individual gene-finders and yielded remarkable improvements in sensitivity and specificity at the gene level compared with the best gene level accuracies achieved by any single gene-finder.Second, we have developed a novel index which precisely derives protein coding regions from cross-species genome alignments. The index is deeply related to frame recovery observed in coding sequence alignments, that is, if insertions or deletions of nucleotides causes frame shifts in coding regions, other in-dels which recover the reading frames will be often observed in the vicinity. In contrast, such frame recoveries are not observed in other conserved regions. We prepared two gene models: a model which finds gene by using sequence similarity and intrinsic gene measures (basic model), and the other model which finds gene by using frame recovery index in addition to sequence similarity and intrinsic gene measures (frame recovery model). We evaluated the prediction accuracies of the two models, and our benchmark test revealed that frame recovery model significantly improved the prediction accuracy in comparison with basic model.Third, we have developed GeneDecoder which is a gene finding technology for eukaryotes, based on HMMs. The algorithm, using dynamic programing method and statistic models trained by annotated genome sequences, divides the input nucleic acid sequence into some meaningful segments. Besides, GeneDecoder has some additional features: (1) multi-stream architecture, (2) incorporation of similarity search and (3) SVM-driven putative splice sites screening. (1) In addition to nucleic acid sequences, GeneDecoder allows any other data streams to be added. Typically, dicodon bigram values can be calculated in advance and be aligned on a 'Direct' stream, which makes state transition networks much simpler. Any other meaningful features extracted in advance can be incorporated to. gene-finding process using this scheme. (2) Combining calculation of coding potential and similarity search with known sequence database realizes more reliable putative exons. For this purpose, GeneDecoder has ability both to embed known motif models in exon models and to use segments with which similarity to known sequence was found by BLAST search. (3) Support Vector Machine (SVM) is one of the pattern re cognition techniques known to have high classification capability and has succes sfully been applied to splice site prediction. In GeneDecoder, this fearure is implemented as well as PWM-based splice site mod els. While parsing, putative splice sites derived from the PWM-based models but have poor support by the SVMs designed as splice site classifiers are excluded. Less
期刊论文(92)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.11234/gi1990.13.112
发表时间:
2002
期刊:
Genome informatics. International Conference on Genome Informatics
影响因子:
--
作者:
[Taishin Kin;K. Tsuda;K. Asai]
通讯作者:
Taishin Kin;K. Tsuda;K. Asai
DOI:
10.1093/bioinformatics/bti339
发表时间:
2005-05-15
期刊:
BIOINFORMATICS
影响因子:
5.8
作者:
[Kato, T, Tsuda, K, Asai, K]
通讯作者:
Asai, K
DOI:
10.1038/nature03001
发表时间:
2004-10-21
期刊:
NATURE
影响因子:
64.8
作者:
[Collins, FS, Lander, ES, Waterston, RH]
通讯作者:
Waterston, RH
Differential display analysis of mutants for the transcription factor pdr1p regulating multidrug resistance in the budding yeast
芽殖酵母多药耐药性转录因子 pdr1p 突变体的差异显示分析
DOI:
--
发表时间:
2001
期刊:
FEBS Letters 505
影响因子:
--
作者:
[Miura, F., Yada, T., Nakai, K., Sakaki, Y., Ito., T.]
通讯作者:
T.
T.Kato, K.Tsuda, K.Tomii, K Asai: "Maximum likelihood superposition of protein structures"Genome Informatics. 14. 488-489 (2003)
T.Kato、K.Tsuda、K.Tomii、K Asai:“蛋白质结构的最大似然叠加”基因组信息学。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 39 条
Designing promoter sequences
-
批准号:22240032
-
项目类别:Grant-in-Aid for Scientific Research (A)
-
资助金额:$32.03万
-
财政年份:2010
-
负责人:YADA Tetsushi
-
依托单位:
Comparative analysis of large scale genome data and knowledge discovery
-
批准号:17018021
-
项目类别:Grant-in-Aid for Scientific Research on Priority Areas
-
资助金额:$44.54万
-
财政年份:2005
-
负责人:YADA Tetsushi
-
依托单位:
海外基金