课题基金 / 基金详情

Modeling and prediction of genome sequence information by using information representation models

Modeling and prediction of genome sequence information by using information representation models
利用信息表示模型对基因组序列信息进行建模和预测
批准号:
12208010
负责人:
YADA Tetsushi
金额:
$46.34万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research on Priority Areas
财政年份:
2000
资助国家:
日本
项目状态:
已结题
起止时间:
2000 至 2004

项目摘要

项目成果

YADA Tetsushi的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
In this research, we have focused on gene models which are capable of finding genes from genome sequences.First, we have developed a general purpose algorithm which finds genes by combining plural existing gene-finders. The algorithm has been implemented into a novel gene-finder named DIGIT. An outline of the algorithm is as follows. First, existing gene-finders are applied to an uncharacterized genomic sequence (input sequence). Next, DIGIT produces all possible exons from the results of gene-finders, and assigns them their exon types, reading frames and exon scores. Finally, DIGIT searches a set of exons whose additive score is maximized under their reading frame constraints. Bayesian procedure and a hidden Markov model (HMM) are used to infer exon scores and search the exon set, respectively. We have designed DIGIT so as to combine the results of FGENESH, GENSCAN and HMMgene, and have assessed its prediction accuracy by using recently compiled benchmark data sets. For all data sets, … More DIGIT successfully discarded many false-positive exons predicted by individual gene-finders and yielded remarkable improvements in sensitivity and specificity at the gene level compared with the best gene level accuracies achieved by any single gene-finder.Second, we have developed a novel index which precisely derives protein coding regions from cross-species genome alignments. The index is deeply related to frame recovery observed in coding sequence alignments, that is, if insertions or deletions of nucleotides causes frame shifts in coding regions, other in-dels which recover the reading frames will be often observed in the vicinity. In contrast, such frame recoveries are not observed in other conserved regions. We prepared two gene models: a model which finds gene by using sequence similarity and intrinsic gene measures (basic model), and the other model which finds gene by using frame recovery index in addition to sequence similarity and intrinsic gene measures (frame recovery model). We evaluated the prediction accuracies of the two models, and our benchmark test revealed that frame recovery model significantly improved the prediction accuracy in comparison with basic model.Third, we have developed GeneDecoder which is a gene finding technology for eukaryotes, based on HMMs. The algorithm, using dynamic programing method and statistic models trained by annotated genome sequences, divides the input nucleic acid sequence into some meaningful segments. Besides, GeneDecoder has some additional features: (1) multi-stream architecture, (2) incorporation of similarity search and (3) SVM-driven putative splice sites screening. (1) In addition to nucleic acid sequences, GeneDecoder allows any other data streams to be added. Typically, dicodon bigram values can be calculated in advance and be aligned on a 'Direct' stream, which makes state transition networks much simpler. Any other meaningful features extracted in advance can be incorporated to. gene-finding process using this scheme. (2) Combining calculation of coding potential and similarity search with known sequence database realizes more reliable putative exons. For this purpose, GeneDecoder has ability both to embed known motif models in exon models and to use segments with which similarity to known sequence was found by BLAST search. (3) Support Vector Machine (SVM) is one of the pattern re cognition techniques known to have high classification capability and has succes sfully been applied to splice site prediction. In GeneDecoder, this fearure is implemented as well as PWM-based splice site mod els. While parsing, putative splice sites derived from the PWM-based models but have poor support by the SVMs designed as splice site classifiers are excluded. Less
期刊论文(92)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1038/nature03001
发表时间: 2004-10-21
期刊: NATURE
影响因子: 64.8
作者: [Collins, FS, Lander, ES, Waterston, RH]
通讯作者: Waterston, RH
DOI: 10.1093/bioinformatics/bti339
发表时间: 2005-05-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者: [Kato, T, Tsuda, K, Asai, K]
通讯作者: Asai, K
DOI: 10.11234/gi1990.13.112
发表时间: 2002
期刊: Genome informatics. International Conference on Genome Informatics
影响因子: --
作者: [Taishin Kin;K. Tsuda;K. Asai]
通讯作者: Taishin Kin;K. Tsuda;K. Asai
DOI: --
发表时间: 2001
期刊: FEBS Letters 505
影响因子: --
作者: [Miura, F., Yada, T., Nakai, K., Sakaki, Y., Ito., T.]
通讯作者: T.
39
    Designing promoter sequences
    • 批准号:
      22240032
    • 项目类别:
      Grant-in-Aid for Scientific Research (A)
    • 资助金额:
      $32.03万
    • 财政年份:
      2010
    • 负责人:
      YADA Tetsushi
    • 依托单位:
    Comparative analysis of large scale genome data and knowledge discovery
    • 批准号:
      17018021
    • 项目类别:
      Grant-in-Aid for Scientific Research on Priority Areas
    • 资助金额:
      $44.54万
    • 财政年份:
      2005
    • 负责人:
      YADA Tetsushi
    • 依托单位:
    海外基金