课题基金 / 基金详情

项目摘要

项目成果

MARK BORODOVSKY的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): We propose to extend the ab initio self-training algorithms for eukaryotic gene finding developed in the previous grant period in several important directions. First we will upgrade this algorithm to a multilevel data mining approach to allow construction of a consistent "genome- transcriptome-proteome" data structure at the early stages of a genome project. Here, we will compensate for an information deficit in various segments of experimental data (such as EST data) by unsupervised machine learning on existing and abundant data segments (an anonymous genomic sequence) with subsequent computational modeling of missing biological information (protein-coding genes and proteins). An important new feature of the self-training algorithm will be the utilization of protein level information to monitor and increase biological relevance of the models derived by the unsupervised iterative algorithm. Second, we will enhance the self-training algorithm developed earlier on a smaller scale and tested on fungal and other "compact" eukaryotic genomes (such as Caenorhabditis elegans and Drosophila melanogaster) to work with most complex eukaryotic genomes. At this higher level of complexity we see species with host genes occupying just a small fraction of genome which can be inhomogeneous in GC composition, populated with transposable elements and pseudogenes (besides animal genomes, genomes of some fungal pathogens as well as human parasites and their vectors fall into this category). Third, for the human microbiome containing bacterial, archaeal, viral and fungal species, situated at yet another end of the genome in homogeneity spectrum, we will develop improved algorithms and tools for ab initio gene identification. This work will be done in close contact with sequencing and annotation groups from leading genome centers both in the US and abroad.
期刊论文(46)
专著(0)
科研奖励(0)
会议论文
Multiple testing in large-scale contingency tables: inferring patterns of pair-wise amino acid association in beta-sheets.
大规模列联表中的多重测试:推断β-折叠中成对氨基酸关联的模式。
DOI: 10.1504/ijbra.2006.009768
发表时间: 2006
期刊: International journal of bioinformatics research and applications
影响因子: --
作者: [Kim,SeoungBum, Tsui,Kwok-Leung, Borodovsky,Mark]
通讯作者: Borodovsky,Mark
DOI: 10.1504/ijbra.2009.027519
发表时间: 2009-01-01
期刊: International journal of bioinformatics research and applications
影响因子: --
作者: [Kislyuk, Andrey, Lomsadze, Alexandre, Borodovsky, Mark]
通讯作者: Borodovsky, Mark
DOI: 10.1089/cmb.1995.2.87
发表时间: 1995-01-01
期刊: Journal of computational biology : a journal of computational molecular cell biology
影响因子: --
作者: [Gelfand, M S]
通讯作者: Gelfand, M S
DOI: 10.1093/nar/gkq275
发表时间: 2010-07
期刊: Nucleic acids research
影响因子: 14.9
作者: [Zhu W, Lomsadze A, Borodovsky M]
通讯作者: Borodovsky M
29
    Addressing Open Challenges of Computational Genome Annotation
    • 批准号:
      9975182
    • 项目类别:
    • 资助金额:
      $34.24万
    • 财政年份:
      2018
    • 负责人:
      MARK BORODOVSKY
    • 依托单位:
    Addressing Open Challenges of Computational Genome Annotation
    • 批准号:
      9761554
    • 项目类别:
    • 资助金额:
      $34.41万
    • 财政年份:
      2018
    • 负责人:
      MARK BORODOVSKY
    • 依托单位:
    NIGMS Administrative Supplements to Support Undergraduate Summer Research
    • 批准号:
      10393964
    • 项目类别:
    • 资助金额:
      $0.53万
    • 财政年份:
      2018
    • 负责人:
      MARK BORODOVSKY
    • 依托单位:
    Improving Accuracy of Gene Prediction Programs of the G*
    • 批准号:
      6581987
    • 项目类别:
    • 资助金额:
      $4.69万
    • 财政年份:
      2002
    • 负责人:
      MARK BORODOVSKY
    • 依托单位:
    海外基金