课题基金 / 基金详情

项目摘要

项目成果

MARK BORODOVSKY的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):我们建议在几个重要方向上扩展在前一次资助期间开发的用于真核基因发现的从头算自训练算法。首先,我们将把这个算法升级到一个多层次的数据挖掘方法,以便在基因组计划的早期阶段建立一个一致的“基因组-转录组-蛋白质组”的数据结构。在这里,我们将通过对现有和丰富的数据段(匿名基因组序列)进行无监督机器学习,并随后对缺失的生物信息(编码蛋白质的基因和蛋白质)进行计算建模,来弥补实验数据(如EST数据)的信息赤字。自训练算法的一个重要新特征将是利用蛋白质水平信息来监控和增加由无监督迭代算法得出的模型的生物相关性。其次,我们将加强早先在较小规模上开发的自我训练算法,并在真菌和其他“紧凑”的真核基因组(如秀丽线虫和黑腹果蝇)上进行测试,以处理大多数复杂的真核基因组。在这种更高的复杂性水平上,我们看到寄主基因只占基因组的一小部分,这在GC组成上可能是不均匀的,充满了转座元件和假基因(除了动物基因组,一些真菌病原体的基因组以及人类寄生虫及其媒介都属于这一类)。第三,对于在同质谱中位于基因组另一端的包含细菌、古生菌、病毒和真菌物种的人类微生物组,我们将开发改进的算法和工具,用于从头开始基因识别。这项工作将与美国和国外领先的基因组中心的测序和注释小组密切联系起来完成。
英文摘要
DESCRIPTION (provided by applicant): We propose to extend the ab initio self-training algorithms for eukaryotic gene finding developed in the previous grant period in several important directions. First we will upgrade this algorithm to a multilevel data mining approach to allow construction of a consistent "genome- transcriptome-proteome" data structure at the early stages of a genome project. Here, we will compensate for an information deficit in various segments of experimental data (such as EST data) by unsupervised machine learning on existing and abundant data segments (an anonymous genomic sequence) with subsequent computational modeling of missing biological information (protein-coding genes and proteins). An important new feature of the self-training algorithm will be the utilization of protein level information to monitor and increase biological relevance of the models derived by the unsupervised iterative algorithm. Second, we will enhance the self-training algorithm developed earlier on a smaller scale and tested on fungal and other "compact" eukaryotic genomes (such as Caenorhabditis elegans and Drosophila melanogaster) to work with most complex eukaryotic genomes. At this higher level of complexity we see species with host genes occupying just a small fraction of genome which can be inhomogeneous in GC composition, populated with transposable elements and pseudogenes (besides animal genomes, genomes of some fungal pathogens as well as human parasites and their vectors fall into this category). Third, for the human microbiome containing bacterial, archaeal, viral and fungal species, situated at yet another end of the genome in homogeneity spectrum, we will develop improved algorithms and tools for ab initio gene identification. This work will be done in close contact with sequencing and annotation groups from leading genome centers both in the US and abroad.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Addressing Open Challenges of Computational Genome Annotation
  • 批准号:
    9975182
  • 项目类别:
  • 资助金额:
    $34.24万
  • 财政年份:
    2018
  • 负责人:
    MARK BORODOVSKY
  • 依托单位:
Addressing Open Challenges of Computational Genome Annotation
  • 批准号:
    9761554
  • 项目类别:
  • 资助金额:
    $34.41万
  • 财政年份:
    2018
  • 负责人:
    MARK BORODOVSKY
  • 依托单位:
NIGMS Administrative Supplements to Support Undergraduate Summer Research
  • 批准号:
    10393964
  • 项目类别:
  • 资助金额:
    $0.53万
  • 财政年份:
    2018
  • 负责人:
    MARK BORODOVSKY
  • 依托单位:
Improving Accuracy of Gene Prediction Programs of the G*
  • 批准号:
    6581987
  • 项目类别:
  • 资助金额:
    $4.69万
  • 财政年份:
    2002
  • 负责人:
    MARK BORODOVSKY
  • 依托单位:
海外基金