课题基金 / 基金详情

Collaborative Research: Algorithms for Threat Detection via Geometry of Virus Genome Space

Collaborative Research: Algorithms for Threat Detection via Geometry of Virus Genome Space
合作研究:通过病毒基因组空间几何进行威胁检测的算法
批准号:
1120824
负责人:
Stephen S. Yau
金额:
$74.52万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-15 至 2015-08-31

项目摘要

项目成果

Stephen S. Yau的其他基金

相似基金

相关文献

中文摘要
翻译
基因组空间是基因组的模空间。在这个空间中,每个点对应于一个基因组。如果基因组空间中的对应点彼此接近,则预期两个基因组密切相关。研究者和他们的合作者获得了一种新的几何表示--DNA序列的自然向量,并证明了DNA序列和自然向量之间是一一对应的。他们在这个空间中对基因组序列进行系统发育和聚类分析。与大多数现有方法不同,本文提出的基因组空间不需要序列比对或任何进化模型,从而避免了计算重复。在对27,643个基因组序列的初步研究中,使用自然向量方法只需几个小时即可计算所有成对差异,而使用经典的多重比对方法则需要四年时间。考虑到已知基因组数据库的指数增长的大小,自然向量方法是唯一已知的可行的方法来聚类整个基因组空间。利用构造的自然向量,研究者使用基于永久过程的分类模型,即随机分类模型,进行分类和聚类。此外,还可以获得每个病毒基因组属于一个簇的概率。例如,研究人员根据全基因组序列对59个甲型H1N1流感病毒基因组和113个人鼻病毒(HRV)基因组进行了聚类分析,结果显示,此次新暴发的甲型H1N1流感病毒与欧亚猪流感病毒和北美猪流感病毒的亲缘关系最近,113个HRV基因组被很好地聚类为5类HRV-A,HRV-B、HRV-C、HEV-B和HEV-C。该方法只需18秒就能得到聚类结果,而常用的多重比对方法需要19小时以上,两种方法得到的聚类结果相同。拟议活动的第一个目标是收集每种病毒的所有现有基因组序列,计算其天然载体,建立和维持一个病毒“天然载体库”。其次,研究人员将探索自然载体的必要维数,以便准确地对基因组进行分类或聚类。第三个目标是根据病毒的自然载体对病毒基因组进行聚类。最终目标是根据基因组序列对任何给定的新病毒进行分类或识别,并预测其功能或行为模式。在该项目中,研究人员构建了一种新颖的,高速的,准确的几何表示,称为自然向量,用于DNA序列。基于这种新的强大的方法,生物学家可以同时对所有基因组进行全局比较,这是任何其他方法都无法实现的。一旦知道了基因组序列,它是非常快速和方便的,这对国土安全至关重要。为了预测来自恐怖组织的新病毒的特征,可以计算新病毒的自然载体,并将其与其他已知病毒的自然载体进行比较。通过这种方式,人们可以通过观察附近病毒的特性来预测这种新病毒的可能特性。快速准确地识别新病毒并预测其功能将非常有助于当局采取预防措施并在其达到大流行状态并在公众中传播之前制造疫苗。
英文摘要
A genome space is a moduli space of genomes. In this space each point corresponds to a genome. It is expected that two genomes are closely related if the corresponding points in the genome space are close to each other. The investigators and their collaborators obtain a new geometric representation-the natural vector for DNA sequences, and show that the correspondence between DNA sequences and the natural vectors is one-to-one. They perform phylogenetic and clustering analysis for genome sequences in this space. Unlike most existing methods, the proposed genome space here does not need sequence alignment or any evolutionary model and thus avoids computational repetition. In a pilot study of 27,643 genome sequences, it takes only a couple of hours using the natural vector method to compute all the pairwise differences, while it will take four years using the classical multiple alignment methods. Considering the exponentially increasing size of the known genome database, the natural vector method is the only known feasible approach to cluster the whole genome space. With the constructed natural vectors, the investigators use the classification model based on a permanental process, a stochastic classification model, to perform classification and clustering. Moreover, the probability of each virus genome belonging to a cluster can also be obtained. For example, the investigators did clustering analysis for 59 Influenza A H1N1 swine flu genomes and 113 human rhinovirus (HRV) genomes based on their whole genome sequences, and showed that the new outbreak of Influenza A H1N1 swine flu virus was most closely related to Eurasian swine flu viruses and North American swine flu viruses, and the 113 HRV genomes were well clustered into 5 classes HRV-A, HRV-B, HRV-C, HEV-B, and HEV-C. It takes only 18 seconds for the proposed method to get the clustering result while it takes more than 19 hours for the commonly used multiple alignment method.Both methods yield the same clustering result. The first goal of the proposed activity is to collect all available genome sequences for each type of virus, compute their natural vectors, set up and maintain a "natural vector bank" for viruses. Secondly, the investigators will explore the necessary number of dimensions of the natural vector such that it accurately classifies or clusters the genomes. The third goal is to do clustering on the virus genomes based on their natural vectors. The final goal is to classify or identify any given new virus based on its genome sequence, and predict its functions or behavior pattern.In this project the investigators construct a novel, high-speed, accurate geometric representation, called the natural vector, for DNA sequences. Based on this new powerful method, the biologists can have a global comparison of all genomes simultaneously, which cannot be achieved by any other method. It is very fast and convenient once the genome sequence is known, which is vital to the homeland security. To predict the characteristics of a new virus coming from terrorist groups, one can compute the natural vector of the new virus and compare it with the natural vectors of other known viruses. In this way one can predict the possible properties of this new virus by looking at the properties of those viruses located nearby. Quickly and accurately identifying a new virus and predicting its functions will be very helpful to authorities taking precautions and manufacturing a vaccine before it reaches a pandemic state and propagates throughout the general public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Global invariants for complex varieties with isolated singularities and applications
  • 批准号:
    0802803
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $16.47万
  • 财政年份:
    2008
  • 负责人:
    Stephen S. Yau
  • 依托单位:
Global Invariants for CR Geometry and Isolated Singularities
  • 批准号:
    0503868
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2005
  • 负责人:
    Stephen S. Yau
  • 依托单位:
U.S.-Hong Kong Joint Workshop: Recent Developments in Several Complex Variables, Cauchy Riemann Geometry and Complex Algebraic Geometry
  • 批准号:
    0224546
  • 项目类别:
    Standard Grant
  • 资助金额:
    $3.16万
  • 财政年份:
    2002
  • 负责人:
    Stephen S. Yau
  • 依托单位:
Computation in Nonlinear Filtering
  • 批准号:
    9975354
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.5万
  • 财政年份:
    1999
  • 负责人:
    Stephen S. Yau
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)