Genome rearrangement with gene families

Genome rearrangement with gene families
复制标题

DOI:
10.1093/bioinformatics/15.11.909
复制
发表时间:
1999-11-01
期刊:
影响因子:
5.8
通讯作者:
Sankoff, D
Sankoff, D
中科院分区:
生物学3区
文献类型:
--
作者:
Sankoff, D

文献摘要

被引文献

相似文献

动机:基因组重排分析的理论和实践在生物学上普遍存在的情况下失败了,在这种情况下,每个基因可能以多个拷贝存在,而不一定是连续的。然而,在这些背景中的某些情况下,询问两个基因组G和H(长度l(G)和l(H))中的每个基因家族的哪些成员是其真正的范例是适当的,即,哪些最好地反映了共同祖先基因组中祖先基因的原始位置。这需要搜索两个长度相同的示例字符串n(=基因家族的数量,包括单例),具有最小可能的重排距离:样本距离。结果:分支和定界算法在基于容易计算的传统重排距离(诸如带符号的反转距离或断点距离)时有效地计算这些距离,其还满足基因数目的单调性的性质。模拟结果表明,在两个随机基因组中,期望样本距离/n对基因家族的数目和大小敏感,但随着单例家族数目的增加而接近1。当基本重排距离只是断点的数量时,计算示例性断点距离(EBD)的预期成本(如通过对基础断点距离例程的总调用所测量的)高度依赖于n和基因家族的构型。另一方面,基于样本反转距离(ERD)的样本距离,预期的计算成本取决于基因家族的配置,但对n不敏感。可用性:EBD和ERD的代码可从作者处获得,或可在http://www.crm.umontreal.ca/(类似于)viart/exemplar_dis.html联系人:sankoff@ere.umontreal.ca访问。
Motivation: The theory and practice of genome rearrangement analysis breaks down in the biologically widespread contexts where each gene may bt present in a number of copies, not necessarily contiguous. In some of these contexts it is, however appropriate to ask which members of each gene family in two genomes G and H, lengths l(G) and l(H), are its true exemplars, i.e. which best reflect the original position of the ancestral gene in the common ancestor genome. This entails a search for the two exemplar strings of same length n (= number of gene families, including singletons), having the smallest possible rearrangement distance: the exemplar distance.Results: A branch and bound algorithm calculates these distances efficiently when based on easily calculated traditional rearrangement distances, such as signed reversals distance or breakpoint distance, which also satisfy a property of monotonicity in the number of genes. Simulations show that in two random genomes, the expected exemplar distance/n is sensitive to the number and size of gene families, but approaches 1 as the number of singleton families increases. When the basic rearrangement distance is just the number of breakpoints, the expected cost of computing the exemplar breakpoints distance (EBD), as measured by total calls to the underlying breakpoint distance routine, is highly dependent on both n and the configuration of gene families. On the other hand, basing exemplar distance on exemplar reversals distance (ERD), the expected computing cost depends on the configuration of gene families but is not sensitive to n.Availability: Code for EBD and ERD is available from the author or may be accessed at http://www.crm.umontreal.ca/(similar to)viart/exemplar_dis.htmlContact: sankoff@ere.umontreal.ca.