Orthology inference in nonmodel organisms using transcriptomes and low-coverage genomes: improving accuracy and matrix occupancy for phylogenomics.

Orthology inference in nonmodel organisms using transcriptomes and low-coverage genomes: improving accuracy and matrix occupancy for phylogenomics.
复制标题

DOI:
10.1093/molbev/msu245
复制
发表时间:
2014-11
影响因子:
10.7
通讯作者:
Smith SA
Smith SA
中科院分区:
生物学1区
文献类型:
--
作者:
Yang Y;Smith SA

文献摘要

参考文献

被引文献

相似文献

正交推论是系统发育学分析的核心。系统基因组数据集通常包括转录本和低覆盖率的基因组,这些基因组是不完整的,包含错误和异构体。这些性质可能严重违反现有启发式正字法推理的基本假设。我们提出了一种使用系统发育来进行同源和直系归属的方法。该程序首先使用相似性分数来推断假定的同源物,然后将其比对,构建成系统发育图,并修剪由深层同源基因、错误组装、移码或重组引起的虚假分支。这些最终的同源基因随后被用来识别同源基因。我们探索了四种可供选择的基于树的正字法推理方法,其中两种是新的。它们适应了基因和基因组的复制以及基因树的不一致。我们在三个已发表的数据集上展示了这些方法,包括葡萄科、膜翅目和千足虫,它们的分歧时间从大约100 Ma到400 Ma以上。该程序显著提高了推断的同源物和直系物的完整性和准确性。我们还发现,最近分歧较大和/或包括更高覆盖率基因组的数据集有更完整的直系同源物集。为了明确地评估相互冲突的系统发育信号的来源,我们应用了保持每个基因座完整的基因区域的一系列刀切分析。这里描述的方法可以扩展到100多个分类群。它们是用Python实现的,每个步骤都有独立的脚本,这使得修改它们或将它们合并到现有管道中变得很容易。所有脚本均可从https://bitbucket.org/yangya/phylogenomic_dataset_construction.获得
Orthology inference is central to phylogenomic analyses. Phylogenomic data sets commonly include transcriptomes and low-coverage genomes that are incomplete and contain errors and isoforms. These properties can severely violate the underlying assumptions of orthology inference with existing heuristics. We present a procedure that uses phylogenies for both homology and orthology assignment. The procedure first uses similarity scores to infer putative homologs that are then aligned, constructed into phylogenies, and pruned of spurious branches caused by deep paralogs, misassembly, frameshifts, or recombination. These final homologs are then used to identify orthologs. We explore four alternative tree-based orthology inference approaches, of which two are new. These accommodate gene and genome duplications as well as gene tree discordance. We demonstrate these methods in three published data sets including the grape family, Hymenoptera, and millipedes with divergence times ranging from approximately 100 to over 400 Ma. The procedure significantly increased the completeness and accuracy of the inferred homologs and orthologs. We also found that data sets that are more recently diverged and/or include more high-coverage genomes had more complete sets of orthologs. To explicitly evaluate sources of conflicting phylogenetic signals, we applied serial jackknife analyses of gene regions keeping each locus intact. The methods described here can scale to over 100 taxa. They have been implemented in python with independent scripts for each step, making it easy to modify or incorporate them into existing pipelines. All scripts are available from https://bitbucket.org/yangya/phylogenomic_dataset_construction.
TreeFam:动物基因家族系统发育树的精选数据库
DOI: 10.1093/nar/gkj118
发表时间: 2006-01-01
影响因子: 14.9
作者:
Li, Heng;Coghlan, Avril;Ruan, Jue;Coin, Lachlan James;Heriche, Jean-Karim;Osmotherly, Lara;Li, Ruiqiang;Liu, Tao;Zhang, Zhang;Bolund, Lars;Wong, Gane Ka-Shu;Zheng, Weimou;Dehal, Paramvir;Wang, Jun;Durbin, Richard
通讯作者: Durbin, Richard
DOI: 10.1186/gb-2008-9-10-235
发表时间: 2008-10-30
期刊: Genome biology
影响因子: 12.3
作者:
Gabaldón T
通讯作者: Gabaldón T
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W
DOI: 10.1093/bioinformatics/btq539
发表时间: 2010-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Larget, Bret R.;Kotha, Satish K.;Ane, Cecile
通讯作者: Ane, Cecile
DOI: 10.1111/evo.12099
发表时间: 2013-08-01
期刊: EVOLUTION
影响因子: 3.3
作者:
Cui, Rongfeng;Schumer, Molly;Rosenthal, Gil G.
通讯作者: Rosenthal, Gil G.