SCaFoS: a tool for selection, concatenation and fusion of sequences for phylogenomics.

SCaFoS: a tool for selection, concatenation and fusion of sequences for phylogenomics.
复制标题

DOI:
10.1186/1471-2148-7-s1-s2
复制
发表时间:
2007-02-08
影响因子:
3.4
通讯作者:
Philippe H
Philippe H
中科院分区:
生物学2区
文献类型:
--
作者:
Roure B;Rodriguez-Ezpeleta N;Philippe H

文献摘要

被引文献

相似文献

基于富含基因和物种的数据集进行的系统发育分析(系统基因组学)正成为解决进化问题的一种标准方法。然而,大型数据集的组装存在一些困难,例如每个物种存在基因的多个拷贝(旁系同源或异源基因)、给定物种缺少某些基因,或者是部分序列。在系统发育推断中使用未检测到的旁系同源或异源基因会导致不准确的结果,使用部分序列会导致分辨率不足。在系统基因组学背景下,需要一种在处理这些问题的同时能够选择序列、物种和基因的工具。 在此,我们介绍SCaFoS,这是一种能够快速组装包含最大系统发育信息的系统基因组数据集的工具,同时在选择物种、序列和基因时调整缺失数据的量。从单个序列比对开始,并使用用户定义的单系群,SCaFoS利用部分序列创建嵌合体,或者在多个序列中选择直系同源和/或进化最慢的序列。一旦代表每个预定义单系群的序列被选中,SCaFoS会根据用户允许的缺失数据水平保留基因,并以与标准系统发育推断软件兼容的几种格式生成用于超矩阵和超树分析的文件。由于不存在明确的序列选择标准,所以提供了一种半自动模式以适应用户的专业知识。 SCaFoS能够处理数百个物种和基因的数据集,无论是在氨基酸还是核苷酸水平。它有一个图形界面,并且可以集成到自动工作流程中。此外,SCaFoS是第一个整合用户知识来选择直系同源序列、创建嵌合序列以减少缺失数据以及根据基因的缺失数据水平选择基因的工具。最后,将SCaFoS应用于不同的数据集,我们表明对基因、物种和序列的明智选择减少了树重建假象,尤其是当数据集包含快速进化的物种时。
Phylogenetic analyses based on datasets rich in both genes and species (phylogenomics) are becoming a standard approach to resolve evolutionary questions. However, several difficulties are associated with the assembly of large datasets, such as multiple copies of a gene per species (paralogous or xenologous genes), lack of some genes for a given species, or partial sequences. The use of undetected paralogous or xenologous genes in phylogenetic inference can lead to inaccurate results, and the use of partial sequences to a lack of resolution. A tool that selects sequences, species, and genes, while dealing with these issues, is needed in a phylogenomics context. Here, we present SCaFoS, a tool that quickly assembles phylogenomic datasets containing maximal phylogenetic information while adjusting the amount of missing data in the selection of species, sequences and genes. Starting from individual sequence alignments, and using monophyletic groups defined by the user, SCaFoS creates chimeras with partial sequences, or selects, among multiple sequences, the orthologous and/or slowest evolving sequences. Once sequences representing each predefined monophyletic group have been selected, SCaFos retains genes according to the user's allowed level of missing data and generates files for super-matrix and super-tree analyses in several formats compatible with standard phylogenetic inference software. Because no clear-cut criteria exist for the sequence selection, a semi-automatic mode is available to accommodate user's expertise. SCaFos is able to deal with datasets of hundreds of species and genes, both at the amino acid or nucleotide level. It has a graphical interface and can be integrated in an automatic workflow. Moreover, SCaFoS is the first tool that integrates user's knowledge to select orthologous sequences, creates chimerical sequences to reduce missing data and selects genes according to their level of missing data. Finally, applying SCaFoS to different datasets, we show that the judicious selection of genes, species and sequences reduces tree reconstruction artefacts, especially if the dataset includes fast evolving species.