RUNNING HEADER: Missing data produces biased loci TITLE: Uneven missing data skews phylogenomic relationships within the lories and lorikeets

RUNNING HEADER: Missing data produces biased loci TITLE: Uneven missing data skews phylogenomic relationships within the lories and lorikeets
复制标题

标题:缺失数据产生有偏差的位点标题:不均匀的缺失数据扭曲了鹦鹉和澳洲鹦鹉内的系统发育关系

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Michael J. Andersen
Michael J. Andersen
中科院分区:
--
文献类型:
--
作者:
B. Smith;William M. Mauck;Brett W. Benz;Michael J. Andersen

文献摘要

参考文献

被引文献

相似文献

随着基因组位点的大规模并行测序,生命之树的解析速度加快。为了在进化枝内实现密集的分类单元采样,通常需要从历史博物馆标本中获得DNA以补充现代遗传样本。这种类型的采样方案所带来的一个特殊挑战是DNA序列中的预期系统性偏差,其中较旧的材料具有更多的缺失数据。在这项研究中,我们评估了缺失的数据如何影响分布在澳大利亚地区的刷舌鹦鹉或吸蜜鹦鹉和吸蜜鹦鹉(部落:吸蜜鹦鹉)的系统基因组关系。我们收集了ultraconserved元素,从现代和历史的材料,代表了大多数描述的类群中的分支。初步的基因组学分析恢复了属内样品的聚类,其中强烈支持基于样品类型形成的组。为了评估异常关系是否是由缺失数据驱动的,我们进行了离群基因座分析,并计算了构建有和没有缺失数据的树的基因似然性。我们产生了一系列比对,其中基于Δ基因对数似然分数和推断的拓扑结构用不同的数据集排除了基因座,以评估是否可以通过排除特定基因座来改变样本类型聚类。我们发现,大多数可疑的关系是由特定的基因座子集驱动的。出乎意料的是,有偏见的位点没有更高的缺失数据,而是更多的简约信息网站。这一违反直觉的结果表明,信息量最大的基因座可能受到最高的偏倚,因为最可变的基因座在样品类型之间的系统发育信号中可能具有最大的差异。在考虑了有偏见的基因座后,我们推断出Loriini的一个更强大的基因组假说。进化枝内的分类学关系现在可以被修改以反映自然分组,但对于某些组来说,额外的工作仍然是必要的。1. CC-BY-NC-ND 4.0国际许可(通过同行评审认证)是作者/资助者,他授予bioRxiv永久展示预印本的许可。它在此预印本的版权保持器下提供(不是2018年8月23日发布的此版本。; www.example.com doi:bioRxiv预印本
​.​—Resolution of the Tree of Life has accelerated with massively parallel sequencing of genomic loci. To achieve dense taxon sampling within clades, it is often necessary to obtain DNA from historical museum specimens to supplement modern genetic samples. A particular challenge that arises with this type of sampling scheme is an expected systematic bias in DNA sequences, where older material has more missing data. In this study, we evaluated how missing data influenced phylogenomic relationships in the brush-tongued parrots, or the lories and lorikeets (Tribe: Loriini), which are distributed across the Australasian region. We collected ultraconserved elements from modern and historical material representing the majority of described taxa in the clade. Preliminary phylogenomic analyses recovered clustering of samples within genera, where strongly supported groups formed based on sample type. To assess if the aberrant relationships were being driven by missing data, we performed an outlier loci analysis and calculated gene-likelihoods for trees built with and without missing data. We produced a series of alignments where loci were excluded based on Δ gene-wise log-likelihood scores and inferred topologies with the different datasets to assess whether sample-type clustering could be altered by excluding particular loci. We found that the majority of questionable relationships were driven by particular subsets of loci. Unexpectedly, the biased loci did not have higher missing data, but rather more parsimony informative sites. This counterintuitive result suggests that the most informative loci may be subject to the highest bias as the most variable loci can have the greatest disparity in phylogenetic signal among sample types. After accounting for biased loci, we inferred a more robust phylogenomic hypothesis for the Loriini. Taxonomic relationships within the clade can now be revised to reflect natural groupings, but for some groups additional work is still necessary. 1 . CC-BY-NC-ND 4.0 International license a certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under The copyright holder for this preprint (which was not this version posted August 23, 2018. ; https://doi.org/10.1101/398297 doi: bioRxiv preprint
DOI: 10.1016/j.ympev.2018.03.033
发表时间: 2018-09
影响因子: 4.1
作者:
Gilbert PS;Wu J;Simon MW;Sinsheimer JS;Alfaro ME
通讯作者: Alfaro ME
DOI: 10.1002/ece3.3065
发表时间: 2017-07
影响因子: 2.6
作者:
Linck EB;Hanna ZR;Sellas A;Dumbacher JP
通讯作者: Dumbacher JP