minotaur : A platform for the analysis and visualization of multivariate results from genome scans with R Shiny

minotaur : A platform for the analysis and visualization of multivariate results from genome scans with R Shiny
复制标题

minotaur :使用 R Shiny 对基因组扫描的多变量结果进行分析和可视化的平台

DOI:
10.1111/1755-0998.12579
复制
发表时间:
2017
影响因子:
7.7
通讯作者:
Lotterhos, Katie E.
Lotterhos, Katie E.
中科院分区:
生物学1区
文献类型:
--
作者:
Verity, Robert;Collins, Caitlin;Card, Daren C.;Schaal, Sara M.;Wang, Liuyang;Lotterhos, Katie E.

文献摘要

相似文献

基因组扫描被广泛用于识别基因组数据中的“异常值”:由于选择或其他非适应性进化力量的作用,与基因组其余部分相比具有不同模式的基因座。这些基因组数据集通常是高维的,变量之间具有复杂的相关性结构,这使得以稳健的方式识别离群值是一个挑战。马氏距离已被广泛使用,但它的主要局限性是假设数据服从简单的参数分布。在这里,我们开发了三个新的度量标准,可以用来识别多变量空间中的异常值,而不会对数据的分布做出强烈的假设。这些指标是在R Packageminotaur中实现的,它还包括一个基于Web的交互式应用程序,用于可视化高维数据集中的离群值。我们说明了如何使用这些度量来从模拟的基因数据中识别异常值,并讨论了它们在应用中可能面临的一些限制。
Genome scans are widely used to identify ‘outliers’ in genomic data: loci with different patterns compared with the rest of the genome due to the action of selection or other nonadaptive forces of evolution. These genomic data sets are often high dimensional, with complex correlation structures among variables, making it a challenge to identify outliers in a robust way. The Mahalanobis distance has been widely used, but has the major limitation of assuming that data follow a simple parametric distribution. Here, we develop three new metrics that can be used to identify outliers in multivariate space, while making no strong assumptions about the distribution of the data. These metrics are implemented in the R packageminotaur, which also includes an interactive web‐based application for visualizing outliers in high‐dimensional data sets. We illustrate how these metrics can be used to identify outliers from simulated genetic data and discuss some of the limitations they may face in application.