Defining window-boundaries for genomic analyses using smoothing spline techniques.

Defining window-boundaries for genomic analyses using smoothing spline techniques.
复制标题

DOI:
10.1186/s12711-015-0105-9
复制
发表时间:
2015-04-17
期刊:
Genetics, selection, evolution : GSE
影响因子:
--
通讯作者:
de Leon N
de Leon N
中科院分区:
其他
文献类型:
--
作者:
Beissinger TM;Rosa GJ;Kaeppler SM;Gianola D;de Leon N

文献摘要

参考文献

被引文献

相似文献

高密度基因组数据通常通过在相邻标记的窗口上组合信息来分析。在窗口中分组的数据与在单个位置处分组的数据的解释可以增加统计功效、简化计算、减少采样噪声并且减少所执行的测试的总数。然而,使用相邻标记信息可能导致过度平滑或欠平滑、不期望的窗口边界规范或高度相关的测试统计。我们介绍了一种方法,用于定义窗口的基础上统计指导断点的数据,作为基础的多个相邻的数据点的分析。该方法首先将三次平滑样条拟合到数据,然后识别拟合样条的拐点,这些拐点用作相邻窗口的边界。该技术不需要连锁不平衡的先验知识,因此可以应用于从单独或合并的测序实验收集的数据。此外,与现有方法相比,窗口大小的任意选择不是必需的,因为这些是凭经验确定的,并且允许沿着基因组变化。进行应用该方法的模拟以从合并的测序FST数据中鉴定选择特征,其中等位基因频率从个体池中估计。真阳性与假阳性的相对比率是现有技术的两倍。该方法与先前的一项研究(涉及来自玉米的合并测序FST数据)的比较表明,与使用标准滑动窗口方法相比,外围窗口与其相邻窗口更清楚地分离。我们已经开发出一种新的技术,以确定后续分析协议的窗口边界。当应用于基于FST数据的选择研究时,该方法提供了高发现率并最小化假阳性。该方法在R包GenWin中实现,该包可从CRAN公开获得。
High-density genomic data is often analyzed by combining information over windows of adjacent markers. Interpretation of data grouped in windows versus at individual locations may increase statistical power, simplify computation, reduce sampling noise, and reduce the total number of tests performed. However, use of adjacent marker information can result in over- or under-smoothing, undesirable window boundary specifications, or highly correlated test statistics. We introduce a method for defining windows based on statistically guided breakpoints in the data, as a foundation for the analysis of multiple adjacent data points. This method involves first fitting a cubic smoothing spline to the data and then identifying the inflection points of the fitted spline, which serve as the boundaries of adjacent windows. This technique does not require prior knowledge of linkage disequilibrium, and therefore can be applied to data collected from individual or pooled sequencing experiments. Moreover, in contrast to existing methods, an arbitrary choice of window size is not necessary, since these are determined empirically and allowed to vary along the genome. Simulations applying this method were performed to identify selection signatures from pooled sequencing FST data, for which allele frequencies were estimated from a pool of individuals. The relative ratio of true to false positives was twice that generated by existing techniques. A comparison of the approach to a previous study that involved pooled sequencing FST data from maize suggested that outlying windows were more clearly separated from their neighbors than when using a standard sliding window approach. We have developed a novel technique to identify window boundaries for subsequent analysis protocols. When applied to selection studies based on FST data, this method provides a high discovery rate and minimizes false positives. The method is implemented in the R package GenWin, which is publicly available from CRAN.
DOI: 10.1371/journal.pone.0049525
发表时间: 2012
期刊: PloS one
影响因子: 3.7
作者:
Qanbari S;Strom TM;Haberer G;Weigend S;Gheyas AA;Turner F;Burt DW;Preisinger R;Gianola D;Simianer H
通讯作者: Simianer H
DOI: 10.1007/bf02162161
发表时间: 1967-01-01
影响因子: 2.1
作者:
REINSCH, CH
通讯作者: REINSCH, CH
DOI: 10.1038/nature08832
发表时间: 2010-03-25
期刊: NATURE
影响因子: 64.8
作者:
Rubin, Carl-Johan;Zody, Michael C.;Andersson, Leif
通讯作者: Andersson, Leif
DOI: 10.1186/1471-2148-13-150
发表时间: 2013-07-12
影响因子: 3.4
作者:
Hider JL;Gittelman RM;Shah T;Edwards M;Rosenbloom A;Akey JM;Parra EJ
通讯作者: Parra EJ
DOI: 10.1093/gbe/evt100
发表时间: 2013
影响因子: 3.3
作者:
Kelly JK;Koseva B;Mojica JP
通讯作者: Mojica JP