Mind the gaps: overlooking inaccessible regions confounds statistical testing in genome analysis

Mind the gaps: overlooking inaccessible regions confounds statistical testing in genome analysis
复制标题

DOI:
10.1186/s12859-018-2438-1
复制
发表时间:
2018-12-14
期刊:
影响因子:
3
通讯作者:
Sandve, Geir Kjetil
Sandve, Geir Kjetil
中科院分区:
生物学4区
文献类型:
--
作者:
Domanska, Diana;Kanduri, Chakravarthi;Sandve, Geir Kjetil

文献摘要

被引文献

相似文献

背景当前版本的参考基因组组件仍然包含由 N 段表示的间隙。由于高通量测序读数无法映射到这些间隙区域,因此这些区域的实验数据已耗尽。此外,一些技术平台分析基因组序列的目标部分,这意味着在这些实验中无法检测到基因组序列的未分析部分的区域。我们在这里将所有此类区域称为不可访问区域,并假设在零模型中忽略这些区域可能会增加基因组特征共定位统计测试中的错误结果。结果我们的探索性分析证实,公共基因组轨迹中的基因组区域与人类参考基因组(hg19 和 hg38)的组装间隙几乎没有交叉。仅在间隙区域的开始和结束部分观察到小交叉。此外,我们通过匹配真实基因组轨迹的属性来模拟一组合成轨迹,这种方式消除了它们之间的任何真实关联。这使我们能够检验我们的假设,即不避免零模型中的不可访问区域(以装配间隙为代表)将导致统计显着性的虚假膨胀。我们对比了基于蒙特卡罗的排列测试的测试统计量和 p 值的分布,在测试一对轨道之间的共定位时,这些排列测试避免或未避免零模型中的组装间隙。我们观察到,没有考虑零模型中组装间隙的统计测试导致测试统计量的分布向右移动,p 值的分布向左移动(表明显着性夸大)。我们在 hg19 和 hg38 中观察到类似水平的显着性夸大,尽管组装间隙覆盖了后者参考基因组的较小比例。结论我们提供的经验证据表明,不可访问的区域,即使只覆盖基因组的几个百分比,如果不在统计共定位分析中考虑,也可能导致大量错误的结果。
BackgroundThe current versions of reference genome assemblies still contain gaps represented by stretches of Ns. Since high throughput sequencing reads cannot be mapped to those gap regions, the regions are depleted of experimental data. Moreover, several technology platforms assay a targeted portion of the genomic sequence, meaning that regions from the unassayed portion of the genomic sequence cannot be detected in those experiments. We here refer to all such regions as inaccessible regions, and hypothesize that ignoring these regions in the null model may increase false findings in statistical testing of colocalization of genomic features.ResultsOur explorative analyses confirm that the genomic regions in public genomic tracks intersect very little with assembly gaps of human reference genomes (hg19 and hg38). The little intersection was observed only at the beginning and end portions of the gap regions. Further, we simulated a set of synthetic tracks by matching the properties of real genomic tracks in a way that nullified any true association between them. This allowed us to test our hypothesis that not avoiding inaccessible regions (as represented by assembly gaps) in the null model would result in spurious inflation of statistical significance. We contrasted the distributions of test statistics and p-values of Monte Carlo-based permutation tests that either avoided or did not avoid assembly gaps in the null model when testing colocalization between a pair of tracks. We observed that the statistical tests that did not account for assembly gaps in the null model resulted in a distribution of the test statistic that is shifted to the right and a distribution of p-values that is shifted to the left (indicating inflated significance). We observed a similar level of inflated significance in hg19 and hg38, despite assembly gaps covering a smaller proportion of the latter reference genome.ConclusionWe provide empirical evidence demonstrating that inaccessible regions, even when covering only a few percentages of the genome, can lead to a substantial amount of false findings if not accounted for in statistical colocalization analysis.