Computing the Statistical Significance of Overlap between Genome Annotations with ISTAT

Computing the Statistical Significance of Overlap between Genome Annotations with ISTAT
复制标题

DOI:
10.1016/j.cels.2019.05.006
复制
发表时间:
2019-06-26
期刊:
影响因子:
9.3
通讯作者:
Bafna, Vineet
Bafna, Vineet
中科院分区:
生物学1区
文献类型:
--
作者:
Sarmashghi, Shahab;Bafna, Vineet

文献摘要

被引文献

相似文献

基因组注释仍然是现代生物学中的一项基本工作。随着成本的降低和新形式的测序技术,不断产生对组织类型和实验条件特异性的注释(例如,组蛋白甲基化标记)。计算两个不同注释之间重叠的统计显著性是许多生物学发现的关键,但以前没有系统地解决。我们将问题形式化如下:让I和I-f各自描述具有特定注释的基因组的n和m个间隔的集合。在原假设I中的基因组间隔相对于I-f随机排列的情况下,I-f的m个间隔中的k与I中的间隔相交的意义是什么?我们描述了一个工具iSTAT,实现了一个组合算法,以准确地计算p值。我们将iSTAT应用于模拟和真实的数据集,以获得精确的估计值,并使用排列或参数检验将其与以前的结果进行对比。
Genome annotation remains a fundamental effort in modern biology. With reducing costs and new forms of sequencing technologies, annotations specific to tissue type and experimental conditions are continually being generated (e.g., histone methylation marks). Computing the statistical significance of overlap between two different annotations is key to many biological findings but has not been systematically addressed previously. We formalize the problem as follows: let I and I-f each describe a collection of n and m intervals of a genome with particular annotation. Under the null hypothesis that genomic intervals in I are randomly arranged with respect to I-f, what is the significance of k of m intervals of I-f intersecting with intervals in I? We describe a tool iSTAT that implements a combinatorial algorithm to accurately compute p values. We applied iSTAT to simulated and real datasets to obtain precise estimates and contrasted them against previous results using permutation or parametric tests.