Computing the Statistical Significance of Overlap between Genome Annotations with ISTAT
Computing the Statistical Significance of Overlap between Genome Annotations with ISTAT
复制标题
DOI:
10.1016/j.cels.2019.05.006
复制
发表时间:
2019-06-26
期刊:
影响因子:
9.3
通讯作者:
Bafna, Vineet
中科院分区:
文献类型:
--
作者:
Sarmashghi, Shahab;Bafna, Vineet
Genome annotation remains a fundamental effort in modern biology. With reducing costs and new forms of sequencing technologies, annotations specific to tissue type and experimental conditions are continually being generated (e.g., histone methylation marks). Computing the statistical significance of overlap between two different annotations is key to many biological findings but has not been systematically addressed previously. We formalize the problem as follows: let I and I-f each describe a collection of n and m intervals of a genome with particular annotation. Under the null hypothesis that genomic intervals in I are randomly arranged with respect to I-f, what is the significance of k of m intervals of I-f intersecting with intervals in I? We describe a tool iSTAT that implements a combinatorial algorithm to accurately compute p values. We applied iSTAT to simulated and real datasets to obtain precise estimates and contrasted them against previous results using permutation or parametric tests.