A statistical methodology for analyzing co-occurrence data from a large sample
A statistical methodology for analyzing co-occurrence data from a large sample
复制标题
DOI:
10.1016/j.jbi.2006.11.003
复制
发表时间:
2007-06-01
影响因子:
4.5
通讯作者:
Markatou, Marianthi
中科院分区:
文献类型:
--
作者:
Cao, Hui;Hripcsak, George;Markatou, Marianthi
Determining important associations among items in a large database is challenging due to multiple simultaneous hypotheses and the ability to select weak associations that are statistically but not clinically significant. The simple application of the 2 test among all possible pairs of items results in mostly inappropriate associations surpassing the traditional (alpha =.05, chi(2) = 3.94) threshold. One can choose a stricter threshold to find stronger associations, but the choice may be arbitrary. We combined the volume test of Diaconis and Efron with 2 a p-value plot to select a more rigorous and less arbitrary threshold. The volume test adjusts the p-value of the Z(2) -statistic. A plot of adjusted p-values (1-p versus N-p), where N-p is the number of test statistics with a p-value greater than p, should be linear if there are no true associations. The point where the plot deviates from a line can be used as a threshold. We used linear regression to select the threshold in a reproducible fashion. In one experiment, we found that the method selected a threshold similar to that previously obtained by manually reviewing associations. (C) 2006 Elsevier Inc. All rights reserved.