Expanding the boundaries of local similarity analysis.

Expanding the boundaries of local similarity analysis.
复制标题

DOI:
10.1186/1471-2164-14-s1-s3
复制
发表时间:
2013
期刊:
影响因子:
4.4
通讯作者:
Hallam SJ
Hallam SJ
中科院分区:
生物学2区
文献类型:
--
作者:
Durno WE;Hanson NW;Konwar KM;Hallam SJ

文献摘要

被引文献

相似文献

时间序列数据的局部和时滞关系的成对比较是一个计算上具有挑战性的问题,涉及到许多领域的调查。局部相似性分析(LSA)统计量识别局部和滞后关系的存在,但是通过p值确定显著性在算法上很麻烦,因为需要进行密集的排列测试、重排行和列以及重复计算统计量。此外,这个p值是在正态性假设下计算的--这是一种与大多数真实的世界数据集无关的统计奢侈。为了提高LSA在大数据集上的性能,在没有正态性假设的情况下推导出p值计算的渐近上界。这种边界计算的变化显着提高了计算速度,从O(pm 2n)到O(m2 n),其中p是排列测试中的排列数,m是时间序列的数量,n是每个时间序列的长度。绑定过程被实现为计算效率高的软件包FASTLSA,该软件包用C编写,并针对多核计算机上的线程进行了优化,从而提高了其实际计算时间。我们计算比较我们的方法,以前的实现LSA,展示了广泛的适用性,通过分析时间序列数据,从公共卫生,微生物生态学和社交媒体,并可视化所产生的网络使用Cytoscape软件。FASTLSA软件包扩展了LSA的边界,允许分析具有数百万个共变时间序列的数据集。将元数据映射到从FASTLSA导出的力导向图上,使调查人员能够查看相关的集团并探索以前未识别的网络关系。该软件可在http://www.cmde.science.ubc.ca/hallam/fastLSA/免费下载。
Pairwise comparison of time series data for both local and time-lagged relationships is a computationally challenging problem relevant to many fields of inquiry. The Local Similarity Analysis (LSA) statistic identifies the existence of local and lagged relationships, but determining significance through a p-value has been algorithmically cumbersome due to an intensive permutation test, shuffling rows and columns and repeatedly calculating the statistic. Furthermore, this p-value is calculated with the assumption of normality -- a statistical luxury dissociated from most real world datasets. To improve the performance of LSA on big datasets, an asymptotic upper bound on the p-value calculation was derived without the assumption of normality. This change in the bound calculation markedly improved computational speed from O(pm2n) to O(m2n), where p is the number of permutations in a permutation test, m is the number of time series, and n is the length of each time series. The bounding process is implemented as a computationally efficient software package, FASTLSA, written in C and optimized for threading on multi-core computers, improving its practical computation time. We computationally compare our approach to previous implementations of LSA, demonstrate broad applicability by analyzing time series data from public health, microbial ecology, and social media, and visualize resulting networks using the Cytoscape software. The FASTLSA software package expands the boundaries of LSA allowing analysis on datasets with millions of co-varying time series. Mapping metadata onto force-directed graphs derived from FASTLSA allows investigators to view correlated cliques and explore previously unrecognized network relationships. The software is freely available for download at: http://www.cmde.science.ubc.ca/hallam/fastLSA/.