Meta-analysis of repository data: impact of data regularization on NIMH schizophrenia linkage results.

Meta-analysis of repository data: impact of data regularization on NIMH schizophrenia linkage results.
复制标题

DOI:
10.1371/journal.pone.0084696
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Vieland VJ
Vieland VJ
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Walters KA;Huang Y;Azaro M;Tobin K;Lehner T;Brzustowicz LM;Vieland VJ

文献摘要

参考文献

被引文献

相似文献

人类遗传学家越来越多地转向基于非常大的样本量的研究设计,以克服研究复杂疾病的困难。这反过来几乎总是需要通过集中存储库进行多站点数据收集和数据处理。虽然此类存储库具有许多优点,包括能够返回以前收集的数据以应用新的分析技术,但它们也有一些局限性。为了说明这一点,我们回顾了 NIMH 资助的精神疾病基因组合作研究中心(也称为人类遗传学计划 (HGI))提供的七项较早的精神分裂症研究的数据,并评估了数据清理和规范化对连锁分析的影响。开发了广泛的数据正则化协议并将其应用于基因型和表型数据。每个研究在数据处理的各个阶段都计算了全基因组非参数连锁(NPL)统计数据。为了评估数据处理对汇总结果的影响,进行了基因组扫描荟萃分析 (GSMA)。当将基于原始 HGI 数据的连锁结果与使用同一组谱系内的后处理数据的结果进行比较时,发现了连锁峰增加、减少和移动的例子。有趣的是,减少受影响个体的数量往往会增加而不是减少连锁峰值。但最重要的是,虽然单个数据集中数据正则化的影响很小,但 GSMA 应用于聚合数据后,在数据正则化后产生了截然不同的情况。这些结果对于基于其他类型数据(例如病例对照 GWAS 或测序数据)以及从其他存储库获得的数据的分析具有影响。
Human geneticists are increasingly turning to study designs based on very large sample sizes to overcome difficulties in studying complex disorders. This in turn almost always requires multi-site data collection and processing of data through centralized repositories. While such repositories offer many advantages, including the ability to return to previously collected data to apply new analytic techniques, they also have some limitations. To illustrate, we reviewed data from seven older schizophrenia studies available from the NIMH-funded Center for Collaborative Genomic Studies on Mental Disorders, also known as the Human Genetics Initiative (HGI), and assessed the impact of data cleaning and regularization on linkage analyses. Extensive data regularization protocols were developed and applied to both genotypic and phenotypic data. Genome-wide nonparametric linkage (NPL) statistics were computed for each study, over various stages of data processing. To assess the impact of data processing on aggregate results, Genome-Scan Meta-Analysis (GSMA) was performed. Examples of increased, reduced and shifted linkage peaks were found when comparing linkage results based on original HGI data to results using post-processed data within the same set of pedigrees. Interestingly, reducing the number of affected individuals tended to increase rather than decrease linkage peaks. But most importantly, while the effects of data regularization within individual data sets were small, GSMA applied to the data in aggregate yielded a substantially different picture after data regularization. These results have implications for analyses based on other types of data (e.g., case-control GWAS or sequencing data) as well as data obtained from other repositories.
DOI: 10.1093/bioinformatics/bti529
发表时间: 2005-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Wigginton, JE;Abecasis, GR
通讯作者: Abecasis, GR
DOI: 10.1046/j.1469-1809.1999.6330263.x
发表时间: 1999-05-01
影响因子: 1.9
作者:
Wise, LH;Lanchbury, JS;Lewis, CN
通讯作者: Lewis, CN
DOI: 10.1176/appi.ajp.163.10.1760
发表时间: 2006-10-01
影响因子: 17.7
作者:
Faraone, Stephen V.;Hwu, Hai-Gwo;Tsuang, Ming T.
通讯作者: Tsuang, Ming T.
DOI: 10.1038/ng786
发表时间: 2002-01-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Abecasis, GR;Cherny, SS;Cardon, LR
通讯作者: Cardon, LR
DOI: 10.1086/301592
发表时间: 1997-11-01
影响因子: 9.8
作者:
Kong, A;Cox, NJ
通讯作者: Cox, NJ