Toward modernizing the systematic review pipeline in genetics: efficient updating via data mining

Toward modernizing the systematic review pipeline in genetics: efficient updating via data mining
复制标题

DOI:
10.1038/gim.2012.7
复制
发表时间:
2012-07-01
影响因子:
8.8
通讯作者:
Trikalinos, Thomas A.
Trikalinos, Thomas A.
中科院分区:
医学1区
文献类型:
--
作者:
Wallace, Byron C.;Small, Kevin;Trikalinos, Thomas A.

文献摘要

被引文献

相似文献

目的:本研究的目的是证明,现代数据挖掘工具可以作为一个步骤,以减少必要的劳动生产和维护systematic reviews.Methods:我们使用了四个不断更新的,人工策划的资源,总结了MEDLINE索引的文章在整个领域使用系统性综述方法(PD基因、AlzGene和SzGene分别用于帕金森病、阿尔茨海默病和精神分裂症的遗传决定因素;和塔夫茨成本-效果分析(CEA)登记处进行成本-效果分析)。在每个数据集中,我们训练了一个分类模型,用于筛选2009年之前的引文。然后,我们使用人类筛选作为金标准,评估了该模型将2010年发表的引文分类为“相关”或“不相关”的能力。在PDGene、AlzGene和SzGene中,分类模型分别没有漏掉104、65和179个合格引文中的任何一个,CEA登记处仅漏诊79例中的1例(前3例的敏感性为100%,第4例为99%)。特异性分别为90%、93%、90%和73%。如果在2010年使用半自动化系统,那么人类只需要阅读605/5,616篇引文就可以更新PDGene注册表(11%),其他三个数据库分别为555/7,298(8%)、717/5,381(13%)和334/1,015(33%)。数据挖掘方法可以减少更新系统评价的负担,而不会比人类错过更多的论文。
Purpose: The aim of this study was to demonstrate that modern data mining tools can be used as one step in reducing the labor necessary to produce and maintain systematic reviews.Methods: We used four continuously updated, manually curated resources that summarize MEDLINE-indexed articles in entire fields using systematic review methods (PDGene, AlzGene, and SzGene for genetic determinants of Parkinson disease, Alzheimer disease, and schizophrenia, respectively; and the Tufts Cost-Effectiveness Analysis (CEA) Registry for cost-effectiveness analyses). In each data set, we trained a classification model on citations screened up until 2009. We then evaluated the ability of the model to classify citations published in 2010 as "relevant" or "irrelevant" using human screening as the gold standard.Results: Classification models did not miss any of the 104, 65, and 179 eligible citations in PDGene, AlzGene, and SzGene, respectively, and missed only 1 of 79 in the CEA Registry (100% sensitivity for the first three and 99% for the fourth). The respective specificities were 90, 93, 90, and 73%. Had the semiautomated system been used in 2010, a human would have needed to read only 605/5,616 citations to update the PDGene registry (11%) and 555/7,298 (8%), 717/5,381 (13%), and 334/1,015 (33%) for the other three databases.Conclusion: Data mining methodologies can reduce the burden of updating systematic reviews, without missing more papers than humans.