Performance of InSilicoVA for assigning causes of death to verbal autopsies: multisite validation study using clinical diagnostic gold standards

Performance of InSilicoVA for assigning causes of death to verbal autopsies: multisite validation study using clinical diagnostic gold standards
复制标题

DOI:
10.1186/s12916-018-1039-1
复制
发表时间:
2018-04-19
期刊:
影响因子:
9.3
通讯作者:
Lopez, Alan D.
Lopez, Alan D.
中科院分区:
医学1区
文献类型:
--
作者:
Flaxman, Abraham D.;Joseph, Jonathan C.;Lopez, Alan D.

文献摘要

被引文献

相似文献

背景:最近,一种新的口述尸检数据计算机自动认证算法InSilicoVA问世。作者给出了他们的算法作为一种统计方法,并使用一组模型预测者和一个年龄组来评估其性能。方法:我们使用相同的数据和作者发布的公开可用的算法实现,执行了一个标准程序来分析言语尸检分类方法的预测精度。我们将原来的分析扩展到包括儿童和新生儿,而不是只包括成人,并使用不同的预测值集合来测试准确率,包括原始论文中使用的集合和与发布的软件匹配的集合。结果:当对数据进行类似于原始研究的预处理时,该算法的总体水平性能(即预测准确率)从2.1%到37.6%不等。当使用与软件默认格式匹配的数据进行训练时,性能范围从-11.5%到17.5%。当使用提供的默认训练数据时,性能范围从-59.4%到-38.5%。总体而言,InSilicoVA的预测精度被发现比另一种算法低11.6-8.2个百分点。此外,InSilicoVA的灵敏度始终低于替代诊断算法(GATIRS 2.0),尽管其特异度相当。结论:该软件提供的默认格式和训练数据导致的结果充其量是次优的,死因预测性能较差。这种方法可能会产生错误的死亡原因预测,即使正确配置,也不像其他自动诊断方法那样准确。
Background: Recently, a new algorithm for automatic computer certification of verbal autopsy data named InSilicoVA was published. The authors presented their algorithm as a statistical method and assessed its performance using a single set of model predictors and one age group.Methods: We perform a standard procedure for analyzing the predictive accuracy of verbal autopsy classification methods using the same data and the publicly available implementation of the algorithm released by the authors. We extend the original analysis to include children and neonates, instead of only adults, and test accuracy using different sets of predictors, including the set used in the original paper and a set that matches the released software.Results: The population-level performance (i.e., predictive accuracy) of the algorithm varied from 2.1 to 37.6% when trained on data preprocessed similarly as in the original study. When trained on data that matched the software default format, the performance ranged from -11.5 to 17.5%. When using the default training data provided, the performance ranged from -59.4 to -38.5%. Overall, the InSilicoVA predictive accuracy was found to be 11.6-8.2 percentage points lower than that of an alternative algorithm. Additionally, the sensitivity for InSilicoVA was consistently lower than that for an alternative diagnostic algorithm (Tariff 2.0), although the specificity was comparable.Conclusions: The default format and training data provided by the software lead to results that are at best suboptimal, with poor cause-of-death predictive performance. This method is likely to generate erroneous cause of death predictions and, even if properly configured, is not as accurate as alternative automated diagnostic methods.