A Multicenter, Scan-Rescan, Human and Machine Learning CMR Study to Test Generalizability and Precision in Imaging Biomarker Analysis

A Multicenter, Scan-Rescan, Human and Machine Learning CMR Study to Test Generalizability and Precision in Imaging Biomarker Analysis
复制标题

DOI:
10.1161/circimaging.119.009214
复制
发表时间:
2019-10-01
影响因子:
7.5
通讯作者:
Manisty, Charlotte H.
Manisty, Charlotte H.
中科院分区:
医学1区
文献类型:
--
作者:
Bhuva, Anish N.;Bai, Wenjia;Manisty, Charlotte H.

文献摘要

被引文献

相似文献

背景技术背景:使用机器学习(ML)对心脏结构和功能进行自动分析具有巨大的潜力,但目前普遍性较差。传统上,比较是针对临床医生作为参考,忽略了固有的人类观察者间和观察者内误差,并确保ML不能证明优越性。测量精度(扫描:再扫描再现性)解决了这一问题。我们比较了ML和人类的精度使用多中心,多疾病,扫描:再扫描心血管磁共振dataset.METHODS:110例患者(5种疾病类别,5个机构,2个扫描仪制造商,2个场强)进行扫描:再扫描心血管磁共振(96%在一周内)。在确定了最精确的人类技术后,由一名专家、一名训练有素的初级临床医生和一个在599个独立的多中心疾病病例上训练的全自动卷积神经网络测量了左心室腔容积、质量和射血分数。扫描:再扫描的变异系数和1000自举95%CI计算和比较使用混合线性effects models.RESULTS:临床医生可以有信心地检测到左心室射血分数的9%的变化,超过一半的变异系数归因于观察者内变异。专家、受过培训的初级和自动扫描:再扫描精度相似(左心室射血分数,变异系数6.1 [5.2%-7.1%],P=0.2581; 8.3 [5.6%-10.3%],P=0.3653; 8.8 [6.1%-11.1%],P=0.8620)。自动化分析比人类快186倍(0.07 vs 13分钟)。结论:自动化ML分析速度更快,精度与最精确的人类技术相似,即使在现实世界的扫描:重新扫描数据的挑战下也是如此。多中心、多供应商、多场强扫描的评估:重新扫描数据(可在www.thevolumesresource.com上获得)允许对ML精密度进行概括性评估,并可促进ML直接转化为临床实践。
BACKGROUND: Automated analysis of cardiac structure and function using machine learning (ML) has great potential, but is currently hindered by poor generalizability. Comparison is traditionally against clinicians as a reference, ignoring inherent human inter- and intraobserver error, and ensuring that ML cannot demonstrate superiority. Measuring precision (scan:rescan reproducibility) addresses this. We compared precision of ML and humans using a multicenter, multi-disease, scan:rescan cardiovascular magnetic resonance data set.METHODS: One hundred ten patients (5 disease categories, 5 institutions, 2 scanner manufacturers, and 2 field strengths) underwent scan:rescan cardiovascular magnetic resonance (96% within one week). After identification of the most precise human technique, left ventricular chamber volumes, mass, and ejection fraction were measured by an expert, a trained junior clinician, and a fully automated convolutional neural network trained on 599 independent multicenter disease cases. Scan:rescan coefficient of variation and 1000 bootstrapped 95% CIs were calculated and compared using mixed linear effects models.RESULTS: Clinicians can be confident in detecting a 9% change in left ventricular ejection fraction, with greater than half of coefficient of variation attributable to intraobserver variation. Expert, trained junior, and automated scan:rescan precision were similar (for left ventricular ejection fraction, coefficient of variation 6.1 [5.2%-7.1%], P=0.2581; 8.3 [5.6%-10.3%], P=0.3653; 8.8 [6.1%-11.1%], P=0.8620). Automated analysis was 186x faster than humans (0.07 versus 13 minutes).CONCLUSIONS: Automated ML analysis is faster with similar precision to the most precise human techniques, even when challenged with real-world scan:rescan data. Assessment of multicenter, multi-vendor, multi-field strength scan:rescan data (available at www.thevolumesresource.com) permits a generalizable assessment of ML precision and may facilitate direct translation of ML to clinical practice.