Machine learning algorithms outperform conventional regression models in predicting development of hepatocellular carcinoma.

Machine learning algorithms outperform conventional regression models in predicting development of hepatocellular carcinoma.
复制标题

DOI:
10.1038/ajg.2013.332
复制
发表时间:
2013-11
期刊:
The American journal of gastroenterology
影响因子:
--
通讯作者:
Waljee AK
Waljee AK
中科院分区:
其他
文献类型:
--
作者:
Singal AG;Mukherjee A;Elmunzer BJ;Higgins PD;Lok AS;Zhu J;Marrero JA;Waljee AK

文献摘要

被引文献

相似文献

肝细胞癌(HCC)的预测模型一直受到准确性不高和缺乏验证的限制。机器学习算法提供了一种新的方法,可以改善肝硬化患者的HCC风险预测。我们研究的目的是使用传统的回归分析和机器学习算法,开发和比较肝癌患者发展的预测模型。我们在2004年1月至2006年9月期间在密歇根大学招募了442例Child A或B肝硬化患者(UM队列),并对他们进行前瞻性随访,直至发生HCC、肝移植、死亡或研究终止。回归分析和机器学习算法用于构建HCC发展的预测模型,这些模型在丙型肝炎抗病毒长期治疗肝硬化(HALT-C)试验的独立验证队列中进行了测试。这两种模型也与先前发表的HALT-C模型进行了比较。使用受试者工作特征曲线分析评估辨别力,使用净重新分类改善和综合辨别力改善统计评估诊断准确性。在中位随访3.5年后,41例患者发生了HCC。在验证队列中,UM回归模型的c统计量为0.61(95%CI 0.56-0.67),而机器学习算法的c统计量为0.64(95%CI 0.60-0.69)。机器学习算法具有显著更好的诊断准确性,如通过净重新分类改善(p<0.001)和综合辨别改善(p=0.04)所评估的。HALT-C模型在验证队列中的c统计量为0.60(95%CI 0.50-0.70),并且优于机器学习算法(p=0.047)。机器学习算法提高了对肝硬化患者进行风险分层的准确性,并可用于准确识别患有HCC的高风险患者。
Predictive models for hepatocellular carcinoma (HCC) have been limited by modest accuracy and lack of validation. Machine learning algorithms offer a novel methodology, which may improve HCC risk prognostication among patients with cirrhosis. Our study's aim was to develop and compare predictive models for HCC development among cirrhotic patients, using conventional regression analysis and machine learning algorithms. We enrolled 442 patients with Child A or B cirrhosis at the University of Michigan between January 2004 and September 2006 (UM cohort) and prospectively followed them until HCC development, liver transplantation, death, or study termination. Regression analysis and machine learning algorithms were used to construct predictive models for HCC development, which were tested on an independent validation cohort from the Hepatitis C Antiviral Long-term Treatment against Cirrhosis (HALT-C) Trial. Both models were also compared to the previously published HALT-C model. Discrimination was assessed using receiver operating characteristic curve analysis and diagnostic accuracy was assessed with net reclassification improvement and integrated discrimination improvement statistics. After a median follow-up of 3.5 years, 41 patients developed HCC. The UM regression model had a c-statistic of 0.61 (95%CI 0.56-0.67), whereas the machine learning algorithm had a c-statistic of 0.64 (95%CI 0.60–0.69) in the validation cohort. The machine learning algorithm had significantly better diagnostic accuracy as assessed by net reclassification improvement (p<0.001) and integrated discrimination improvement (p=0.04). The HALT-C model had a c-statistic of 0.60 (95%CI 0.50-0.70) in the validation cohort and was outperformed by the machine learning algorithm (p=0.047). Machine learning algorithms improve the accuracy of risk stratifying patients with cirrhosis and can be used to accurately identify patients at high-risk for developing HCC.