A roadmap for semi-automatically extracting predictive and clinically meaningful temporal features from medical data for predictive modeling.

A roadmap for semi-automatically extracting predictive and clinically meaningful temporal features from medical data for predictive modeling.
复制标题

DOI:
10.1016/j.glt.2018.11.001
复制
发表时间:
2019-01-01
期刊:
影响因子:
--
通讯作者:
Luo, Gang
Luo, Gang
中科院分区:
其他
文献类型:
--
作者:
Luo, Gang

文献摘要

被引文献

相似文献

基于机器学习的医疗数据预测建模在改善医疗保健和降低成本方面具有巨大的潜力。然而,除其他外,有两个障碍阻碍了它在医疗保健中的广泛采用。首先,医学数据本质上是纵向的。对它们进行前处理,特别是对于特征工程,是劳动密集型的,通常需要50%-80%的模型构建工作。预测时间特征是建立精确模型的基础,但很难识别。这是有问题的。医疗保健系统用于建模的资源有限,而不准确的模型会产生不太理想的结果,而且往往毫无用处。其次,大多数机器学习模型没有提供对其预测结果的解释。然而,提供这样的解释对于模型在通常的临床实践中使用是必不可少的。为了解决这两个障碍,本文概述了:1)数据驱动的方法,用于从医疗数据中半自动地提取预测性和临床意义的时间特征用于预测建模;2)使用这些特征来自动解释机器学习的预测结果并建议定制的干预措施。这为未来的研究提供了路线图。
Predictive modeling based on machine learning with medical data has great potential to improve healthcare and reduce costs. However, two hurdles, among others, impede its widespread adoption in hdealthcare. First, medical data are by nature longitudinal. Pre-processing them, particularly for feature engineering, is labor intensive and often takes 50-80% of the model building effort. Predictive temporal features are the basis of building accurate models, but are difficult to identify. This is problematic. Healthcare systems have limited resources for model building, while inaccurate models produce sub-optimal outcomes and are often useless. Second, most machine learning models provide no explanation of their prediction results. However, offering such explanations is essential for a model to be used in usual clinical practice. To address these two hurdles, this paper outlines: 1) a data-driven method for semi-automatically extracting predictive and clinically meaningful temporal features from medical data for predictive modeling; and 2) a method of using these features to automatically explain machine learning prediction results and suggest tailored interventions. This provides a roadmap for future research.