Building the Model.

Building the Model.
复制标题

DOI:
10.5858/arpa.2021-0635-ra
复制
发表时间:
2023-07-01
影响因子:
4.6
通讯作者:
Wang F
Wang F
中科院分区:
医学2区
文献类型:
--
作者:
Yang HS;Rhoads DD;Sepulveda J;Zang C;Chadburn A;Wang F

文献摘要

参考文献

相似文献

机器学习(ML)允许分析大量的高维临床实验室数据,从而揭示复杂的模式和趋势。因此,ML可以潜在地提高临床数据解释和实验室医学实践的效率。但是,应认识到产生偏倚或不具代表性的模型的风险,这可能导致误导性的临床结论或高估模型性能。讨论创建ML模型的主要组件,包括数据收集、数据预处理、模型开发和模型评估。我们还强调了开发ML模型的许多挑战和陷阱,这些挑战和陷阱可能会导致误导性的临床印象或不准确的模型性能,并就如何规避这些挑战提供建议和指导。通过检索PubMed数据库、美国食品药品监督管理局的白色论文和指南、会议摘要和在线预印本确定了本综述的参考文献。随着在临床实践中开发和实施ML模型的兴趣越来越大,实验室人员和临床医生需要接受教育,以收集足够大和高质量的数据,正确报告数据集特征,并将来自多个机构的联合收割机数据与适当的标准化相结合。他们还需要评估缺失值的原因,确定是否包含或排除离群值,并评估数据集的完整性。此外,他们需要必要的知识来为特定的临床问题选择合适的ML模型,并根据客观标准准确评估ML模型的性能。特定领域的知识在开发ML模型的整个工作流程中至关重要。
Machine learning (ML) allows for the analysis of massive quantities of high-dimensional clinical laboratory data, thereby revealing complex patterns and trends. Thus, ML can potentially improve the efficiency of clinical data interpretation and the practice of laboratory medicine. However, the risks of generating biased or unrepresentative models, which can lead to misleading clinical conclusions or overestimation of the model performance, should be recognized. To discuss the major components for creating ML models, including data collection, data preprocessing, model development, and model evaluation. We also highlight many of the challenges and pitfalls in developing ML models, which could result in misleading clinical impressions or inaccurate model performance, and provide suggestions and guidance on how to circumvent these challenges. The references for this review were identified through searches of the PubMed database, the US Food and Drug Administration white papers and guidelines, conference abstracts, and online preprints. With the growing interest in developing and implementing ML models in clinical practice, laboratorians and clinicians need to be educated in order to collect sufficiently large and high-quality data, properly report the data set characteristics, and combine data from multiple institutions with proper normalization. They will also need to assess the reasons for missing values, determine the inclusion or exclusion of outliers, and evaluate the completeness of a data set. In addition, they require the necessary knowledge to select a suitable ML model for a specific clinical question and accurately evaluate the performance of the ML model, based on objective criteria. Domain-specific knowledge is critical in the entire workflow of developing ML models.
DOI: 10.1093/clinchem/hvaa168
发表时间: 2020-09-01
期刊: Clinical chemistry
影响因子: 9.3
作者:
Ganetzky RD;Master SR
通讯作者: Master SR