课题基金 / 基金详情

Multi-modal unsupervised embeddings to advance machine learning in healthcare

Multi-modal unsupervised embeddings to advance machine learning in healthcare
多模式无监督嵌入促进医疗保健领域的机器学习
批准号:
10445794
负责人:
Thomas J Fuchs
金额:
$35.91万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-04-01 至 2026-01-31

项目摘要

项目成果

Thomas J Fuchs的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY Integrating high-dimensional and heterogenous biomedical data, such as electronic health records (EHRs), molecular data, imaging, and free text, is a key challenge for making robust discoveries that transform healthcare. Current work in the literature commonly analyze biomedical data types separately, focus on small disease-related cohorts of patients, and rely on domain experts and manual clinical feature selection in an ad hoc manner. Although appropriate in some situations, supervised definitions of the feature space scale poorly, do not generalize well, include inherent biases, and miss opportunities to discover novel patterns and features. To address these issues, we will develop novel methods based on unsupervised machine learning to derive low-dimensional vector-based representations, i.e., “embeddings”, of medical concepts and patient clinical histories from large- scale, multi-modal and domain-free biomedical datasets. These pre-computed representations aim to overcome common biases due to population, supervised labeling, and specific hospital operation processes. These multi-modal embeddings can be fine-tuned and applied to a number of specific predictive tasks, improving scalability, generalizability and effectiveness of machine learning models in healthcare. In particular, we will first develop methods based on unsupervised learning to create multi-modal embeddings of medical concepts using heterogeneous EHRs, linked biobanks and electrocardiogram waveform data, from the diverse population of five hospitals within the Mount Sinai Health System in New York, NY, and publicly available medical knowledge. We will then create a scalable framework to compute unsupervised multi-modal embeddings that can summarize patient clinical histories and lead to subtyping and patient stratification. We will also develop a federated learning system to share, visualize, and combine embeddings generated separately at different medical institutes to capture a larger and more diverse population and clinical landscape. We will apply embeddings to advance methods for EHR-based disease phenotyping, onset prediction, and subtyping. While tested on EHRs, genetic and waveform data from linked repositories, and medical knowledge, the proposed approaches will be easily extendable to include other data, such as clinical images. This project will represent a step towards the next generation of ML in healthcare ML that can (i) scale to billions of patients, (ii) embed complex relationships of multi-modal data, and (iii) create less biased disease representations by securely learning from patients across institutions via federated learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-modal unsupervised embeddings to advance machine learning in healthcare
海外基金