Scalable Biomedical Pattern Recognition Via Deep Learning
Scalable Biomedical Pattern Recognition Via Deep Learning
批准号:
9302040
负责人:
Thomas Lasko
金额:
$0.77万
依托单位国家:
美国
项目类别:
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-07-01 至 2016-06-30
中文摘要
描述(由申请人提供):从电子医疗记录(EMRS)和其他生物医学数据集中提取的模式可以为学习医疗系统提供有价值的反馈,但我们找到它们的能力受到某些手动步骤的限制。寻找模式的主要方法是使用监督学习,即计算算法在为结果变量(或标签)建模的输入变量(或特征)中搜索模式。这通常需要专家指定学习任务,构建输入特征,并准备结果标签。几十年来,这种工作流程一直很好地为我们服务,但对人类努力的依赖阻碍了它的扩展,而且它错过了最具信息量的模式,而从定义上讲,这些模式几乎是没有人预料到的。它不太适合于人口规模数据的新兴时代,在这个时代,我们可以设想大量的新任务,如监测所有新出现的疾病,检测所有意想不到的药物效应,或推断所有基因变异的完整临床表型。
无监督特征学习方法克服了这些限制,通过在海量的、几乎没有人参与的未标记数据集中识别有意义的模式来克服这些限制。虽然有大量关于特征创建的文献,但对非监督方法的兴趣出现了新的热潮
这是由深度学习的最新发展推动的,在深度学习中,表达特征的紧凑层次是从大型未标记数据集中学习的。在图像和语音识别领域,深度学习已经产生了达到或超过(高达70%)困难标准化任务的先前技术水平的特征。
不幸的是,EMR中通常存在的噪声、稀疏和不规则的数据不是深度学习的不良基础。我们的方法使用高斯过程回归将这种不规则的观测序列转换为适合于深层体系结构使用的纵向概率密度。通过这种方法,我们可以学习连续的非监督特征,这些特征捕捉到稀疏和不规则观测的纵向结构。在我们的初步结果中,在意想不到的分类任务中,非监督特征与由完全了解领域、分类任务和类别标签的专家设计的黄金标准特征一样强大(0.96AUC)。
在这个项目中,我们将学习非监督功能,用于识别EMR图像中所有个人的记录,用于与1型或2型糖尿病相关的100项实验室测试和200种药物的每一项。我们将使用三个特征学习算法未知的模式识别任务来评估这些特征:1)区分糖尿病患者和非糖尿病患者的简单监督分类任务,2)区分1型糖尿病患者和2型糖尿病患者的更困难的任务,以及3)将这些特征视为微表型并测量它们与29种不同单核苷酸多态之间的关联的遗传关联任务,这些单核苷酸多态与1型或2型糖尿病存在已知关联。
英文摘要
DESCRIPTION (provided by applicant): Patterns extracted from Electronic Medical Records (EMRs) and other biomedical datasets can provide valuable feedback to a learning healthcare system, but our ability to find them is limited by certain manual steps. The dominant approach to finding the patterns uses supervised learning, where a computational algorithm searches for patterns among input variables (or features) that model an outcome variable (or label). This usually requires an expert to specify the learning task, construct input features, and prepare the outcome labels. This workflow has served us well for decades, but the dependence on human effort prevents it from scaling and it misses the most informative patterns, which are almost by definition the ones that nobody anticipates. It is poorly suited to the emerging era of population-scale data, in which we can conceive of massive new undertakings such as surveiling for all emerging diseases, detecting all unanticipated medication effects, or inferring the complete clinical phenotype of all genetic variants.
The approach of unsupervised feature learning overcomes these limitations by identifying meaningful patterns in massive, unlabeled datasets with little or no human involvement. While there is a large literature on feature creation, a new surge of interest in unsupervised methods is
being driven by the recent development of deep learning, in which a compact hierarchy of expressive features is learned from large unlabeled datasets. In the domains of image and speech recognition, deep learning has produced features that meet or exceed (by as much as 70%) the previous state of the art on difficult standardized tasks.
Unfortunately, the noisy, sparse, and irregular data typically found in an EMR is a poor substrate for deep learning. Our approach uses Gaussian process regression to convert such an irregular sequence of observations into a longitudinal probability density that is suitable for use with a deep architecture. With this approach, we can learn continuous unsupervised features that capture the longitudinal structure of sparse and irregular observations. In our preliminary results unsupervised features were as powerful (0.96 AUC) in an unanticipated classification task as gold-standard features engineered by an expert with full knowledge of the domain, the classification task, and the class labels.
In this project we will learn unsupervised features for records of all individuals in our deidentifed EMR image, for each of 100 laboratory tests and 200 medications of relevance to type 1 or type 2 diabetes. We will evaluate the features using three pattern recognition tasks that were unknown to the feature-learning algorithm: 1) an easy supervised classification task of distinguishing diabetics vs. nondiabetics, 2) a much more difficult task of distinguishing type 1 vs. type 2 diabetics, and 3) a genetic association task that considers the features as micro-phenotypes and measures their association with 29 different single nucleotide polymorphisms with known associations to type 1 or type 2 diabetes.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Nonstationary Gaussian Process Regression for Evaluating Clinical Laboratory Test Sampling Strategies.
用于评估临床实验室测试采样策略的非平稳高斯过程回归。
DOI:
--
发表时间:
2015
期刊:
Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
[Lasko,ThomasA]
通讯作者:
Lasko,ThomasA
Efficient Inference of Gaussian-Process-Modulated Renewal Processes with Application to Medical Event Data.
高斯过程调制更新过程的有效推理及其应用于医疗事件数据。
DOI:
--
发表时间:
2014
期刊:
Uncertainty in artificial intelligence : proceedings of the ... conference. Conference on Uncertainty in Artificial Intelligence
影响因子:
--
作者:
[Lasko,ThomasA]
通讯作者:
Lasko,ThomasA
Data-Driven Guidance for Timing Repeated Inpatient Laboratory Tests
-
批准号:10450243
-
项目类别:
-
资助金额:$22.58万
-
财政年份:2022
-
负责人:Thomas Lasko
-
依托单位:
Data-Driven Guidance for Timing Repeated Inpatient Laboratory Tests
-
批准号:10599337
-
项目类别:
-
资助金额:$18.81万
-
财政年份:2022
-
负责人:Thomas Lasko
-
依托单位:
Identification, Extraction and Display of Clinical Data Patterns with Application to Anesthesia Workflows
-
批准号:9248768
-
项目类别:
-
资助金额:$35.17万
-
财政年份:2016
-
负责人:Thomas Lasko
-
依托单位:
Identification, Extraction and Display of Clinical Data Patterns with Application to Anesthesia Workflows
-
批准号:9051683
-
项目类别:
-
资助金额:$7.05万
-
财政年份:2016
-
负责人:Thomas Lasko
-
依托单位:
Identification, Extraction and Display of Clinical Data Patterns with Application to Anesthesia Workflows
-
批准号:9420613
-
项目类别:
-
资助金额:$40.68万
-
财政年份:2016
-
负责人:Thomas Lasko
-
依托单位:
Scalable Biomedical Pattern Recognition Via Deep Learning
-
批准号:8689173
-
项目类别:
-
资助金额:$20.27万
-
财政年份:2013
-
负责人:Thomas Lasko
-
依托单位:
海外基金