LATTE: Label-efficient incident phenotyping from longitudinal electronic health records.
LATTE: Label-efficient incident phenotyping from longitudinal electronic health records.
复制标题
DOI:
10.1016/j.patter.2023.100906
复制
发表时间:
2024-01-12
期刊:
影响因子:
6.5
通讯作者:
Cai, Tianxi
中科院分区:
文献类型:
--
作者:
Wen, Jun;Hou, Jue;Bonzel, Clara -Lea;Zhao, Yihan;Castro, Victor M.;Gainer, Vivian S.;Weisenfeld, Dana;Cai, Tianrun;Ho, Yuk-Lam;Panickan, Vidul A.;Costa, Lauren;Hong, Chuan;Gaziano, J. Michael;Liao, Katherine P.;Lu, Junwei;Cho, Kelly;Cai, Tianxi
Electronic health record (EHR) data are increasingly used to support real-world evidence studies but are limited by the lack of precise timings of clinical events. Here, we propose a label-efficient incident phenotyping (LATTE) algorithm to accurately annotate the timing of clinical events from longitudinal EHR data. By leveraging the pre-trained semantic embeddings, LATTE selects predictive features and compresses their information into longitudinal visit embeddings through visit attention learning. LATTE models the sequential dependency between the target event and visit embeddings to derive the timings. To improve label efficiency, LATTE constructs longitudinal silver-standard labels from unlabeled patients to perform semi-supervised training. LATTE is evaluated on the onset of type 2 diabetes, heart failure, and relapses of multiple sclerosis. LATTE consistently achieves substantial improvements over benchmark methods while providing high prediction interpretability. The event timings are shown to help discover risk factors of heart failure among patients with rheumatoid arthritis. An incident phenotyping method is proposed to identify timings of clinical events We achieve label efficiency by exploiting EHR embeddings and predictive surrogates Model is validated on incident type 2 diabetes, heart failure, and multiple sclerosis Results facilitate assessing cardiac risks among patients with rheumatoid arthritis Electronic health record (EHR) data collected during routine clinical care are increasingly used by translational and clinical researchers to address a variety of questions, such as identifying associations between diseases or phenotypes, predicting disease risk or prognosis, or supporting the safety and efficacy of treatments. The feasibility of these studies relies on precisely inferring the timing and ordering of clinical events from EHR data to define baseline eligibility and patient outcomes. Rule-based extraction methods can be inaccurate, and existing machine-learning approaches generally require large-scale labels for training. Better methods for identifying the timing of clinical events could help expand the use of EHR data to address important medical questions and could improve the quality of the resulting analyses. A label-efficient method, LATTE, is proposed to identify the timings of clinical events from longitudinal electronic health records. It achieves significantly improved performance in identifying the onset of type 2 diabetes, heart failure, and relapses of multiple sclerosis. LATTE has strong cross-site portability and is highly interpretable by indicating the important features and visits that drive the predictions.
登录
查看更多内容
影响因子:
8.4
作者:
Eisenhauer, E. A.;Therasse, P.;Verweij, J.
通讯作者:
Verweij, J.
影响因子:
13.8
作者:
通讯作者:
--
影响因子:
1.9
作者:
Hou, Jue;Chan, Stephanie F.;Wang, Xuan;Cai, Tianxi
通讯作者:
Cai, Tianxi
DOI:
10.1093/jamia/ocaa079
发表时间:
2020-08-01
影响因子:
6.4
作者:
Ahuja, Yuri;Zhou, Doudou;Cai, Tianxi
通讯作者:
Cai, Tianxi
影响因子:
3
作者:
Hassett MJ;Uno H;Cronin AM;Carroll NM;Hornbrook MC;Ritzwoller D
通讯作者:
Ritzwoller D