Sleep apnea phenotyping and relationship to disease in a large clinical biobank.
Sleep apnea phenotyping and relationship to disease in a large clinical biobank.
复制标题
DOI:
10.1093/jamiaopen/ooab117
复制
发表时间:
2022-04
期刊:
影响因子:
2.1
通讯作者:
Karlson EW
中科院分区:
文献类型:
--
作者:
Cade BE;Hassan SM;Dashti HS;Kiernan M;Pavlova MK;Redline S;Karlson EW
Sleep apnea is associated with a broad range of pathophysiology. While electronic health record (EHR) information has the potential for revealing relationships between sleep apnea and associated risk factors and outcomes, practical challenges hinder its use. Our objectives were to develop a sleep apnea phenotyping algorithm that improves the precision of EHR case/control information using natural language processing (NLP); identify novel associations between sleep apnea and comorbidities in a large clinical biobank; and investigate the relationship between polysomnography statistics and comorbid disease using NLP phenotyping. We performed clinical chart reviews on 300 participants putatively diagnosed with sleep apnea and applied International Classification of Sleep Disorders criteria to classify true cases and noncases. We evaluated 2 NLP and diagnosis code-only methods for their abilities to maximize phenotyping precision. The lead algorithm was used to identify incident and cross-sectional associations between sleep apnea and common comorbidities using 4876 NLP-defined sleep apnea cases and 3× matched controls. The optimal NLP phenotyping strategy had improved model precision (≥0.943) compared to the use of one diagnosis code (≤0.733). Of the tested diseases, 170 disorders had significant incidence odds ratios (ORs) between cases and controls, 8 of which were confirmed using polysomnography (n = 4544), and 281 disorders had significant prevalence OR between sleep apnea cases versus controls, 41 of which were confirmed using polysomnography data. An NLP-informed algorithm can improve the accuracy of case-control sleep apnea ascertainment and thus improve the performance of phenome-wide, genetic, and other EHR analyses of a highly prevalent disorder. Sleep apnea is a common disease in which breathing partially or completely pauses during sleep, leading to less oxygen in the blood, repeated awakenings, and increased risk of developing multiple diseases. Current studies of sleep apnea often have relatively few participants due to the challenge of performing overnight sleep recordings. Electronic health record (EHR) billing code diagnoses of sleep apnea could be repurposed to increase the size of research studies, but the accuracy of the diagnoses is reduced. We developed a reusable algorithm that improves the accuracy of EHR sleep apnea diagnoses using natural language processing to extract information from clinical notes. As a proof of concept, we used the algorithm to identify hundreds of diseases that are increased among participants with sleep apnea compared to similar patients without sleep apnea. Many of these disease relationships with sleep apnea have not been previously recognized. This improved algorithm will help to accelerate future large-scale investigations of the causes and consequences of sleep apnea.
登录
查看更多内容
影响因子:
3.7
作者:
Liao KP;Ananthakrishnan AN;Kumar V;Xia Z;Cagan A;Gainer VS;Goryachev S;Chen P;Savova GK;Agniel D;Churchill S;Lee J;Murphy SN;Plenge RM;Szolovits P;Kohane I;Shaw SY;Karlson EW;Cai T
通讯作者:
Cai T
影响因子:
46.9
作者:
通讯作者:
--
影响因子:
4.3
作者:
Keenan, Brendan T.;Kirchner, H. Lester;Derose, Stephen F.
通讯作者:
Derose, Stephen F.
影响因子:
3.5
作者:
Gellen, Barnabas;Canoui-Poitrine, Florence;Damy, Thibaud
通讯作者:
Damy, Thibaud
影响因子:
--
作者:
Karlson, Elizabeth W.;Boutin, Natalie T.;Allen, Nicole L.
通讯作者:
Allen, Nicole L.