A broadly applicable approach to enrich electronic-health-record cohorts by identifying patients with complete data: a multisite evaluation.
A broadly applicable approach to enrich electronic-health-record cohorts by identifying patients with complete data: a multisite evaluation.
复制标题
DOI:
10.1093/jamia/ocad166
复制
发表时间:
2023-11-17
影响因子:
6.4
通讯作者:
Murphy, Shawn N.
中科院分区:
文献类型:
--
作者:
Klann, Jeffrey G.;Henderson, Darren W.;Morris, Michele;Estiri, Hossein;Weber, Griffin M.;Visweswaran, Shyam;Murphy, Shawn N.
关键词:
Patients who receive most care within a single healthcare system (colloquially called a “loyalty cohort” since they typically return to the same providers) have mostly complete data within that organization’s electronic health record (EHR). Loyalty cohorts have low data missingness, which can unintentionally bias research results. Using proxies of routine care and healthcare utilization metrics, we compute a per-patient score that identifies a loyalty cohort. We implemented a computable program for the widely adopted i2b2 platform that identifies loyalty cohorts in EHRs based on a machine-learning model, which was previously validated using linked claims data. We developed a novel validation approach, which tests, using only EHR data, whether patients returned to the same healthcare system after the training period. We evaluated these tools at 3 institutions using data from 2017 to 2019. Loyalty cohort calculations to identify patients who returned during a 1-year follow-up yielded a mean area under the receiver operating characteristic curve of 0.77 using the original model and 0.80 after calibrating the model at individual sites. Factors such as multiple medications or visits contributed significantly at all sites. Screening tests’ contributions (eg, colonoscopy) varied across sites, likely due to coding and population differences. This open-source implementation of a “loyalty score” algorithm had good predictive power. Enriching research cohorts by utilizing these low-missingness patients is a way to obtain the data completeness necessary for accurate causal analysis. i2b2 sites can use this approach to select cohorts with mostly complete EHR data.
登录
查看更多内容
影响因子:
7.4
作者:
Kohane IS;Aronow BJ;Avillach P;Beaulieu-Jones BK;Bellazzi R;Bradford RL;Brat GA;Cannataro M;Cimino JJ;García-Barrio N;Gehlenborg N;Ghassemi M;Gutiérrez-Sacristán A;Hanauer DA;Holmes JH;Hong C;Klann JG;Loh NHW;Luo Y;Mandl KD;Daniar M;Moore JH;Murphy SN;Neuraz A;Ngiam KY;Omenn GS;Palmer N;Patel LP;Pedrera-Jiménez M;Sliz P;South AM;Tan ALM;Taylor DM;Taylor BW;Torti C;Vallejos AK;Wagholikar KB;Consortium For Clinical Characterization Of COVID-19 By EHR (4CE);Weber GM;Cai T
通讯作者:
Cai T
DOI:
10.13063/2327-9214.1074
发表时间:
2014
期刊:
EGEMS (Washington, DC)
影响因子:
--
作者:
Murphy S;Wilcox A
通讯作者:
Wilcox A
DOI:
10.1056/nejmsr1809937
发表时间:
2019-08-15
期刊:
The New England journal of medicine
影响因子:
--
作者:
All of Us Research Program Investigators;Denny JC;Rutter JL;Goldstein DB;Philippakis A;Smoller JW;Jenkins G;Dishman E
通讯作者:
Dishman E
影响因子:
4
作者:
Gianfrancesco MA;Goldstein ND
通讯作者:
Goldstein ND
影响因子:
15.2
作者:
Brat, Gabriel A.;Weber, Griffin M.;Kohane, Isaac S.
通讯作者:
Kohane, Isaac S.