Predicting Cardiovascular Risk Using Social Media Data: Performance Evaluation of Machine-Learning Models.

Predicting Cardiovascular Risk Using Social Media Data: Performance Evaluation of Machine-Learning Models.
复制标题

DOI:
10.2196/24473
复制
发表时间:
2021-02-19
期刊:
影响因子:
--
通讯作者:
Merchant RM
Merchant RM
中科院分区:
其他
文献类型:
--
作者:
Andy AU;Guntuku SC;Adusumalli S;Asch DA;Groeneveld PW;Ungar LH;Merchant RM

文献摘要

被引文献

相似文献

目前的动脉粥样硬化性心血管疾病(ASCVD)预测模型具有局限性;因此,正在努力提高ASCVD模型的区分能力。我们试图评估社交媒体帖子与集合队列风险方程(PCE)相比在预测ASCVD 10年风险方面的区分性能力。我们同意在城市学术急诊科接受治疗的患者分享他们在Facebook上的帖子和电子医疗记录(EMR)。我们检索了所有同意研究的患者在研究登记前5年内的Facebook状态更新。我们确定了181名患者(N=181),他们的EMR中没有冠心病病史,ASCVD评分,他们在Facebook上的帖子超过200个单词。使用这些患者在Facebook上的帖子,我们应用了机器学习模型来预测10年内ASCVD的风险评分。使用机器学习模型和心理语言学词典、语言查询和字数统计,我们评估了仅来自帖子的语言是否可以分别预测风险分数的差异和某些单词与风险类别的关联。机器学习模型预测了<lt;5%,5%-7.4%,7.5%-9.9%和≥10%类别的10年内ASCVD风险评分,曲线下面积(AUC值)分别为0.78,0.57,0.72和0.61。机器学习模型区分低风险(&lt;10%)和高风险(&gt;10%),AUC为0.69。此外,机器学习模型预测了皮尔逊r=0.26的ASCVD风险评分。通过语言查询和字数统计,ASCVD得分较高的患者更有可能使用与悲伤相关的单词(r=0.32)。社交媒体上使用的语言可以提供对个人ASCVD风险的洞察,并为风险调整提供方法。
Current atherosclerotic cardiovascular disease (ASCVD) predictive models have limitations; thus, efforts are underway to improve the discriminatory power of ASCVD models. We sought to evaluate the discriminatory power of social media posts to predict the 10-year risk for ASCVD as compared to that of pooled cohort risk equations (PCEs). We consented patients receiving care in an urban academic emergency department to share access to their Facebook posts and electronic medical records (EMRs). We retrieved Facebook status updates up to 5 years prior to study enrollment for all consenting patients. We identified patients (N=181) without a prior history of coronary heart disease, an ASCVD score in their EMR, and more than 200 words in their Facebook posts. Using Facebook posts from these patients, we applied a machine-learning model to predict 10-year ASCVD risk scores. Using a machine-learning model and a psycholinguistic dictionary, Linguistic Inquiry and Word Count, we evaluated if language from posts alone could predict differences in risk scores and the association of certain words with risk categories, respectively. The machine-learning model predicted the 10-year ASCVD risk scores for the categories <5%, 5%-7.4%, 7.5%-9.9%, and ≥10% with area under the curve (AUC) values of 0.78, 0.57, 0.72, and 0.61, respectively. The machine-learning model distinguished between low risk (<10%) and high risk (>10%) with an AUC of 0.69. Additionally, the machine-learning model predicted the ASCVD risk score with Pearson r=0.26. Using Linguistic Inquiry and Word Count, patients with higher ASCVD scores were more likely to use words associated with sadness (r=0.32). Language used on social media can provide insights about an individual’s ASCVD risk and inform approaches to risk modification.