Beyond opinion classification: Extracting facts, opinions and experiences from health forums

Beyond opinion classification: Extracting facts, opinions and experiences from health forums
复制标题

DOI:
10.1371/journal.pone.0209961
复制
发表时间:
2019-01-09
期刊:
影响因子:
3.7
通讯作者:
Plaza, Laura
Plaza, Laura
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Carrillo-de-Albornoz, Jorge;Aker, Ahmet;Plaza, Laura

文献摘要

被引文献

相似文献

前言调查表明,患者,特别是那些患有慢性病的患者,从社交网络和在线论坛中找到的信息中受益匪浅。获取在线健康信息的一个挑战是区分事实信息和更主观的信息。在这项工作中,我们评估了利用文本的词汇、句法、语义、网络和情感属性来使用机器学习算法将患者生成的内容自动分类为三种类型:经验、事实和观点的可行性。在这种情况下,我们的目标是开发自动化方法,使患者、专业人员和研究人员更容易获取和使用在线健康信息。材料和方法我们使用一组3000篇帖子到在线健康论坛,内容涉及乳腺癌、Morbus Crohn和不同的过敏。帖子中的每一句话都被手动贴上了“经验”、“事实”或“观点”的标签。利用这些数据,我们训练了支持向量机算法来执行分类。结果总体而言,我们发现,使用简单的文本表示,如单词嵌入和词袋,可以非常高的准确率(超过80%)预测论坛帖子中包含的信息类型。我们还分析了更复杂的特征,如基于网络属性、词的极性和句子的动词时态的特征,并表明当它们与前面的特征相结合时,可以提高结果。
IntroductionSurveys indicate that patients, particularly those suffering from chronic conditions, strongly benefit from the information found in social networks and online forums. One challenge in accessing online health information is to differentiate between factual and more subjective information. In this work, we evaluate the feasibility of exploiting lexical, syntactic, semantic, network-based and emotional properties of texts to automatically classify patient-generated contents into three types: "experiences", "facts" and "opinions", using machine learning algorithms. In this context, our goal is to develop automatic methods that will make online health information more easily accessible and useful for patients, professionals and researchers.Material and methodsWe work with a set of 3000 posts to online health forums in breast cancer, morbus crohn and different allergies. Each sentence in a post is manually labeled as "experience", "fact" or "opinion". Using this data, we train a support vector machine algorithm to perform classification. The results are evaluated in a 10-fold cross validation procedure.ResultsOverall, we find that it is possible to predict the type of information contained in a forum post with a very high accuracy (over 80 percent) using simple text representations such as word embeddings and bags of words. We also analyze more complex features such as those based on the network properties, the polarity of words and the verbal tense of the sentences and show that, when combined with the previous ones, they can boost the results.