Scaling up the evaluation of psychotherapy: evaluating motivational interviewing fidelity via statistical text classification

Scaling up the evaluation of psychotherapy: evaluating motivational interviewing fidelity via statistical text classification
复制标题

DOI:
10.1186/1748-5908-9-49
复制
发表时间:
2014-04-24
影响因子:
7.2
通讯作者:
Smyth, Padhraic
Smyth, Padhraic
中科院分区:
医学1区
文献类型:
--
作者:
Atkins, David C.;Steyvers, Mark;Smyth, Padhraic

文献摘要

被引文献

相似文献

背景:行为干预,如心理治疗是领先的,以证据为基础的做法,为各种问题(例如,药物滥用),但对提供者对行为干预的忠诚度的评估受到人类判断需求的限制。目前的研究评估的准确性,统计文本分类在复制人为基础的判断提供者的忠诚度在一个特定的心理治疗动机interviewing(MI).Method:参与者(n = 148)来自五个以前进行的随机试验,无论是在一个安全网医院或大学生的初级保健患者。为了有资格参加最初的研究,参与者符合有问题的药物或酒精使用的标准。所有参与者都接受了一种简短的动机访谈,这是一种针对酒精和物质使用障碍的循证干预。动机访谈技能代码是基于用于评估所有治疗会话的人类评级的MI提供者忠诚度的标准度量。一种被称为标记主题模型的文本分类方法被用来学习基于人类的保真度评级和MI会话成绩单之间的关联。然后,它被用来生成新会话的代码。主要比较的是基于模型的代码与基于人类的codes.Results的准确性:基于模型的代码的受试者工作特征(ROC)分析显示出相当强的灵敏度和特异性与那些从人类raters(范围ROC曲线下面积(AUC)得分:0.62 - 0.81,平均AUC:0.72)。与人类评分员的一致性是根据整个会话的谈话次数和代码计数来评估的。生成的代码具有较高的可靠性与人类代码会话tallies,也有很大的不同,个别codes.Conclusion:要扩大行为干预措施的评价,技术解决方案将是必需的。目前的研究表明,初步的,令人鼓舞的调查结果,统计文本分类在弥合这一方法上的差距的效用。
Background: Behavioral interventions such as psychotherapy are leading, evidence-based practices for a variety of problems (e.g., substance abuse), but the evaluation of provider fidelity to behavioral interventions is limited by the need for human judgment. The current study evaluated the accuracy of statistical text classification in replicating human-based judgments of provider fidelity in one specific psychotherapy-motivational interviewing (MI).Method: Participants (n = 148) came from five previously conducted randomized trials and were either primary care patients at a safety-net hospital or university students. To be eligible for the original studies, participants met criteria for either problematic drug or alcohol use. All participants received a type of brief motivational interview, an evidence-based intervention for alcohol and substance use disorders. The Motivational Interviewing Skills Code is a standard measure of MI provider fidelity based on human ratings that was used to evaluate all therapy sessions. A text classification approach called a labeled topic model was used to learn associations between human-based fidelity ratings and MI session transcripts. It was then used to generate codes for new sessions. The primary comparison was the accuracy of model-based codes with human-based codes.Results: Receiver operating characteristic (ROC) analyses of model-based codes showed reasonably strong sensitivity and specificity with those from human raters (range of area under ROC curve (AUC) scores: 0.62 - 0.81; average AUC: 0.72). Agreement with human raters was evaluated based on talk turns as well as code tallies for an entire session. Generated codes had higher reliability with human codes for session tallies and also varied strongly by individual code.Conclusion: To scale up the evaluation of behavioral interventions, technological solutions will be required. The current study demonstrated preliminary, encouraging findings regarding the utility of statistical text classification in bridging this methodological gap.