PREDOSE: a semantic web platform for drug abuse epidemiology using social media.

PREDOSE: a semantic web platform for drug abuse epidemiology using social media.
复制标题

Predose:使用社交媒体流行病学的语义网络平台。

DOI:
10.1016/j.jbi.2013.07.007
复制
发表时间:
2013-12
影响因子:
4.5
通讯作者:
Falck, Russel
Falck, Russel
中科院分区:
医学3区
文献类型:
--
作者:
Cameron, Delroy;Smith, Gary A.;Daniulaityte, Raminta;Sheth, Amit P.;Dave, Drashti;Chen, Lu;Anand, Gaurish;Carlson, Robert;Watkins, Kera Z.;Falck, Russel

文献摘要

参考文献

被引文献

相似文献

近年来,社交媒体在生物医学知识挖掘中的作用,包括临床,医疗和保健信息学,处方药滥用流行病学和药物药理学,变得越来越重要。社交媒体为人们提供了在在线社区自由分享意见和经验的机会,这可能会提供超出领域专业人员知识范围的信息。本文介绍了一种新的语义Web平台PREDOSE(处方药物滥用在线监测和流行病学),其目的是促进使用社交媒体的处方(及相关)药物滥用行为的流行病学研究的发展。PREDOSE使用Web论坛帖子和领域知识,在手动创建的药物滥用本体(DAO)(发音为dow)中建模,以促进从用户生成内容(UGC)中提取语义信息。结合词汇、模式和语义等技术,结合领域知识,从用户生成内容中提取细粒度的语义信息。在以前的一项研究中,PREDOSE被用来获得数据集,从中获得药物滥用研究的新知识。在这里,我们报告了各种平台增强功能,包括更新的DAO,用于关系和三元组提取的新组件,以及用于内容分析,趋势检测和新兴模式探索的工具,这些都增强了PREDOSE平台的功能。鉴于这些改进,PREDOSE现在更有能力通过减轻传统的劳动密集型内容分析任务来影响药物滥用研究。使用自定义网络爬虫从公开的网络论坛中抓取UGC,PREDOSE首先自动收集基于网络的社交媒体内容,以便进行后续的语义注释。注释方案在DAO中建模,并且包括特定于领域的知识,例如处方(和相关)药物、制备方法、副作用、给药途径等。DAO还用于帮助识别三种类型的数据,即:1)实体、2)关系和3)三元组。PREDOSE然后使用基于词汇和语义的技术的组合来从抓取的内容中提取实体和关系,以及使用DAO中表达的模式进行三重提取的自顶向下方法。此外,PREDOSE使用公开可用的词典来识别文本中的初始情感表达,然后使用概率优化算法(来自相关研究)来提取最终的情感表达。这些技术结合在一起,可以从用户生成内容中捕获精细的语义信息,并对与处方药滥用相关的社交媒体进行查询、搜索、趋势分析和整体内容分析。此外,提取的数据也提供给领域专家,用于创建培训和测试集,用于评估和改进信息提取技术。最近对PREDOSE平台中应用的信息提取技术进行的评估表明,在手动创建的黄金标准数据集上,实体识别的准确率为85%,召回率为72%。在另一项研究中,通过领域专家的手动评估,PREDOSE在关系识别方面达到了36%的精确度,在三重提取方面达到了33%的精确度。鉴于关系和三重提取任务的复杂性以及社交媒体文本的深奥性质,我们将这些解释为有利的初步结果。提取的语义信息目前正用于在线发现支持系统,由赖特州立大学干预、治疗和成瘾研究中心(CITAR)的处方药滥用研究人员使用。一个全面的平台,实体,关系,三元组和情感提取从这样的深奥的文本从来没有开发用于药物滥用研究。PREDOSE已经证明了挖掘社交媒体的重要性,它提供的数据揭示了药物滥用研究的新发现。鉴于最近对平台进行了改进,包括改进了DAO、关系和三重提取的组成部分以及内容、趋势和新出现的模式分析工具,预计PREDOSE将在今后推动药物滥用流行病学方面发挥重要作用。
The role of social media in biomedical knowledge mining, including clinical, medical and healthcare informatics, prescription drug abuse epidemiology and drug pharmacology, has become increasingly significant in recent years. Social media offers opportunities for people to share opinions and experiences freely in online communities, which may contribute information beyond the knowledge of domain professionals. This paper describes the development of a novel Semantic Web platform called PREDOSE (PREscription Drug abuse Online Surveillance and Epidemiology), which is designed to facilitate the epidemiologic study of prescription (and related) drug abuse practices using social media. PREDOSE uses web forum posts and domain knowledge, modeled in a manually created Drug Abuse Ontology (DAO) (pronounced dow), to facilitate the extraction of semantic information from User Generated Content (UGC). A combination of lexical, pattern-based and semantics-based techniques is used together with the domain knowledge to extract fine-grained semantic information from UGC. In a previous study, PREDOSE was used to obtain the datasets from which new knowledge in drug abuse research was derived. Here, we report on various platform enhancements, including an updated DAO, new components for relationship and triple extraction, and tools for content analysis, trend detection and emerging patterns exploration, which enhance the capabilities of the PREDOSE platform. Given these enhancements, PREDOSE is now more equipped to impact drug abuse research by alleviating traditional labor-intensive content analysis tasks. Using custom web crawlers that scrape UGC from publicly available web forums, PREDOSE first automates the collection of web-based social media content for subsequent semantic annotation. The annotation scheme is modeled in the DAO, and includes domain specific knowledge such as prescription (and related) drugs, methods of preparation, side effects, routes of administration, etc. The DAO is also used to help recognize three types of data, namely: 1) entities, 2) relationships and 3) triples. PREDOSE then uses a combination of lexical and semantic-based techniques to extract entities and relationships from the scraped content, and a top-down approach for triple extraction that uses patterns expressed in the DAO. In addition, PREDOSE uses publicly available lexicons to identify initial sentiment expressions in text, and then a probabilistic optimization algorithm (from related research) to extract the final sentiment expressions. Together, these techniques enable the capture of fine-grained semantic information from UGC, and querying, search, trend analysis and overall content analysis of social media related to prescription drug abuse. Moreover, extracted data are also made available to domain experts for the creation of training and test sets for use in evaluation and refinements in information extraction techniques. A recent evaluation of the information extraction techniques applied in the PREDOSE platform indicates 85% precision and 72% recall in entity identification, on a manually created gold standard dataset. In another study, PREDOSE achieved 36% precision in relationship identification and 33% precision in triple extraction, through manual evaluation by domain experts. Given the complexity of the relationship and triple extraction tasks and the abstruse nature of social media texts, we interpret these as favorable initial results. Extracted semantic information is currently in use in an online discovery support system, by prescription drug abuse researchers at the Center for Interventions, Treatment and Addictions Research (CITAR) at Wright State University. A comprehensive platform for entity, relationship, triple and sentiment extraction from such abstruse texts has never been developed for drug abuse research. PREDOSE has already demonstrated the importance of mining social media by providing data from which new findings in drug abuse research were uncovered. Given the recent platform enhancements, including the refined DAO, components for relationship and triple extraction, and tools for content, trend and emerging pattern analysis, it is expected that PREDOSE will play a significant role in advancing drug abuse epidemiology in future.
DOI: 10.15288/jsad.2008.69.703
发表时间: 2008-09-01
影响因子: 3.4
作者:
Boyer, Edward W.;Wines, James D., Jr.
通讯作者: Wines, James D., Jr.
DOI: 10.1080/10550490701525368
发表时间: 2007-01-01
影响因子: 3.7
作者:
Boyer, Edward W.;Babu, Kavita M.;Compton, Wilson
通讯作者: Compton, Wilson
DOI: 10.1016/0002-9343(90)90269-j
发表时间: 1990-06-20
影响因子: 5.9
作者:
ERICSSON, CD;JOHNSON, PC
通讯作者: JOHNSON, PC
DOI: 10.1145/1519103.1519105
发表时间: 2008-12-01
期刊: SIGMOD RECORD
影响因子: 1.1
作者:
Krishnamurthy, Rajasekar;Li, Yunyao;Zhu, Huaiyu
通讯作者: Zhu, Huaiyu
DOI: 10.1016/j.drugalcdep.2009.11.010
发表时间: 2010-04-01
影响因子: 4.2
作者:
Lange, James E.;Daniel, Jason;Clapp, John D.
通讯作者: Clapp, John D.