Automated classification of author's sentiments in citation using machine learning techniques: A preliminary study

Automated classification of author's sentiments in citation using machine learning techniques: A preliminary study
复制标题

使用机器学习技术对引文中作者的情绪进行自动分类:初步研究

DOI:
10.1109/cibcb.2015.7300319
复制
发表时间:
2015
期刊:
2015 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB)
影响因子:
--
通讯作者:
G. Thoma
G. Thoma
中科院分区:
--
文献类型:
--
作者:
In;G. Thoma

文献摘要

被引文献

相似文献

科学论文通常包括引用外部来源,如期刊文章,书籍或网络链接,以引用与研究重要相关的作品。引用的原因出现在正文中引用标记周围的句子中,并表示引用与被引用作品之间的关系,如支持,对比,纠正等,这可能是研究人员为某一研究目的寻找相关先前工作或方法的重要线索。我们建议开发一种自动化的方法来识别引用作者的情感对引用句子中表达的引用外部来源,使用机器学习技术和语言线索。作为一个初步的研究,本文提出了一种基于支持向量机(SVM)的文本分类技术来分类作者的情绪,特别是对评论(CON)的文章。CON是一个MEDLINE引用字段,表示给定文章的作者评论的先前发表的文章,表达了可能的赞美或矛盾的观点。实现了具有径向基核函数(RBF)的SVM,并且基于表示CON句子中的词的分布的n-gram词统计来创建用于SVM的输入特征向量。对从414种不同的在线生物医学期刊标题中收集的一组CON句子进行的实验表明,具有RBF的SVM对于结合单元语法和双元语法单词统计的输入特征向量产生最佳结果。
Scientific papers generally include citations to external sources such as journal articles, books, or Web links to refer to works that are related in an important way to the research. The reason for the citation appears within the sentences surrounding the citation tag in the body text, and represents the relationship between the citation and cited works as supportive, contrastive, corrective, etc. This could be an important clue for researchers seeking relevant previous work or approaches for a certain research purpose. We propose to develop an automated method to identify the citing author's sentiments toward the cited external sources expressed in citation sentences using machine-learning techniques and linguistic cues. As a preliminary study, this paper presents a support vector machine (SVM)-based text categorization technique to classify the author's sentiments specifically toward Comment-on (CON) articles. CON, a MEDLINE citation field, indicates previously published articles commented on by authors of a given article expressing possibly complimentary or contradictory opinions. An SVM with a radial basis kernel function (RBF) is implemented, and Input feature vectors for the SVM are created based on n-grams word statistics representing the distribution of words in CON sentences. Experiments conducted on a set of CON sentences collected from 414 different online biomedical journal titles show that the SVM with a RBF yields the best result for an input feature vector combining uni-gram and bi-gram word statistics.