TeamX: A Sentiment Analyzer with Enhanced Lexicon Mapping and Weighting Scheme for Unbalanced Data

TeamX: A Sentiment Analyzer with Enhanced Lexicon Mapping and Weighting Scheme for Unbalanced Data
复制标题

DOI:
10.3115/v1/s14-2111
复制
发表时间:
2014-08
期刊:
--
影响因子:
--
通讯作者:
Yasuhide Miura;Shigeyuki Sakaki;K. Hattori;Tomoko Ohkuma
Yasuhide Miura;Shigeyuki Sakaki;K. Hattori;Tomoko Ohkuma
中科院分区:
其他
文献类型:
--
作者:
Yasuhide Miura;Shigeyuki Sakaki;K. Hattori;Tomoko Ohkuma

文献摘要

被引文献

相似文献

本文描述了TeamX在SemEval-2014任务9子任务B中使用的系统。该系统是一个基于监督文本分类方法的情感分析器,设计了以下两个概念。首先,由于词汇特征在SemEval-2013任务2中被证明是有效的,因此引入了各种词汇和预处理器来增强词汇信息。其次,由于已知推文上的情感分布是不平衡的,因此引入加权方案来偏置机器学习器的输出。对于测试运行,该系统针对Twitter文本进行了调整,并成功地在Twitter数据上获得了高分结果,Twitter 2014上的平均F1为70.96,Twitter 2014上的平均F1为56.50。
This paper describes the system that has been used by TeamX in SemEval-2014 Task 9 Subtask B. The system is a sentiment analyzer based on a supervised text categorization approach designed with following two concepts. Firstly, since lexicon features were shown to be effective in SemEval-2013 Task 2, various lexicons and pre-processors for them are introduced to enhance lexical information. Secondly, since a distribution of sentiment on tweets is known to be unbalanced, an weighting scheme is introduced to bias an output of a machine learner. For the test run, the system was tuned towards Twitter texts and successfully achieved high scoring results on Twitter data, average F1 70.96 on Twitter2014 and average F1 56.50 on Twitter2014Sarcasm.