FineSum: Target-Oriented, Fine-Grained Opinion Summarization

FineSum: Target-Oriented, Fine-Grained Opinion Summarization
复制标题

DOI:
10.1145/3539597.3570397
复制
发表时间:
2023-02
期刊:
Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining
影响因子:
--
通讯作者:
Suyu Ge;Jiaxin Huang;Yu Meng;Jiawei Han
Suyu Ge;Jiaxin Huang;Yu Meng;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Suyu Ge;Jiaxin Huang;Yu Meng;Jiawei Han

文献摘要

相似文献

面向目标的意见摘要是通过从多个相关文档中提取用户意见来描述目标。而不是简单地挖掘关于目标的意见评级(例如,餐馆)或多个方面(例如,食物,服务),则期望更深入,以挖掘关于细粒度子方面的意见(例如,鱼)。然而,在这种细粒度的尺度上获得高质量的注释是昂贵的。这导致我们提出了一个新的框架,FineSum,它在三个方面推进了意见分析的前沿:(1)最小监督,其中没有提供文档摘要对,只有方面名称和一些方面/情感关键字可用;(2)细粒度的意见分析,其中情感分析深入到每个一般方面中的特定主题或特征;(3)基于短语的文摘,以短语为基本单位进行文摘,并将语义连贯的短语聚集起来,提高文摘的一致性和全面性。给定一个没有注释的大型语料库,FineSum首先自动识别意见短语的潜在跨度,并使用方面和情感分类器进一步减少识别结果中的噪声。然后在每个方面和情感下构建多个细粒度的意见聚类。每个聚类表达对某些子方面的统一意见(例如,“食物”方面的“鱼”)或特征(例如,“墨西哥”在“食物”方面)。为了实现这一点,我们训练了一个球形单词嵌入空间来显式地表示不同的方面和情感。然后,我们从嵌入到上下文短语分类器中提取知识,并使用上下文感知短语嵌入进行聚类。自动评估的基准和定量的人类评估验证了我们的方法的有效性。
Target-oriented opinion summarization is to profile a target by extracting user opinions from multiple related documents. Instead of simply mining opinion ratings on a target (e.g., a restaurant) or on multiple aspects (e.g., food, service) of a target, it is desirable to go deeper, to mine opinion on fine-grained sub-aspects (e.g., fish). However, it is expensive to obtain high-quality annotations at such fine-grained scale. This leads to our proposal of a new framework, FineSum, which advances the frontier of opinion analysis in three aspects: (1) minimal supervision, where no document-summary pairs are provided, only aspect names and a few aspect/sentiment keywords are available; (2) fine-grained opinion analysis, where sentiment analysis drills down to a specific subject or characteristic within each general aspect; and (3) phrase-based summarization, where short phrases are taken as basic units for summarization, and semantically coherent phrases are gathered to improve the consistency and comprehensiveness of summary. Given a large corpus with no annotation, FineSum first automatically identifies potential spans of opinion phrases, and further reduces the noise in identification results using aspect and sentiment classifiers. It then constructs multiple fine-grained opinion clusters under each aspect and sentiment. Each cluster expresses uniform opinions towards certain sub-aspects (e.g., "fish" in "food" aspect) or characteristics (e.g., "Mexican" in "food" aspect). To accomplish this, we train a spherical word embedding space to explicitly represent different aspects and sentiments. We then distill the knowledge from embedding to a contextualized phrase classifier, and perform clustering using the contextualized opinion-aware phrase embedding. Both automatic evaluations on the benchmark and quantitative human evaluation validate the effectiveness of our approach.