课题基金 / 基金详情

Generating Descriptive Sentence Labels for Multinomial Sentiment-bearing Topics (GenSent)

Generating Descriptive Sentence Labels for Multinomial Sentiment-bearing Topics (GenSent)
为多项情感主题生成描述性句子标签 (GenSent)
批准号:
EP/P005810/1
负责人:
Chenghua Lin
金额:
$12.84万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
情感主题模型是一套算法,其目的是挖掘并从文本中揭示丰富的意见结构。情感主题模型的实用性源于这样一个事实,即推断出的隐藏的带有情感的主题,以单词的多项分布表示,类似于集合的意见信息,可以用作探索和理解大型非结构化文本档案中的意见的镜头。然而,将情感主题模型应用于探索性目的的一个主要挑战是解释发现的情感主题的含义,到目前为止,这完全依赖于人工解释。此外,目前的情绪承载模型无法促进准确的意见和情绪理解。例如,通过检查带有情绪的话题“亚马逊订单退货船收到退款损坏失望政策不满意”,可以解释为这个话题捕捉了与“不满意的网上购物体验”相关的意见。但是我们无法深入了解这个意见,即不满意的情绪是仅仅针对订购的产品,还是也与亚马逊的政策有关。情感主题的自动解释和标签解决方案是最及时的,因为:(i)当应用情感主题模型进行数据探索时,用户被迫手动解释推断的情感主题,这在分析高动态或大规模数据时是缓慢和不切实际的;(ii)促进准确意见理解的自动化工具对于许多实际应用(例如网络安全和商业智能)至关重要,因为它允许人们从大量文本数据中获取知识并制定决策,将数据转化为可操作的知识。该项目旨在通过开发一种新的句子标签自动生成框架来推动情感主题建模的前沿,该框架可以准确地描述多项情感主题的观点,并且在清晰度,简洁性和信息丰富性方面最适合人类。主要的挑战将是如何准确地解释蕴含在情感话题中的观点,以及如何生成简洁的句子标签,尽可能地传达情感话题的本质。这既雄心勃勃又冒险,因为:(i)已经证明,仅就主题信息自动标记标准主题是一项具有挑战性的任务(现有证据似乎支持这一点)。标记带有情感的主题涉及从情感和主题维度以及它们之间的依赖关系中捕获和解释语义,从而为标记任务增加了额外的复杂性维度;(ii)句子标签生成的两个要求,即最大的意见覆盖率和高度的简洁性,自然是相互冲突的。如何优化这两个正交目标之间的权衡,以生成最合适的句子标签是一个重要的科学问题。
英文摘要
Sentiment-topic models are a suite of algorithms whose aim is mine and uncover rich opinion structures from text. The utility of sentiment topic models stems from the fact that the inferred hidden sentiment-bearing topics, represented as a multinomial distribution over words, resemble the opinion information of a collection, which can be used as a lens for exploring and understanding opinions from large archives of unstructured text. However, a major challenge in applying sentiment-topic models for exploratory purposes is to interpret the meaning of the discovered sentiment-bearing topics, which, so far, relied entirely on manual interpretation. In addition, current sentiment-bearing models are not able to facilitate accurate opinion and sentiment understanding. For example, by examining the sentiment-bearing topic "amazon order return ship receive refund damaged disappointed policy unhappy", one can interpret that this topic captures opinions relating to "unsatisfactory online shopping experience". But it is impossible to gain deep insight of the opinion, i.e., whether the sentiment unhappy is only targeted to the product being ordered, or it is also related to Amazon's policy.A solution to automatic interpretation and labelling of sentiment-bearing topics is most timely because: (i) when applying sentiment-topic models for data exploration, users are forced to interpret the inferred sentiment-bearing topics manually, which is slow and impractical when analysing highly dynamic or large scale data; and (ii) automated tools facilitating accurate opinion understanding is crucial for many practical applications (e.g. cybersecurity and business intelligence), as it allows one to derive knowledge from large amounts of text data and to formulate decisions, converting data into actionable knowledge.The project aims to push the frontier of sentiment-topic modelling through the development of a novel framework for automated generation of sentence labels that can accurately describe the opinions of multinomial sentiment-bearing topics and are optimally suitable for humans in terms of clarity, brevity and information-richness. The main challenges will be the accurate interpretation of opinions encoded in sentiment-bearing topics and the generation of concise sentence labels which convey the essences of sentiment-bearing topics as much as possible. This is both ambitious and adventurous because: (i) it has already been demonstrated to be a challenging task to automatically labelling standard topics concerning topical information alone (as existing evidence seems to support). Labelling sentiment-bearing topics involves capturing and interpreting semantics from both sentiment and topic dimensions and the dependencies between them, thus adding an additional dimension of complexity for the labelling task; (ii) the two requirements for sentence label generation, i.e., maximal opinion coverage and high conciseness, naturally conflict with each other. How to optimise the trade-off between these two orthogonal objectives for generating a most suitable sentence label is an important scientific question.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/s11704-018-7073-5
发表时间: 2019-08
期刊: Frontiers of Computer Science
影响因子: 4.2
作者: [Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi]
通讯作者: Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi
DOI: --
发表时间: 2021-08
期刊: ArXiv
影响因子: --
作者: [M. Barawi;Chenghua Lin;Advaith Siddharthan;Yinbing Liu]
通讯作者: M. Barawi;Chenghua Lin;Advaith Siddharthan;Yinbing Liu
DOI: --
发表时间: 2017-11
期刊:
影响因子: --
作者: [N. Yusof;Chenghua Lin;Frank Guerin]
通讯作者: N. Yusof;Chenghua Lin;Frank Guerin
DOI: 10.18653/v1/p18-1113
发表时间: 2018-07
期刊:
影响因子: --
作者: [Rui Mao;Chenghua Lin;Frank Guerin]
通讯作者: Rui Mao;Chenghua Lin;Frank Guerin
共 8 条
    海外基金