课题基金 / 基金详情

Generating Descriptive Sentence Labels for Multinomial Sentiment-bearing Topics (GenSent)

Generating Descriptive Sentence Labels for Multinomial Sentiment-bearing Topics (GenSent)
为多项情感主题生成描述性句子标签 (GenSent)
批准号:
EP/P005810/1
负责人:
Chenghua Lin
金额:
$12.84万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
情感主题模型是一套算法,其目的是挖掘并从文本中发现丰富的观点结构。情感主题模型的效用源于这样的事实,即推断出的隐藏的情感承载主题,表示为多项分布的话,类似于意见信息的集合,它可以被用作一个透镜,用于探索和理解意见,从大型档案的非结构化文本。然而,将情感主题模型应用于探索性目的的一个主要挑战是解释所发现的情感主题的含义,到目前为止,这完全依赖于人工解释。此外,目前的情感承载模型无法促进准确的意见和情感理解。例如,通过检查带有情感的主题“亚马逊订单退货收到退款损坏失望政策不满意”,可以解释该主题捕获与“不满意的在线购物体验”有关的意见。但是,要深入了解意见是不可能的,即,情绪不愉快是否仅针对所订购的产品,或者它也与亚马逊的政策有关。自动解释和标记带有情绪的主题的解决方案是最及时的,因为:(i)当应用情绪主题模型进行数据探索时,用户被迫手动解释推断的带有情绪的主题,这在分析高度动态或大规模数据时是缓慢且不切实际的;以及(ii)自动化工具促进准确的意见理解对于许多实际应用至关重要(例如网络安全和商业智能),因为它允许人们从大量文本数据中获取知识并制定决策,将数据转化为可操作的知识。该项目旨在推动情感的前沿-主题建模,通过开发用于自动生成句子标签的新框架,该框架可以准确地描述多项情感承载主题的意见,并且在清晰度方面最适合人类,简洁和信息丰富。主要的挑战将是准确地解释表达情感的主题中编码的意见,并生成尽可能传达情感主题本质的简洁句子标签。这既雄心勃勃又具有冒险精神,因为:(i)已经证明,自动标记仅涉及主题信息的标准主题是一项具有挑战性的任务(现有证据似乎支持这一点)。标记带有情感的主题涉及从情感和主题维度以及它们之间的依赖关系捕获和解释语义,从而为标记任务增加了额外的复杂性维度;(ii)句子标签生成的两个要求,即,最大的意见覆盖率和高度的简洁性,自然是相互冲突的。如何优化这两个正交目标之间的权衡,以生成最合适的句子标签是一个重要的科学问题。
英文摘要
Sentiment-topic models are a suite of algorithms whose aim is mine and uncover rich opinion structures from text. The utility of sentiment topic models stems from the fact that the inferred hidden sentiment-bearing topics, represented as a multinomial distribution over words, resemble the opinion information of a collection, which can be used as a lens for exploring and understanding opinions from large archives of unstructured text. However, a major challenge in applying sentiment-topic models for exploratory purposes is to interpret the meaning of the discovered sentiment-bearing topics, which, so far, relied entirely on manual interpretation. In addition, current sentiment-bearing models are not able to facilitate accurate opinion and sentiment understanding. For example, by examining the sentiment-bearing topic "amazon order return ship receive refund damaged disappointed policy unhappy", one can interpret that this topic captures opinions relating to "unsatisfactory online shopping experience". But it is impossible to gain deep insight of the opinion, i.e., whether the sentiment unhappy is only targeted to the product being ordered, or it is also related to Amazon's policy.A solution to automatic interpretation and labelling of sentiment-bearing topics is most timely because: (i) when applying sentiment-topic models for data exploration, users are forced to interpret the inferred sentiment-bearing topics manually, which is slow and impractical when analysing highly dynamic or large scale data; and (ii) automated tools facilitating accurate opinion understanding is crucial for many practical applications (e.g. cybersecurity and business intelligence), as it allows one to derive knowledge from large amounts of text data and to formulate decisions, converting data into actionable knowledge.The project aims to push the frontier of sentiment-topic modelling through the development of a novel framework for automated generation of sentence labels that can accurately describe the opinions of multinomial sentiment-bearing topics and are optimally suitable for humans in terms of clarity, brevity and information-richness. The main challenges will be the accurate interpretation of opinions encoded in sentiment-bearing topics and the generation of concise sentence labels which convey the essences of sentiment-bearing topics as much as possible. This is both ambitious and adventurous because: (i) it has already been demonstrated to be a challenging task to automatically labelling standard topics concerning topical information alone (as existing evidence seems to support). Labelling sentiment-bearing topics involves capturing and interpreting semantics from both sentiment and topic dimensions and the dependencies between them, thus adding an additional dimension of complexity for the labelling task; (ii) the two requirements for sentence label generation, i.e., maximal opinion coverage and high conciseness, naturally conflict with each other. How to optimise the trade-off between these two orthogonal objectives for generating a most suitable sentence label is an important scientific question.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/s11704-018-7073-5
发表时间: 2019-08
期刊: Frontiers of Computer Science
影响因子: 4.2
作者: [Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi]
通讯作者: Ebuka Ibeke;Chenghua Lin;A. Wyner;M. Barawi
DOI: --
发表时间: 2021-08
期刊: ArXiv
影响因子: --
作者: [M. Barawi;Chenghua Lin;Advaith Siddharthan;Yinbing Liu]
通讯作者: M. Barawi;Chenghua Lin;Advaith Siddharthan;Yinbing Liu
DOI: --
发表时间: 2017-11
期刊:
影响因子: --
作者: [N. Yusof;Chenghua Lin;Frank Guerin]
通讯作者: N. Yusof;Chenghua Lin;Frank Guerin
DOI: 10.18653/v1/p18-1113
发表时间: 2018-07
期刊:
影响因子: --
作者: [Rui Mao;Chenghua Lin;Frank Guerin]
通讯作者: Rui Mao;Chenghua Lin;Frank Guerin
共 8 条
    海外基金