课题基金 / 基金详情

Machine Learning Methods for Personalised, Abstractive Summarisation of Consumer-Generated Media

Machine Learning Methods for Personalised, Abstractive Summarisation of Consumer-Generated Media
用于消费者生成媒体的个性化、抽象总结的机器学习方法
批准号:
EP/I004327/1
负责人:
Kalina Bontcheva
金额:
$75.4万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --

项目摘要

项目成果

Kalina Bontcheva的其他基金

相似基金

相关文献

中文摘要
翻译
Web 2.0和CGM的成功是基于利用人类互动的社会性质,使人们能够表达自己的意见,成为虚拟社区的一部分,并远程协作。如果我们以微博为例,2008年至2009年Twitter访问量的增长超过1,000%,预计到2010年,所有互联网用户中约有10%将使用Twitter。在线内容的数量和重要性前所未有地增加,导致公司和个人花费越来越多的时间试图跟上相关的CGM。据估计,每年700人小时是公司和公共服务需要花费在CGM监测、在线用户参与和发现新信息上的绝对最低限度。这个奖学金是关于帮助人们科普由此产生的信息过载,通过自动方法,能够适应个人的信息寻求目标,并简要总结相关媒体,从而支持信息解释和决策。自动文本摘要是我们目标的关键,包括压缩文本文档的含义,同时保留其中包含的相关信息。虽然已经有很多关于新闻等精心撰写的文本的研究,但社交媒体的摘要仍处于起步阶段,研究重点是产品评论。一个关键的实验发现是,由于社交媒体的特点,(特别是产品评论)最好首先从不同的文档和网站中提取相关信息,然后使用自然语言生成基于这些信息创建流畅的文本。在这个奖学金中,我将研究和评估新的机器学习方法,用于个性化,跨不同社交媒体的抽象多文档摘要。例如,将关于给定主题的联合收割机Twitter帖子、博客文章和Facebook留言墙消息相结合的历时摘要。与以前的工作相比,我们将采取跨学科的方法,这将有助于我们研究CGM摘要的社会层面,并建立实际的用户需求。第二个研究挑战是,算法需要在面对这种嘈杂的,充满行话和动态的内容,以及需要能够代表CGM的矛盾和强烈的时间性质的模型是强大的。我们工作的一个关键的新贡献是基于用户兴趣,目标和社会背景的模型来个性化摘要。可信度、隐私和在线社区(及其中心和权威)等问题也将发挥重要作用。第四个研究挑战是生成个性化的抽象摘要,可以帮助用户理解和解释内容。我的研究的一个令人兴奋的元素将是在研究不同类型的摘要,是有用的各种真实的用户(公司,记者,和一般公众)通过多学科合作与新闻协会,英国电信,牛津互联网研究所,和谢菲尔德的新闻系。一个关键的项目交付将是一个公开的浏览器插件,提供方便地访问自动生成的摘要。这将使我能够与真实的用户一起大规模地评估项目结果。它还将为自然语言生成社区提供新的评估挑战,因为研究人员将能够将他们的摘要与我们的开源算法提供的摘要进行比较。最后但并非最不重要的是,该奖学金不仅涵盖基础多学科研究,而且还测试了涉及商业合作伙伴(新闻协会,英国电信,Fizzback)的几个数字经济试点实验的结果。
英文摘要
The success of Web 2.0 and CGM is based on tapping into the social nature of human interactions, by making it possible for people to voice their opinion, become part of a virtual community and collaborate remotely. If we take micro-blogging as an example, the growth in Twitter visits between 2008 and 2009 was over 1,000% and it is projected that by 2010 around 10% of all internet users will be on Twitter. This unprecedented rise in the volume and importance of online content has resulted in companies and individuals spending ever increasing amounts of time trying to keep up with relevant CGM. It is estimated that 700 person hours per year is the absolute minimum that companies and public services need to spend on CGM monitoring, online user engagement, and discovery of new information. This fellowship is about helping people to cope with the resulting information overload, through automatic methods that are capable of adapting to individual's information seeking goals and summarising briefly the relevant media and thus supporting information interpretation and decision making. Automatic text summarisation is key to our goal and consists of compressing the meaning of text documents while preserving the relevant information contained within them. While there has been a lot of research on well-authored texts such as news, summarisation of social media is still in its infancy, with research focused on product reviews. A key experimental finding has been that due to the characteristics of social media (product reviews in particular) it is better first to abstract the relevant information from the different documents and sites and then to use natural language generation to create a fluent text based on this information.In this fellowship I will investigate and evaluate new machine learning methods for personalised, abstractive multi-document summarisation across different social media. For example, diachronic summaries that combine Twitter posts, blog articles, and Facebook wall messages on a given topic. In contrast to previous work, we will pursue an inter-disciplinary approach, which will help us study the social dimension of CGM summarisation and establish actual user needs. The second research challenge is that the algorithms need to be robust in the face of this noisy, jargon-full and dynamic content, as well as needing models capable of representing the contradictory and strongly temporal nature of CGM. A key novel contribution of our work is personalising the summaries, based on a model of user interests, goals, and social context. Issues such as trustworthiness, privacy, and online communities (with their hubs and authorities) will also play an important role. The fourth research challenge is to generate personalised abstractive summaries that can help users with sensemaking and content interpretation. An exciting element of my research will be in studying the different kinds of summaries that are useful for a variety of real users (companies, journalists, and the general public) through multi-disciplinary collaborations with the Press Association, British Telecom, the Oxford Internet Institute, and Sheffield's Department of Journalism. A key project deliverable will be a publicly available browser plugin that provides easy access to the automatically generated summaries. This will allow me to evaluate the project results with real users, on a large scale. It will also provide a new evaluation challenge for the Natural Language Generation community, as researchers will be able to compare their summarisers against those delivered by our open-source algorithms. Last but not least, the fellowship covers not only foundational multi-disciplinary research but it also tests the results in several Digital Economy pilot experiments involving commercial partners (The Press Association, British Telecom, Fizzback).
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.3233/sw-130110
发表时间: 2014
期刊: Semantic Web
影响因子: 3
作者: [Kalina Bontcheva;D. Rout]
通讯作者: Kalina Bontcheva;D. Rout
DOI: 10.1016/j.csl.2017.01.012
发表时间: 2017-07-01
期刊: COMPUTER SPEECH AND LANGUAGE
影响因子: 4.3
作者: [Augenstein, Isabelle, Derczynski, Leon, Bontcheva, Kalina]
通讯作者: Bontcheva, Kalina
Stance Detection with Bidirectional Conditional Encoding
使用双向条件编码进行姿态检测
DOI: 10.48550/arxiv.1606.05464
发表时间: 2016
期刊: arXiv e-prints
影响因子: --
作者: [Augenstein Isabelle]
通讯作者: Augenstein Isabelle
Working with Text: Tools, Techniques and Approaches for Text Mining
处理文本:文本挖掘的工具、技术和方法
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者: [Bontcheva K]
通讯作者: Bontcheva K
XAIvsDisinfo: eXplainable AI Methods for Categorisation and Analysis of COVID-19 Vaccine Disinformation and Online Debates
  • 批准号:
    EP/W011212/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $29.71万
  • 财政年份:
    2021
  • 负责人:
    Kalina Bontcheva
  • 依托单位:
Responsible AI for Inclusive, Democratic Societies: A cross-disciplinary approach to detecting and countering abusive language online
  • 批准号:
    ES/T012714/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $64.75万
  • 财政年份:
    2020
  • 负责人:
    Kalina Bontcheva
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: