Prodigy: Probabilistic Deep Generation
Prodigy: Probabilistic Deep Generation
批准号:
EP/E029116/1
负责人:
Anya Belz
金额:
$26.91万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2007
资助国家:
英国
项目状态:
已结题
起止时间:
2007 至 --
中文摘要
生成语言的计算方法在几个方面落后于分析语言的计算方法,最明显的是它们没有在商业上使用。主要原因是生成语言的系统需要花费大量的时间来构建,但是一旦构建就不能重用,并且往往严重缺乏语言变化,这很容易被认为是缺乏质量的。语言生成研究的现状让人想起20世纪80年代末的语言分析研究,当时符号方法和统计方法短暂地形成了完全独立的研究范式。语言分析很快转向范式合并,意识到符号方法缺乏概率方法所能提供的效率和鲁棒性,而概率方法反过来又受益于符号方法的准确性和微妙性。在机器翻译领域也出现了类似的发展,在纯统计方法主导该领域数年之后,研究人员现在开始将语言学知识重新引入该领域。这些研究领域的经验表明,当符号范式和统计范式相结合时,可以实现更高的质量。最近的研究表明,语言生成可能也是如此。Prodigy项目的目的是首次开发一种全面的、语言信息灵通的、概率的方法,用于生成语言,从而大大改善语言生成系统的开发时间、可重用性和语言变化,从而提高它们的商业可行性。以首席研究员之前在epsrc资助下对概率NLG的研究为起点,Prodigy项目将探索概率和语言学的结合是否能像对语言分析一样有益于语言生成领域。我们将特别关注两个方面:(i)开发可重用的数据表示和编码策略,以及(ii)开发用于指导语言生成过程的特定概率技术。我们将在五个不同的数据集上测试和评估我们的表示和技术,这些数据集是从现实世界的文本生产任务中收集的,包括天气预报、博物馆展品描述和护士报告。Prodigy项目将产生对工业、研究界和个人最终用户都有潜在好处的研究成果。研究将主要受益于我们对可重用语言生成技术的理解的进步,工业将主要受益于商业可行性的提高,技术本身可以通过加速文本生成来帮助个人用户,以及通过提供一种并不总是存在的模式(例如,使视障读者能够访问图形信息)。
英文摘要
Computational methods for generating language are lagging behind computational methods for analysing language in several ways, most obviously in that they are not used commercially. The main reasons are that systems for generating language take inordinate amounts of time to build, yet once built cannot be reused, and tend to be severely lacking in language variation, something that is easily perceived as a lack of quality. The current situation in language generation research is reminiscent of language analysis research in the late 1980s, when symbolic and statistical methods briefly formed entirely separate research paradigms. Language analysis soon moved towards a paradigm merger, realising that symbolic methods lacked the efficiency and robustness that probabilistic methods could provide, which in turn would benefit from the accuracy and subtlety of symbolic methods. A similar development is currently underway in the field of machine translation where - after several years of purely statistical methods dominating the field - researchers are now beginning to bring linguistic knowledge back in. The experience from these research fields suggests that higher quality can be achieved when the symbolic and statistical paradigms join forces. Recent research shows that this is likely to be true for language generation too. The purpose of the Prodigy project is to develop, for the first time, a comprehensive, linguistically informed, probabilistic methodology for generating language that substantially improves development time, reusability and language variation in language generation systems, and thereby enhances their commercial viability. Taking the principal investigator's previous EPSRC-funded research on probabilistic NLG as a starting point, the Prodigy project will explore whether the combination of the probabilistic and the linguistic can be as beneficial for the field of language generation as it has been for language analysis. We will focus on two aspect in particular: (i) developing reusable data representation and encoding strategies, and (ii) developing specific probabilistic techniques for guiding language generation processes.We will test and evaluate our representations and techniques on five different data sets which have been collected from real-world text production tasks and include weather forecasts, descriptions of museum exhibits, and nurses' reports.The Prodigy project will produce research outcomes that are of potential benefit to industry, the research community and individual end-users. Research will primarily benefit through advances in our understanding of reusable language generation technology, industry through improvements in commercial viability, and the technology itself can help individual users by speeding up text production, as well as by making available a modality that does not always exist (e.g. enabling visually impaired readers to access graphical information).
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.3115/1610195.1610198
发表时间:
2009-03
期刊:
影响因子:
--
作者:
[A. Belz;Eric Kow]
通讯作者:
A. Belz;Eric Kow
Extracting Parallel Fragments from Comparable Corpora for Data-to-text Generation
从可比较的语料库中提取并行片段以生成数据到文本
DOI:
--
发表时间:
2010
期刊:
影响因子:
--
作者:
[AS Belz]
通讯作者:
AS Belz
DOI:
--
发表时间:
2010
期刊:
影响因子:
--
作者:
[AS Belz]
通讯作者:
AS Belz
Assessing the Trade-Off between System Building Cost and Output Quality in Data-to-Text Generation
评估数据到文本生成中系统构建成本和输出质量之间的权衡
DOI:
--
发表时间:
2010
期刊:
影响因子:
--
作者:
[AS Belz]
通讯作者:
AS Belz
ReproHum: Investigating Reproducibility of Human Evaluations in Natural Language Processing
-
批准号:EP/V05645X/1
-
项目类别:Research Grant
-
资助金额:$28.95万
-
财政年份:2022
-
负责人:Anya Belz
-
依托单位:
Generation Challenges 2011: Towards a Surface Realisation Shared Task
-
批准号:EP/I032320/1
-
项目类别:Research Grant
-
资助金额:$8.65万
-
财政年份:2011
-
负责人:Anya Belz
-
依托单位:
EPSRC Network on Vision and Language (V&L Net)
-
批准号:EP/H018557/1
-
项目类别:Research Grant
-
资助金额:$13.28万
-
财政年份:2010
-
负责人:Anya Belz
-
依托单位:
Generation Challenges 2010
-
批准号:EP/H032886/1
-
项目类别:Research Grant
-
资助金额:$5.4万
-
财政年份:2010
-
负责人:Anya Belz
-
依托单位:
Generation Challenges 2009
-
批准号:EP/G03995X/1
-
项目类别:Research Grant
-
资助金额:$4.6万
-
财政年份:2009
-
负责人:Anya Belz
-
依托单位:
REG Challenge 2008: A Shared Task Evaluation Event for Referring Expression Generation
-
批准号:EP/F059760/1
-
项目类别:Research Grant
-
资助金额:$2.22万
-
财政年份:2008
-
负责人:Anya Belz
-
依托单位:
海外基金