课题基金 / 基金详情

CAREER: Content and Cohesion Models, with Applications to Text Summarization and Natural Language Generation

CAREER: Content and Cohesion Models, with Applications to Text Summarization and Natural Language Generation
职业:内容和衔接模型,及其在文本摘要和自然语言生成中的应用
批准号:
0448168
负责人:
Regina Barzilay
金额:
$40.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-02-15 至 2012-01-31

项目摘要

项目成果

Regina Barzilay的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年中,概率方法已经成功地提供了自然语言文本的分析,反过来,又实现了广泛的有价值的实际应用,如机器翻译、问题回答和摘要。尽管取得了这样的成功,但是现有的方法有一个基本的限制:它们处理每个文档时很少或根本没有利用其全局结构的能力。这通常会导致手头任务的性能不够理想。这个项目的目标是为文本、内容和衔接这两个基本的、正交的维度开发概率模型。基于第一个维度(内容)的模型描述文本中出现的主题及其组织。第二个维度,衔接,关注的是在给定的文本中信息是如何实现的。这些模型的发展需要新的无监督技术,能够捕获复杂的文本属性,以及用于主题离散化和话语语法归纳的新算法。获得一个好的文本构成的计算度量将在人文科学和计算机科学的边缘开辟新的研究途径。概率文本模型将成为文本摘要和生成新方法的基础,这将使在线信息比目前的情况更容易获得。这将极大地影响人们体验多种形式的在线文本信息的方式,包括新闻报道、消费者健康信息和政府文件。学生将通过实践项目、拓展项目和本科和研究生阶段的课程参与到这项研究中来。
英文摘要
Within the last decade, probabilistic methods have delivered successful analyses of natural language texts that, in turn, have enabled a broad range of valuable and practical applications, such as machine translation, question answering, and summarization. Despite this success, existing methods suffer from a fundamental limitation: they process each document with little or no ability to take advantage of its global structure. All too often, this results in suboptimal performance for the task at hand.The goal of this project is to develop probabilistic models for two fundamental, orthogonal dimensions of text, content and cohesion. A model based on the first dimension, content, describes the topics present in a text and their organization. The second dimension, cohesion, is concerned with how information is realized in a given text. Development of these models requires new unsupervised techniques able to capture complex text properties and novel algorithms for topical discretization and discourse grammar induction.Gaining a computational measure of what constitutes a good text will open new research avenues on the edge of humanities and computer science. Probabilistic text models will form a basis for novel approaches to text summarization and generation that will make on-line information much more accessible than is currently the case. This will substantially affect the way people experience the many forms of textual on-line information, including news reports, consumer health information, and government documents. Students will become involved in this research through hands-on projects, outreach programs, and courses at both the undergraduate and graduate level.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: Reconstructing the Tower of Babel: Cross-lingual Language Learning
Student Research Workshop in Computational Linguistics, at the Association for Computational Linguistics (ACL) 2005 Conference; June 27, 2005; Ann Arbor, MI
Automatic Processing of Spoken and Written Lecture Material
海外基金