课题基金 / 基金详情

CISE Research Resources: Discourse Penn Treebank and Multimodal FORM: Development of Two Richly Annotated Corpora

CISE Research Resources: Discourse Penn Treebank and Multimodal FORM: Development of Two Richly Annotated Corpora
CISE 研究资源:Discourse Penn Treebank 和 Multimodal FORM:两个注释丰富的语料库的开发
批准号:
0224417
负责人:
Aravind Joshi
金额:
$99.78万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2002
资助国家:
美国
项目状态:
已结题
起止时间:
2002-10-15 至 2006-09-30

项目摘要

项目成果

Aravind Joshi的其他基金

相似基金

相关文献

中文摘要
翻译
EIA-0224417 Aravind K. JoshiMark Liberman宾夕法尼亚大学CISE RR:Discourse Penn Trebank and Multimodal FORM:Development of Two Richly Annotated CorporaThis project,provides critical resources for research discourse modeling and conversational interaction,aims to develop new technologies and systems for information retrieval and human computer interaction. 围绕着标注语料库的建设,我们将建立两个大规模的语料库,一个是语篇领域的语料库,一个是对话领域的语料库。话语宾州树库(DPTB)和2. MultiFORM:用肢体动作、语音和语调扩充FORM语料库。前者开发了一个大规模的、注释可靠的语料库,它将编码与话语联系语相关的连贯关系,包括它们的论元结构和照应联系,从而揭示一个清晰定义的话语结构层次,并支持提取与话语联系语相关的一系列推理。 该注释将位于Penn Treebank(PTB)注释以及PTB的谓词-论元注释(称为命题库或Prop库)之上。 后者涉及一个手势注释的视频语料库,FORM被设计为可扩展的,以最终代表整个多模式的会话交互体验。 这种多模态形式,多形式,将通过添加身体运动,语音和句法结构,和语调。 大规模注释语料库在语音和自然语言研究中发挥了关键作用,它使统计知识(来自语料库)与语言知识(如注释中所表示的)大规模整合,从而导致科学和技术进步。 代表性的例子构成强大的分析和自动提取的关系和共指及其应用信息提取,问答,摘要和机器翻译。 PTB是十年前开发的一种资源,它代表了这种影响全球自然语言处理的资源的一个例子。 PTB在句子层面上处理语料库,提供一个新的大规模、可靠的话语和对话结构注释语料库。 虽然话语结构和对话结构的研究之间存在着知识和实践的联系,但研究这些领域的资源的初始要求在概念上重叠的同时也存在分歧。 在语篇方面,我们需要语料库来处理新闻文章等写作文本中的各种结构。 对话方需要关注人与人之间的互动,以及即兴而不是预先编写的材料。
英文摘要
EIA-0224417Aravind K. JoshiMark LibermanUniversity of PennsylvaniaCISE RR: Discourse Penn Trebank and Multimodal FORM: Development of Two Richly Annotated CorporaThis project, providing critical resources for research discourse modeling and conversational interaction, aims at developing new technologies and systems for information retrieval and human computer interaction. Centering on the construction of annotated corpora, two large-scale resources, one in the discourse domain and one in the dialog domain will be built:1. Discourse Penn Treebank (DPTB) and2. MultiFORM: Augmenting the FORM corpus with body movements, speech, and intonation.The former project develops a large scale and reliably annotated corpus that will encode coherence relations associated with discourse connectives, including their argument structure and anaphoric links, thus exposing a clearly defined level of discourse structure and supporting the extraction of a range of inferences associated with discourse connectives. This annotation will be "on top of" the Penn Treebank (PTB) annotations as well as the predicate-argument annotations of PTB (called the Proposition Bank or Prop Bank). The latter involves a corpus of gesture-annotated videos, FORM that was designed to be extensible in order to eventually represent the entire multimodal experience of conversational interaction. This multimodal FORM , MultiFORM, will be created by adding body movement, speech and syntactic structure, and intonation. Large-scale annotated corpora have played a critical role in speech and natural language research by enabling large-scale integration of statistical knowledge (derived from the corpora) with linguistic knowledge (as represented in annotations) leading to scientific and technological advances. Representative examples constitute robust parsing and automatic extraction of relations and coreferences and their applications to information extraction, question answering, summarization, and machine translation. PTB, a resource developed a decade ago, represents an example of such a resource that impacts natural language processing worldwide. PTB deals with corpora at the sentence level warranting a new large scale and reliable discourse and dialog structure annotated corpora. Although intellectual and practical connections exist between studies of the structures of discourse and dialog, the initial requirements for resources to study these areas diverge while overlapping in conception. On the discourse side, we need for corpora that deals with the kinds of structures found in composed text such as journalistic articles. The dialog side needs to focus on interactions among people and on extemporized rather than pre-composed material.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI: ADDO-EN: Significant Enhancement of the Exisitng Penn Discourse Treebank
  • 批准号:
    1059353
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2011
  • 负责人:
    Aravind Joshi
  • 依托单位:
RI: Exploiting and Exploring Discourse Connectivity: Deriving New Technology and Knowledge from the Penn Discourse Treebank
  • 批准号:
    0705671
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $95.5万
  • 财政年份:
    2007
  • 负责人:
    Aravind Joshi
  • 依托单位:
Metagrammatical Knowledge for Grammars and Corpora
  • 批准号:
    0414409
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2004
  • 负责人:
    Aravind Joshi
  • 依托单位:
ITR: Mining the Bibliome -- Information Extraction from the Biomedical Literature
  • 批准号:
    0205448
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $349.98万
  • 财政年份:
    2002
  • 负责人:
    Aravind Joshi
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)