Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus

Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus
复制标题

DOI:
10.1007/s10579-007-9040-x
复制
发表时间:
2007-05-01
影响因子:
2.7
通讯作者:
Carletta, Jean
Carletta, Jean
中科院分区:
计算机科学4区
文献类型:
--
作者:
Carletta, Jean

文献摘要

被引文献

相似文献

AMI会议语料库包含使用许多同步记录设备捕获的100小时会议,旨在支持语音和视频处理、语言工程、语料库语言学和组织心理学方面的工作。它是按正字法转录的,带有注释的子集,从命名实体、对话行为、摘要到简单的凝视和头部移动。在LREC会议主旨演讲的这个书面版本中,我描述了数据以及它是如何创建的。如果这是“杀手级”数据,则前提是它将“销售”一个平台;在本例中,它是Nite XML工具包,它允许一组分布式用户为相同的基本数据创建、存储、浏览和搜索批注,这些批注既与信号时间一致,又在结构上相互关联。
The AMI Meeting Corpus contains 100 h of meetings captured using many synchronized recording devices, and is designed to support work in speech and video processing, language engineering, corpus linguistics, and organizational psychology. It has been transcribed orthographically, with annotated subsets for everything from named entities, dialogue acts, and summaries to simple gaze and head movement. In this written version of an LREC conference keynote address, I describe the data and how it was created. If this is "killer'' data, that presupposes a platform that it will "sell''; in this case, that is the NITE XML Toolkit, which allows a distributed set of users to create, store, browse, and search annotations for the same base data that are both time-aligned against signal and related to each other structurally.