The evolution of evaluation: Lessons from the Message Understanding Conferences

The evolution of evaluation: Lessons from the Message Understanding Conferences
复制标题

DOI:
10.1006/csla.1998.0102
复制
发表时间:
1998-10-01
影响因子:
4.3
通讯作者:
Hirschman, L
Hirschman, L
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hirschman, L

文献摘要

被引文献

相似文献

消息理解会议(MUC)代表了评估语言理解技术的最早和最长的努力之一。本文回顾了MUC的历史,以及它们向使用通用训练和盲测集、自动评分、任务分解为模块化构建块以及跨语言和应用程序可移植性工具的演变。既然评估已经成为开发人员工具箱中被接受的一部分,那么理解评估方法和研究状态之间的相互作用就很重要了。MUC成功地产生了关于文本处理问题的兴奋,并吸引了有才华的研究人员到该地区。它还将信息提取问题分解为一系列更简单的问题,从而使研究人员能够展示成功的系统并衍生出商业产品。然而,准确信息提取的最终目标一直难以实现,系统的构建速度越来越快,成本越来越低,评估变得越来越困难,但信息提取的总体准确性只得到了适度的提高。多国部队的经验与其他评价的经验形成对照。例如,航空旅行信息系统(ATIS)中的口头评估显示,随着时间的推移,错误率显着改善,但这些评估仅限于单个领域,并且即使实时交互系统可用,指标也没有解决交互问题。在相关评估的背景下,纵观MUC的历史,我们可以得出重要的经验教训,即评估需要随着它评估的技术而发展,平衡成本与收益,并权衡多个利益持有者-开发人员,资助者和用户的不同需求,以提供连续性,同时也为研究社区提供下一组挑战。(C)北京:科学出版社.
The Message Understanding Conferences (MUCs) represent one of the earliest and longest running efforts to evaluate language understanding technology. This article reviews the history of the MUCs and their evolution towards the use of common training and blind test sets, automated scoring, task decomposition into modular building blocks and tools for portability across languages and applications. Now that evaluation has become an accepted part of the developer's toolkit, it is important to understand the interplay between evaluation methods and the state of research. MUC was successful in generating excitement about text processing problems and in attracting talented researchers to the area. It also provided a functional decomposition of the information extraction problem into a series of simpler problems, thus allowing researchers to demonstrate successful systems and to spin off commercial products. However, the ultimate goal of accurate information extraction has been elusive, systems have become faster and cheaper to build, the evaluations have become harder, but overall accuracy in information extraction has improved only modestly. The MUC experience contrasts with experiences in other evaluations. For example, the spoken evaluation in the Air Travel Information System (ATIS) has shown dramatic improvement in error rate over time, but those evaluations were limited to a single domain and the metrics did not address interaction, even though real-time interactive systems were available. Looking across the history of MUC in the context of related evaluations, we can draw important lessons about the need for evaluation to evolve with the technology it evaluates, to balance costs against benefits and to weigh the divergent needs of the multiple stake-holders-developers, funders and users-in order to provide continuity while also providing the next set of challenges to the research community. (C) 1998 Academic Press.