Automated Text Markup for Information Retrieval from an Electronic Textbook of Infectious Disease

Automated Text Markup for Information Retrieval from an Electronic Textbook of Infectious Disease
复制标题

用于从传染病电子教科书中检索信息的自动文本标记

DOI:
--
复制
发表时间:
1998
期刊:
American Medical Informatics Association Annual Symposium
影响因子:
--
通讯作者:
Lawrence M. Fagan
Lawrence M. Fagan
中科院分区:
--
文献类型:
--
作者:
D. Berrios;A. Kehler;David K. Kim;V. Yu;Lawrence M. Fagan

文献摘要

被引文献

相似文献

摘要 临床执业医生的信息需求经常需要课本或期刊搜索。以电子形式提供这些来源提高了这些搜索的速度,但精确度(即相关文件占检索文件总数的比例)仍然很低。通过将搜索词转换为规范概念来改进传统的关键字搜索,并不能大幅提高搜索精度。 Kim等人。设计并建立了一个从即将出版的传染病电子教科书中进行计算机信息检索的原型系统(MYCIN II)。该系统需要专家以复杂文本标记的形式进行手动索引。然而,这一标记过程非常耗时(为218章中的每一章生成、审查和转录索引大约需要3个人小时)。 我们设计并实现了一个半自动标记过程的系统。该系统名为《文件半自动索引信息提取》(ISAID),它使用查询模型和现有的信息提取工具,为任何用户,包括原始材料的作者,提供支持,以便快速、准确地标记三级信息源。
Abstract The information needs of practicing clinicians frequently require textbook or journal searches. Making these sources available in electronic form improves the speed of these searches, but precision (i.e., the fraction of relevant to total documents retrieved) remains low. Improving the traditional keyword search by transforming search terms into canonical concepts does not improve search precision greatly. Kim et al. have designed and built a prototype system (MYCIN II) for computer-based information retrieval from a forthcoming electronic textbook of infectious disease. The system requires manual indexing by experts in the form of complex text markup. However, this mark-up process is time consuming (about 3 person-hours to generate, review, and transcribe the index for each of 218 chapters). We have designed and implemented a system to semiautomate the markup process. The system, information extraction for semiautomated indexing of documents (ISAID), uses query models and existing information-extraction tools to provide support for any user, including the author of the source material, to mark up tertiary information sources quickly and accurately.