A method for determining the number of documents needed for a gold standard corpus

A method for determining the number of documents needed for a gold standard corpus
复制标题

DOI:
10.1016/j.jbi.2011.12.010
复制
发表时间:
2012-06-01
影响因子:
4.5
通讯作者:
Juckett, David
Juckett, David
中科院分区:
医学3区
文献类型:
--
作者:
Juckett, David

文献摘要

被引文献

相似文献

医学中的非结构化叙述越来越多地成为使用自然语言处理(NLP)技术进行内容提取的目标。在大多数情况下,这些工作是通过创建一个手动注释的一组包含地面真相的叙述来促进的;通常被称为黄金标准语料库。该语料库用于建模、微调和测试NLP软件,并为机器学习的训练提供基础。确定这个语料库的注释文档的数量(大小)是很重要的,但很少描述;相反,成本和时间的因素似乎占主导地位的语料库大小的决策。在这份报告中,概述了一种方法来确定黄金标准大小的基础上捕获概率的唯一的话在目标语料库。为了证明这种方法,从密歇根州疼痛顾问(MPC)疼痛管理诊所的口述信件语料库进行了描述和分析。首先构建了一个包含10,000个听写的格式良好的工作语料库,以提供总数的代表性子集,每个患者不超过一个听写字母。每一次听写都被分成单词,常见的单词被删除。泊松函数用于确定从工作语料库中获取的样本内的单词捕获概率,然后在单词长度上进行积分以给出作为样本大小的函数的单个捕获概率。对于这些MPC听写,预测500个文档的样本大小给出约0.95的捕获概率。继续样本选择的演示,选择了500个文档的临时金标准语料库,并检查其与MPC结构化编码的相似性和每个患者可用的人口统计数据。结果表明,一个有代表性的样本,合理的大小,可以选择用作金标准。(C)2012 Elsevier Inc. All rights reserved.
The unstructured narratives in medicine have been increasingly targeted for content extraction using the techniques of natural language processing (NLP). In most cases, these efforts are facilitated by creating a manually annotated set of narratives containing the ground truth; commonly referred to as a gold standard corpus. This corpus is used for modeling, fine-tuning, and testing NLP software as well as providing the basis for training in machine learning. Determining the number of annotated documents (size) for this corpus is important, but rarely described; rather, the factors of cost and time appear to dominate decision-making about corpus size. In this report, a method is outlined to determine gold standard size based on the capture probabilities for the unique words within a target corpus. To demonstrate this method, a corpus of dictation letters from the Michigan Pain Consultant (MPC) clinics for pain management are described and analyzed. A well-formed working corpus of 10,000 dictations was first constructed to provide a representative subset of the total, with no more than one dictation letter per patient. Each dictation was divided into words and common words were removed. The Poisson function was used to determine probabilities of word capture within samples taken from the working corpus, and then integrated over word length to give a single capture probability as a function of sample size. For these MPC dictations, a sample size of 500 documents is predicted to give a capture probability of approximately 0.95. Continuing the demonstration of sample selection, a provisional gold standard corpus of 500 documents was selected and examined for its similarity to the MPC structured coding and demographic data available for each patient. It is shown that a representative sample, of justifiable size, can be selected for use as a gold standard. (C) 2012 Elsevier Inc. All rights reserved.