DISCIPLINING OF MEDICAL DATA

DISCIPLINING OF MEDICAL DATA
复制标题

DOI:
10.1093/oxfordjournals.bmb.a070637
复制
发表时间:
1968-01-01
影响因子:
6.7
通讯作者:
HEALY, MJR
HEALY, MJR
中科院分区:
医学2区
文献类型:
--
作者:
HEALY, MJR

文献摘要

被引文献

相似文献

计算机在医学上的许多应用,特别是在医学研究上的应用,可归为两个主要标题之一:汇集和获取大量数据,以及随后用统计和数学技术对这些数据进行处理。在整个过程中,有一个默认的假设,即数据的意思是他们所说的,可以这么说,例如,历史记录,病人的血清电解质水平有一定的值,或者记录的特定范围的症状,可以被信任,至少是接近事实的。更微妙的是,如果正在研究一种特定的情况,研究人员将需要确保记录保存系统检索到的病例在某种可定义的意义上是相同的,以便为概括和预测提供一个安全的基础。这些基本假设可能会在太多的情况下失败,甚至那些非常熟悉数据分析问题的人也并不总是意识到检查这些假设的必要性,或者认识到现有各种检查技术的力量。常规收集的数据的不可靠性得到了相当广泛的承认,但即使是在研究条件下仔细获得的材料,也比通常认识到的更容易受到测量、记录和转录的某个阶段发生的严重错误的影响(例如,参见Healy, 1952)。即使是年龄和性别等“明显”的项目也很少被错误记录,从而表明80岁的患者患有儿童疾病或男性患子宫癌。最后这些例子表明,在分析数据之前,需要对数据的可信度进行检查。坦率地说,这两个例子令人难以置信,但更常见的是可信度低的例子——这些案例与材料的主体相去甚远,需要进行特别调查。其中不可避免地涉及到个人判断的因素,并没有暗示孤立的项目仅仅是被“拒绝”:它们通常比更正统的案例包含更多的兴趣,但这使得它们的识别更重要,而不是不重要。关于异常值的识别有大量的统计文献,但其中大部分价值相当小。许多建议的技术容易给出愚蠢的答案,而且它们的理论背景往往是不充分的
Many applications of computers in medicine, particularly in medical research, come under one of two main headings: the bringing together and making accessible of bodies of data and the subsequent processing of these data by statistical and mathematical techniques. It is a tacit assumption throughout this process that the data mean what they say, so to speak—that, for example, the historical record that a patient's serum electrolyte levels had certain values or that a particular range of symptoms were recorded can be trusted as being at least a close approximation to the facts. Rather more subtly, if a particular condition is being studied, the research worker will need assurance that the cases retrieved by the recordkeeping system are homogeneous in some definable sense so as to provide a secure basis for generalization and prediction. There are all too many ways in which these basic assumptions may fail, and even those closely acquainted with the problems of data analysis do not always appreciate the need to check them or the power of the various checking techniques now available. The fallibility of routinely collected data is fairly widely recognized, but even material carefully obtained under research conditions is more susceptible than is generally realized to gross errors occurring at one of the stages of measurement, recording and transcription (see, for example, Healy, 1952). Even" obvious" items such as age and sex are not seldom misrecorded, thus indicating a disease of childhood in an 80-year-old patient or cancer of the uterus in a male.These last examples point te the need for checks on the credibility of data before they are analysed. The two examples are frankly incredible, but far more common are instances of low credibility—cases which stand so far from the main body of the material that they demand special investigation. An element of personal judgement is inevitably involved, and there is no kind of implication that the outlying items are merely to be" rejected": they often contain more of interest than more orthodox cases, but this makes their identification more rather than less important. There is a substantial statistical literature on the identification of outliers, but much of it is of rather little value. Many suggested techniques are liable to give foolish answers and their theoretical background is often inadequate (see an im-