课题基金 / 基金详情

Diagnostics for Structured Data and Quality Improvement

Diagnostics for Structured Data and Quality Improvement
结构化数据诊断和质量改进
批准号:
9803622
负责人:
Douglas Hawkins
金额:
$9.21万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1998
资助国家:
美国
项目状态:
已结题
起止时间:
1998-08-15 至 2002-07-31

项目摘要

项目成果

Douglas Hawkins的其他基金

相似基金

相关文献

中文摘要
翻译
9803622道格拉斯·M·霍金斯结构化数据集是指其中的数据不是独立的同分布标量的数据集。例如多元回归和多变量数据集,以及时间序列。诊断涉及识别偏离某个基线(通常为高斯)模型的病例。作为数据分析的一部分,这一一般框架涵盖了寻找离群值,也是统计过程控制(SPC)方法的基础。本项目的一个主线扩展了PI以前在识别离群值方面的工作,特别是在离群值众多或放置不当的情况下。最近很明显,在教科书大小的问题上效果很好的方法在数据集上毫无用处,即使在几十个维度上只有几千个案例。由于这类数据集越来越普遍,这重新打开了本项目的一个主要重点-在大型数据集中找出离奇案例的可行办法的整个问题。一个比较明显的问题是检测按时间排序的标量和多变量数据中的持续变化。这就是变点、指数加权移动平均和累积和方法所解决的问题。工作方案还包括这一领域的一项重大努力。这两个问题领域的结合导致了抗孤立异常值的统计过程控制方法的设计。在大型数据库中,不可能使用当前方法来核实条目的正确性或内部一致性。在小数据集上工作良好的方法在计算上是不可想象的,在以兆字节或更大范围内建立的数据集中,数据库中的信息质量会受到无法检测到的错误的影响。这个项目正在开发识别“离群值”的方法-大型数据库中的非典型条目-具有可接受的、尽管仍然很大的计算量。单靠更高的处理器能力并不能解决问题,但更强大的处理器与本工作计划中开发的改进算法相结合,可能会解决问题。这个问题本质上是分布式处理的-以前的工作表明,如何通过使用中央处理器阵列来加速异常值识别。这项工作的另一个主线是累积和(累计)图表,这是统计过程控制(SPC)家族中的一个工具。经典的休哈特Xbar和R控制图无法检测到微小但持续的变化。这种移位很快就会被发现,并被诊断为充血。与休哈特图表结合使用,可诊断制造问题,从而大幅提高质量。Cusum在许多其他监测情况下也是有效的,从在线医疗监测到检测空气或水中的污染羽流。例如,垃圾填埋场周围的地下水监测就是为了在高度可变的背景下检测污染物水平的增加这一问题。Cusum已经被认为是用于检测和诊断泄漏的强大工具,它们对处理未检测到的化学数据的扩展对于将其适用于低浓度下有害的重金属等污染物具有重要意义。
英文摘要
9803622Douglas M. HawkinsStructured data sets are those in which the data are other than independent identically distributed scalar quantities. Examples are multiple regression and multivariate data sets, and time series. Diagnostics involves identifying cases that depart from some baseline (usually Gaussian) model. This general framework covers finding outliers as a part of data analysis, and is also the basis for statistical process control (SPC) methodologies. One thread of the present project extends the PI's previous work in outlier identification, particularly in situations where the outliers are numerous or badly placed. It has recently become apparent that methods that work well on text-book-sized problems are useless in data sets with even a few thousand cases in a few dozen dimensions. As such data sets are increasingly common, this has reopened a major emphasis of the present project -- the whole question of workable approaches for finding outlying cases in large data sets. A somewhat distinct problem is detection of persistent changes in time-ordered scalar and multivariate data. This is the problem addressed by change-point, exponentially weighted moving average, and cumulative sum methodologies. The program of work includes a major effort in this area also. The union of the two problem areas leads to the design of statistical process control methodologies that are resistant to isolated outliers.In large data bases it is impossible to verify the correctness or internal consistency of the entries using current methodology. Methods that work well on small data sets are computationally unthinkable in data sets up in the megabyte and beyond range leaving the quality of information in data bases hostage to undetected errors. This project is developing methods to identify "outliers" --- atypical entries in large data bases --- with an acceptable, though still large, amount of computational effort. More processor power alone will not solve the problem, but more powerful processors combined with the improved algorithms developed in this program of work may do so. The problem is inherently amenable to distributed processing --- previous work showed how outlier identification could be speeded up by using an array of central processors. Another thread of the work is cumulative sum (cusum) charting, a tool in the statistical process control (SPC) family. The classic Shewhart Xbar and R control charts are incapable of detecting small but persistent shifts. Such shifts are found and diagnosed rapidly with cusums. Used in conjunction with Shewhart charts, cusums can diagnose manufacturing problems, leading to substantial quality improvement. Cusums are also effective in many other monitoring situations, from online medical monitoring to detecting plumes of pollution in air or water. Groundwater monitoring around a landfill, for example, aims at exactly this problem of detecting an increased level of pollutants against a highly variable background. Cusums are already recognized as a powerful tool to use in the detection and diagnosis of leakages, and their extension to handle non-detect chemical data is important in extending their applicability to pollutants like heavy metals that are harmful at low concentrations.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Diagnostics for structured data and quality improvement
  • 批准号:
    0306304
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.61万
  • 财政年份:
    2003
  • 负责人:
    Douglas Hawkins
  • 依托单位:
Mathematical Sciences: Diagnostics in Structured Data and Quality Improvement
  • 批准号:
    9505440
  • 项目类别:
    Standard Grant
  • 资助金额:
    $11.4万
  • 财政年份:
    1995
  • 负责人:
    Douglas Hawkins
  • 依托单位:
Mathematical Sciences: Diagnostics and Graphics for Structured Data
  • 批准号:
    9208819
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.5万
  • 财政年份:
    1992
  • 负责人:
    Douglas Hawkins
  • 依托单位:
Mathematical Sciences: Location of Outliers in Structured Data
  • 批准号:
    9010983
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $6.5万
  • 财政年份:
    1990
  • 负责人:
    Douglas Hawkins
  • 依托单位:
海外基金