课题基金 / 基金详情

Pilot Projects to Explore Large Data Sets

Pilot Projects to Explore Large Data Sets
探索大数据集的试点项目
批准号:
9700867
负责人:
Jerome Sacks
金额:
$81.06万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-06-15 至 2001-05-31

项目摘要

项目成果

Jerome Sacks的其他基金

相似基金

相关文献

中文摘要
翻译
大数据集是现代工业、技术和科学的必然要求,需要统计学的关注,以可靠和计算上可行的方式来处理诸如检测新模式和关系等核心问题。虽然需要是显而易见的,但将统计科学应用于大型数据集的途径并没有明确的规划,部分原因是数据的庞大数量阻碍了标准技术的直接使用。该提案包括两个相互关联的试点项目,旨在刺激统计科学进入这个快速变化的地形。每个城市都有自己的主要产业;L合伙人,它已经承诺了大量的资源。项目和合作伙伴是药物发现(Glaxo Wellcome, Research Traingle Park, NC)和电信欺诈(at&t实验室,Murray Hill, NJ)。在这两种情况下,都存在对整个行业具有高风险影响的特定科学问题;每一种都是推测性的,也就是说,从数据到信息再到知识的路径是事先不知道的。在药物发现中,关键问题是找到新的有效的、无毒的化合物。新的统计方法将被开发出来,用于搜索由机器人合成和筛选的最新进展产生的大量数据集,以便在高度复杂(和非常高维)的分子描述中识别化合物的关键特征,从而得出更好的化合物。在每天数以亿计的长途电话中,有一小部分是非法的或未经授权的,但每年造成数亿美元的损失。问题是尽可能快地检测和描述不寻常的和潜在欺诈的交易模式,特别注意控制错误率(假警报、检测失败),持续做出的大量决策加剧了这个问题。这两个问题领域具有共同的特点:对数据中重要或不寻常模式的顺序统计搜索;在绝对数量上、在每个数据点的高维描述上或在其来源的多样性上高度复杂的数据。这些问题将由分布在不同地点的跨学科研究小组完成,并由NISS密切管理。该GOALI项目由MPS多学科活动办公室(OMA)和数学科学司(DMS)联合支持。
英文摘要
Large data sets are a given in modern industry, technology and science, and demand statistical attention to treat such central issues as detecting new patterns and relationships in a reliable and computationally feasible way. While the need is apparent, paths to bringing statistical science to bear on large data sets are not clearly mapped, in part because the sheer volume of the data prevents direct use of standard techniques. This proposal comprises a pair of interconnected pilot projects targeted to stimulate statistical science entree to this rapidly changing terrain. Each has a major industria;l partner, which has committed substantial resources. The projects and partners are Drug Discovery (Glaxo Wellcome, Research Traingle Park, NC) and Telecommunications Fraud (AT&T Laboratories, Murray Hill, NJ). In both instances there are specific scientific isues with high-stakes implications for the industry at large; each is speculative, in the sense that the path from data to information to knowledge is not known in advance. In drug discovery, the critical problem is to find new potent, non-toxic compounds. New statistical methods will be developed to search substantial data sets generated by recent advances in robotic synthesis and screening in order to identify key features of compounds, leading to better ones, in the presence of highly complex (and very high-dimensional) descriptions of the molecules. Within the hundreds of millions of long distance calls per day a small fraction are illegal or unauthorized, but result in costs of hundreds of millions of dollars annually. The problem is to detect and characterize, as rapidly as possible, patterns of transactions that are unusual, and potentially fraudulent, with special attention to controlling error rates (false alarms, failures to detect), an issue exacerbated by the enormous multiplicity of decisions made on a continual basis. The two problem areas have common features: sequential statistical search for important or unusual patters in the data; data that are highly complex in sheer quantity, in high-dimensional description of each data point, or in the diversity of their sources. Approaching these issues will be done by teams of cross-disciplinary researchers at distributed sites and closely managed by NISS. This GOALI project is jointly supported by the MPS Office of Multidisciplinary Activities (OMA) and the Division of Mathematical Sciences (DMS).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Framework for Statistical Evaluation of Complex Computer Models
Workshop on Statistics and Information Technology
Postdoctoral Fellows at the National Institute of Statistical Sciences
Analysis, Exploration and Inference in Large Educational Data Sets
海外基金