Informatics and hypothesis‐driven research

Informatics and hypothesis‐driven research
复制标题

信息学和假设驱动的研究

DOI:
10.1093/embo-reports/kvf164
复制
发表时间:
2002
期刊:
影响因子:
7.7
通讯作者:
N. Smalheiser
N. Smalheiser
中科院分区:
生物学2区
文献类型:
--
作者:
N. Smalheiser

文献摘要

被引文献

相似文献

数据的价值,而不是提供显著的‘附加值’。考虑一个由信用卡交易组成的商业数据库:它的目的是跟踪个人帐户,而对数据库的大多数查询都是特定的、集中的和单独启动的。相比之下,自动化数据挖掘技术允许同一数据库以提供丰富的市场研究数据的重大大规模相关性为特征。更重要的是,人们可以在持续的基础上搜索增加欺诈可能性的异常活动模式;事实上,如果商业数据库不执行这种自动化的“数据驱动的发现”,甚至可能被视为疏忽。我建议根据具体情况填充和分析的研究数据库
of data, but rather as providing significant ‘added value’. Consider a commercial database consisting of credit-card transactions: its purpose is to keep track of individual accounts, and most of the queries to the database are specific, focused and initiated individually. In contrast, automated data-mining techniques permit the same database to be characterised in terms of significant large-scale correlations that provide a rich array of market research data. More importantly, one can search on an ongoing basis for anomalous patterns of activity that raise the possibility of fraud; in fact, a commercial database that does not carry out such automated ‘data-driven discovery’ might even be considered negligent. I suggest that research databases that are populated and analysed according to specific