On reasoning from data

On reasoning from data
复制标题

从数据推理

DOI:
10.1145/212094.212127
复制
发表时间:
1995
期刊:
ACM Comput. Surv.
影响因子:
--
通讯作者:
S. Kasif
S. Kasif
中科院分区:
--
文献类型:
--
作者:
D. Waltz;S. Kasif

文献摘要

被引文献

相似文献

我们的社会目前正在进入一个新的阶段,在这个阶段中,千兆字节的信息随时可以通过学术网络、数字图书馆、商业信息服务以及专有的商业和政府数据库进行探索。这一重要的技术发展提出了一个巨大的挑战,因为未来的智能系统必须能够存储非常大的数据流,使用简洁高效的模型总结和索引这些数据,并随后执行非常有效的检索和推理,以响应实时查询和更新。我们非正式地把这个具有挑战性的任务称为数据推理。大多数以前的人工智能研究和应用都集中在相对简单的操作上,例如,对相对静态、不可变的知识系统(如数学、国际象棋和硬件组件清单)进行高度约束的查询,在这些系统中,可以抽象出可以被视为真实有效的规则。还有许多其他领域的数据变化或快或慢,其中抽象的真理充其量是暂时的或偶然的,例如,机器人环境、软件环境、人口数据库和公共卫生数据、生态和经济学(生态系统、化学过程、营销和销售点数据库、金融时间序列、视频和文本数据库)。此外,这些领域与对意外查询的快速响应和对不确定、动态、交互式和快速变化的环境的持续更新的需求相关联。这些领域对纯符号的、基于规则的人工智能方法提出了挑战。例如,对诸如重要的电子信息、公平的日程安排、紧急电话、前往夏威夷的优质旅行套餐、关于贝叶斯推理的有趣新论文、高风险汽车、良好的房地产投资或有趣的经济趋势等概念给出正式的逻辑规范似乎很困难。统计决策理论[Pearl 1988]为在随机和快速发展的领域中建立自适应智能代理模型提供了一个有用的框架。此外,它提供了精确的标准(损失函数,期望效用)来评估这些代理的性能。i30然而,当环境很大时,将良好的模型(寻找最大后验模型甚至最大似然模型)拟合到环境产生的数据的过程通常在计算上是难以处理的。当环境很小的时候,我们常常难以获得足够的统计数据。因此,我们可以有效地设计的模型很少是准确的,无论大小
Our society is currently entering a new phase in which gigabytes of’ information are becoming readily available for exploration over academic networks, digital libraries, and commercial information services as well as in proprietary commercial and governmental databases. This important technological development presents a substantial challenge, as future intelligent systems must be able to store very large streams of data, summarize and index this data using concise and efficient models, and subsequently perform very efficient retrieval and reasoning in response to real-time queries and updates. We informally refer to this challenging task as reasoning from data. Most previous AI research and applications have concentrated on relatively simple operations, for example, highly constrained queries on relatively static, immutable systems of knowledge such as mathematics, chess, and hardware components inventories, where it is possible to abstract rules that can be viewed as true and valid. There are many other domains in which data changes more or less rapidly and in which abstract truths are at best temporary or contingent, for example, robot environments, software environments, demographic databases and public-health data, ecological and economics (ecosystems, chemical processes, marketing and point-of-sale databases, financial time series, and video and text databases. In addition, these domains are associated with a demand for very fast response to unanticipated queries and continuous updates over uncertain, dynamic, interactive, and rapidly changing environments. These domains present a challenge for purely symbolic, rule-based approaches to AI. For instance, it appears to be difficult to give a formal logical specification of concepts such as an important electronic message, a fair scheduler, an urgent phone call, a good travel package to Hawaii, an intriguing new paper about Bayesian reasoning, a high-risk car, a good real-estate investment, or an interesting economic trend. Statistical decision theory [Pearl 1988] provides a useful framework to model adaptive intelligent agents in stochastic and rapidly evolving domains. Moreover, it provides precise criteria (loss functions, expected utility) to evaluate the performance of such agents. i30wever, when the environment is large, the process of fitting good models (finding maximum u posterior models or even maximum likelihood models) to data generated by the environment is typically computationally intractable. When the envwonment is small, we often have trouble getting sufficient statistics. Thus the models we can devise effectively are rarely accurate, regardless of the size of