Reconfigurable Web wrapper agents for biological information integration

Reconfigurable Web wrapper agents for biological information integration
复制标题

DOI:
10.1002/asi.20139
复制
发表时间:
2005-03-01
影响因子:
--
通讯作者:
Chang, CC
Chang, CC
中科院分区:
其他
文献类型:
--
作者:
Hsu, CN;Chang, CH;Chang, CC

文献摘要

被引文献

相似文献

各种生物数据在万维网上以压倒性的数量进行传输和交换。如何快速捕获、利用和整合互联网上的信息以发现有价值的生物知识是生物信息学中最关键的问题之一。已经提出了许多信息集成系统用于集成生物数据。这些系统通常依赖于称为包装器的中间软件层来访问连接的信息源。Web数据源的包装器构造通常是专门手工编写的,以适应每个Web站点之间的差异。然而,编写Web包装器需要大量的编程技能,并且很耗时且难以维护。在这篇文章中,我们提供了一个快速构建软件代理,可以作为生物信息集成的Web包装器的解决方案。我们定义了一个基于XML的语言称为Web导航描述语言(WNDL),建模Web浏览会话。WNDL脚本描述了如何定位数据、提取数据以及联合收割机组合数据。通过执行不同的WNDL脚本,我们可以自动化几乎所有类型的Web浏览会话。我们还描述了IEPAD(基于模式发现的信息提取),一个基于模式发现技术的数据提取器。IEPAD允许我们的软件代理自动发现提取规则,以提取结构化格式化网页的内容。使用示例编程创作工具,用户可以通过浏览目标Web站点来生成完整的Web包装器代理。我们构建了各种生物应用程序来证明我们方法的可行性。
A variety of biological data is transferred and exchanged in overwhelming volumes on the World Wide Web. How to rapidly capture, utilize, and integrate the information on the Internet to discover valuable biological knowledge is one of the most critical issues in bioinformatics. Many information integration systems have been proposed for integrating biological data. These systems usually rely on an intermediate software layer called wrappers to access connected information sources. Wrapper construction for Web data sources is often specially hand coded to accommodate the differences between each Web site. However, programming a Web wrapper requires substantial programming skill, and is time-consuming and hard to maintain. In this article we provide a solution for rapidly building software agents that can serve as Web wrappers for biological information integration. We define an XMIL-based language called Web Navigation Description Language (WNDL), to model a Web-browsing session. A WNDL script describes how to locate the data, extract the data, and combine the data. By executing different WNDL scripts, we can automate virtually all types of Web-browsing sessions. We also describe IEPAD (information Extraction Based on Pattern Discovery), a data extractor based on pattern discovery techniques. IEPAD allows our software agents to automatically discover the extraction rules to extract the contents of a structurally formatted Web page. With a programming-by-example authoring tool, a user can generate a complete Web wrapper agent by browsing the target Web sites. We built a variety of biological applications to demonstrate the feasibility of our approach.