Olfactory Receptor Database: a metadata-driven automated population from sources of gene and protein sequences

Olfactory Receptor Database: a metadata-driven automated population from sources of gene and protein sequences
复制标题

DOI:
10.1093/nar/30.1.354
复制
发表时间:
2002-01-01
影响因子:
14.9
通讯作者:
Shepherd, G
Shepherd, G
中科院分区:
生物学2区
文献类型:
--
作者:
Crasto, C;Marenco, L;Shepherd, G

文献摘要

被引文献

相似文献

嗅觉受体数据库(ORDB;http://senselab.med.yale.edu/senselab/ordb))是嗅觉受体(OR)和嗅觉受体样基因和蛋白质序列的中央储存库。为了处理非常大的OR基因家族,我们构建了一个算法,可以自动从GenBank和Swiss-Prot等网站下载序列到数据库中。该算法使用超文本标记语言(HTML)解析技术来提取与ORDB相关的信息。然后将信息与ORDB知识库中的元数据相关联,以将提取的非结构化文本编码为符合数据库体系结构、实体属性值与类和关系(EAV/CR)的结构化格式,这作为一个整体支持SenseLab项目。讨论了批量填充、自动填充和半自动填充三种方法。使用可扩展标记语言(XML)将数据导入数据库。
The Olfactory Receptor Database (ORDB; http://senselab.med.yale.edu/senselab/ordb) is a central repository of olfactory receptor (OR) and olfactory receptor-like gene and protein sequences. To deal with the very large OR gene family, we have constructed an algorithm that automatically downloads sequences from web sources such as GenBank and SWISS-PROT into the database. The algorithm uses hypertext markup language (HTML) parsing techniques that extract information relevant to ORDB. The information is then correlated with the metadata in the ORDB knowledge base to encode the unstructured text extracted into the structured format compliant with the database architecture, entity attribute value with classes and relationship (EAV/CR), which supports the SenseLab project as a whole. Three population methods: batch, automatic and semi-automatic population are discussed. The data is imported into the database using extensible markup language (XML).