Pancreatic Expression database: a generic model for the organization, integration and mining of complex cancer datasets

Pancreatic Expression database: a generic model for the organization, integration and mining of complex cancer datasets
复制标题

DOI:
10.1186/1471-2164-8-439
复制
发表时间:
2007-11-28
期刊:
影响因子:
4.4
通讯作者:
Crnogorac-Jurcevic, Tatjana
Crnogorac-Jurcevic, Tatjana
中科院分区:
生物学2区
文献类型:
--
作者:
Chelala, Claude;Hahn, Stephan A.;Crnogorac-Jurcevic, Tatjana

文献摘要

被引文献

相似文献

背景:胰腺癌是男性和女性癌症死亡的第五大原因。近年来,大量的基因和蛋白质表达研究已经发表,拓宽了我们对胰腺癌生物学的理解。由于来自多个不同来源的公开数据的爆炸性增长,单个研究人员越来越难以将这些数据纳入其当前的研究计划。胰腺表达数据库是一个通用的基于网络的系统,旨在通过为研究界提供一个开放获取的工具来缩小这一差距,不仅可以挖掘当前可用的胰腺癌数据集,还可以将他们自己的数据纳入数据库。目前,数据库保存32个数据集,包括从20个不同的已发表基因或蛋白质表达研究中提取的7636个基因表达测量值,胰腺癌类型、胰腺前驱病变(PanIN)和慢性胰腺炎。胰腺数据与人类基因组基因和蛋白质注释、序列、同源物、SNP和抗体数据一起存储在基于BioMart技术的数据管理系统中。数据库的查询可以通过基于网络的查询界面和通过网络服务使用来自胰腺(疾病阶段、调节、差异表达、表达、平台技术、出版物)和/或公共数据(抗体、基因组区域、基因相关的登记、本体、表达模式、多物种比较、蛋白质数据、SNP)的组合标准来实现。因此,我们的数据库,否则不同的数据源之间的连接,并允许相对简单的导航之间的所有数据类型和annotations.Conclusion:数据库的结构和内容提供了一个强大的和高速的数据挖掘工具,癌症研究。它可用于目标发现,即来自体液的生物标志物,与癌症进展相关的基因的鉴定和分析,跨平台荟萃分析,用于胰腺癌关联研究的SNP选择,癌症基因启动子分析以及挖掘癌症本体信息。该数据模型是通用的,可以很容易地扩展和应用于其他类型的癌症。该数据库可在网上查阅,对科学界没有任何限制,网址是http://www.胰腺表达。org/.
Background: Pancreatic cancer is the 5th leading cause of cancer death in both males and females. In recent years, a wealth of gene and protein expression studies have been published broadening our understanding of pancreatic cancer biology. Due to the explosive growth in publicly available data from multiple different sources it is becoming increasingly difficult for individual researchers to integrate these into their current research programmes. The Pancreatic Expression database, a generic web-based system, is aiming to close this gap by providing the research community with an open access tool, not only to mine currently available pancreatic cancer data sets but also to include their own data in the database.Description: Currently, the database holds 32 datasets comprising 7636 gene expression measurements extracted from 20 different published gene or protein expression studies from various pancreatic cancer types, pancreatic precursor lesions (PanINs) and chronic pancreatitis. The pancreatic data are stored in a data management system based on the BioMart technology alongside the human genome gene and protein annotations, sequence, homologue, SNP and antibody data. Interrogation of the database can be achieved through both a web-based query interface and through web services using combined criteria from pancreatic (disease stages, regulation, differential expression, expression, platform technology, publication) and/or public data (antibodies, genomic region, gene-related accessions, ontology, expression patterns, multi-species comparisons, protein data, SNPs). Thus, our database enables connections between otherwise disparate data sources and allows relatively simple navigation between all data types and annotations.Conclusion: The database structure and content provides a powerful and high-speed data-mining tool for cancer research. It can be used for target discovery i.e. of biomarkers from body fluids, identification and analysis of genes associated with the progression of cancer, cross-platform meta-analysis, SNP selection for pancreatic cancer association studies, cancer gene promoter analysis as well as mining cancer ontology information. The data model is generic and can be easily extended and applied to other types of cancer. The database is available online with no restrictions for the scientific community at http:// www. pancreasexpression. org/.