CamurWeb: a classification software and a large knowledge base for gene expression data of cancer

CamurWeb: a classification software and a large knowledge base for gene expression data of cancer
复制标题

CamurWeb:癌症基因表达数据的分类软件和大型知识库

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
3
通讯作者:
G. Felici
G. Felici
中科院分区:
生物学4区
文献类型:
--
作者:
Emanuel Weitschek;Silvia Di Lauro;Eleonora Cappelli;P. Bertolazzi;G. Felici

文献摘要

被引文献

相似文献

下一代测序数据的高速增长目前需要新的知识提取方法。特别是,RNA测序基因表达实验技术在癌症病例对照研究中脱颖而出,可以通过监督机器学习技术来解决这一问题,该技术能够提取由基因组成的人类可解释模型及其与所研究疾病的关系。最先进的基于规则的分类器被设计为提取单个分类模型,可能由很少的相关基因组成。相反,我们的目标是创建一个由许多基于规则的模型组成的大型知识库,从而确定哪些基因可能参与所分析的肿瘤。这一全面和开放的知识库是传播有关癌症的新见解所必需的。我们提出CamurWeb,一种新的方法和基于Web的软件,能够提取多个和等价的分类模型的形式的逻辑公式(“如果然后”规则),并创建一个知识库,这些规则,可以查询和分析。该方法是基于迭代分类过程和自适应特征消除技术,使许多基于规则的模型相关的癌症研究的计算。此外,CamurWeb还包括一个用户友好的界面,用于运行软件、查询结果和管理已执行的实验。用户可以创建自己的个人资料,上传自己的基因表达数据,运行分类分析,并使用预定义的查询解释结果。为了验证该软件,我们将其应用于来自癌症基因组图谱数据库的所有公开可用的RNA测序数据集,从而获得关于癌症的大型开放获取知识库。CamurWeb可在http://bioinformatics.iasi.cnr.it/camurweb上找到。实验证明了CamurWeb的有效性,获得了许多分类模型,从而获得了与21种不同癌症类型相关的几个基因。最后,关于癌症的全面知识库和软件工具在网上发布;感兴趣的研究人员可以免费使用它们进行进一步研究和设计癌症研究中的生物实验。
The high growth of Next Generation Sequencing data currently demands new knowledge extraction methods. In particular, the RNA sequencing gene expression experimental technique stands out for case-control studies on cancer, which can be addressed with supervised machine learning techniques able to extract human interpretable models composed of genes, and their relation to the investigated disease. State of the art rule-based classifiers are designed to extract a single classification model, possibly composed of few relevant genes. Conversely, we aim to create a large knowledge base composed of many rule-based models, and thus determine which genes could be potentially involved in the analyzed tumor. This comprehensive and open access knowledge base is required to disseminate novel insights about cancer. We propose CamurWeb, a new method and web-based software that is able to extract multiple and equivalent classification models in form of logic formulas (“if then” rules) and to create a knowledge base of these rules that can be queried and analyzed. The method is based on an iterative classification procedure and an adaptive feature elimination technique that enables the computation of many rule-based models related to the cancer under study. Additionally, CamurWeb includes a user friendly interface for running the software, querying the results, and managing the performed experiments. The user can create her profile, upload her gene expression data, run the classification analyses, and interpret the results with predefined queries. In order to validate the software we apply it to all public available RNA sequencing datasets from The Cancer Genome Atlas database obtaining a large open access knowledge base about cancer. CamurWeb is available at http://bioinformatics.iasi.cnr.it/camurweb. The experiments prove the validity of CamurWeb, obtaining many classification models and thus several genes that are associated to 21 different cancer types. Finally, the comprehensive knowledge base about cancer and the software tool are released online; interested researchers have free access to them for further studies and to design biological experiments in cancer research.