Combined analysis of chromosomal instabilities and gene expression for colon cancer progression inference.

Combined analysis of chromosomal instabilities and gene expression for colon cancer progression inference.
复制标题

DOI:
10.1186/2043-9113-4-2
复制
发表时间:
2014-01-24
期刊:
Journal of clinical bioinformatics
影响因子:
--
通讯作者:
Antoniotti M
Antoniotti M
中科院分区:
其他
文献类型:
--
作者:
Cava C;Zoppis I;Gariboldi M;Castiglioni I;Mauri G;Antoniotti M

文献摘要

被引文献

相似文献

拷贝数改变(CNA)是遗传变异的重要组成部分。这种改变与某些类型的癌症有关,包括胰腺癌、结肠癌和乳腺癌等。CNAs作为肿瘤预后的生物标志物已在多项研究中得到应用,但关于CNAs与疾病进展关系的研究报道较少。此外,大多数研究没有考虑以下两个重要问题。(I)鉴定负责表达调控的基因中的CNA对于确定导致恶性转化和进展的遗传事件是至关重要的。(II)大多数真实的域最好用结构化数据来描述,其中多种类型的实例以复杂的方式相互关联。我们的主要兴趣是检查结直肠癌(CRC)进展推断是否在考虑(I)具有CNA的基因的表达水平和(II)由于改变的基因的表达水平差异而导致的患者之间的关系(即相异性)时受益。当受试者仅通过可用属性值集(即基因表达水平)表示时,我们首先评估最先进的推理方法(支持向量机)的准确性性能。然后,我们检查是否推理精度提高,明确利用上述信息。我们的研究结果表明,CRC进展的推断改善时,组合的数据(即CNA和表达水平)和所考虑的相异性措施。通过我们的方法,分类是直观的吸引力,可以方便地获得在由此产生的相异度空间。使用来自基因表达综合数据库(GEO)的不同公共数据集来验证结果。
Copy number alterations (CNAs) represent an important component of genetic variations. Such alterations are related with certain type of cancer including those of the pancreas, colon, and breast, among others. CNAs have been used as biomarkers for cancer prognosis in multiple studies, but few works report on the relation of CNAs with the disease progression. Moreover, most studies do not consider the following two important issues. (I) The identification of CNAs in genes which are responsible for expression regulation is fundamental in order to define genetic events leading to malignant transformation and progression. (II) Most real domains are best described by structured data where instances of multiple types are related to each other in complex ways. Our main interest is to check whether the colorectal cancer (CRC) progression inference benefits when considering both (I) the expression levels of genes with CNAs, and (II) relationships (i.e. dissimilarities) between patients due to expression level differences of the altered genes. We first evaluate the accuracy performance of a state-of-the-art inference method (support vector machine) when subjects are represented only through sets of available attribute values (i.e. gene expression level). Then we check whether the inference accuracy improves, when explicitly exploiting the information mentioned above. Our results suggest that the CRC progression inference improves when the combined data (i.e. CNA and expression level) and the considered dissimilarity measures are applied. Through our approach, classification is intuitively appealing and can be conveniently obtained in the resulting dissimilarity spaces. Different public datasets from Gene Expression Omnibus (GEO) were used to validate the results.