A supervised approach for identifying discriminating genotype patterns and its application to breast cancer data

A supervised approach for identifying discriminating genotype patterns and its application to breast cancer data
复制标题

识别区分基因型模式的监督方法及其在乳腺癌数据中的应用

DOI:
--
复制
发表时间:
2007
期刊:
Bioinform.
影响因子:
--
通讯作者:
R. Sharan
R. Sharan
中科院分区:
--
文献类型:
--
作者:
N. Yosef;Z. Yakhini;A. Tsalenko;V. Kristensen;A. Børresen;E. Ruppin;R. Sharan

文献摘要

参考文献

被引文献

相似文献

动机 大规模关联研究调查了感兴趣的表型的遗传决定因素,正在产生越来越多的人类群体基因组变异数据。这些研究的一个基本挑战是检测基因型模式,将表现出所研究表型的个体与不具有该表型的个体区分开来。困难源于需要测试大量的单核苷酸多态性(SNP)组合。当同一队列获得额外的高通量数据(例如基因表达数据)时,歧视问题变得更加复杂。 结果 我们开发了一种图论方法来识别基因分型群体中给定表型的区分模式(DP)。该方法基于将 SNP 数据表示为个体及其 SNP 状态的二分图,并识别该图的完全连接的子图,这些子图将针对给定表型组富集的个体相关联。该方法可以处理其他数据类型,例如基因分型群体的表达谱。它让人想起双聚类方法,其关键区别在于其搜索过程是由所考虑的表型以监督方式指导的。我们在模拟和真实数据中测试了我们的方法。在模拟中,我们的方法能够以高成功率检索植入的模式。然后,我们将我们的方法应用于 72 名乳腺癌患者的数据集,这些患者具有可用的基因表达谱,对超过 695 个 SNP 进行了基因分型。我们检测到了几个对各种临床表型高度显着的 DP,并研究了患者组和他们定义的基因组。我们发现患者群体的其他表型高度丰富,并且在其概况之间表现出表达一致性。这些基因组表现出功能一致性,并涉及在癌症中具有已知作用的基因,为它们的参与提供了额外的支持。 可用性 该程序可根据要求提供。
MOTIVATION Large-scale association studies, investigating the genetic determinants of a phenotype of interest, are producing increasing amounts of genomic variation data on human cohorts. A fundamental challenge in these studies is the detection of genotypic patterns that discriminate individuals exhibiting the phenotype under study from individuals that do not possess it. The difficulty stems from the large number of single nucleotide polymorphism (SNP) combinations that have to be tested. The discrimination problem becomes even more involved when additional high-throughput data, such as gene expression data, are available for the same cohort. RESULTS We have developed a graph theoretic approach for identifying discriminating patterns (DPs) for a given phenotype in a genotyped population. The method is based on representing the SNP data as a bipartite graph of individuals and their SNP states, and identifying fully connected subgraphs of this graph that relate individuals enriched for a given phenotypic group. The method can handle additional data types such as expression profiles of the genotyped population. It is reminiscent of biclustering approaches with the crucial difference that its search process is guided by the phenotype under consideration in a supervised manner. We tested our approach in simulations and on real data. In simulations, our method was able to retrieve planted patterns with high success rate. We then applied our approach to a dataset of 72 breast cancer patients with available gene expression profiles, genotyped over 695 SNPs. We detected several DPs that were highly significant with respect to various clinical phenotypes, and investigated the groups of patients and the groups of genes they defined. We found the patient groups to be highly enriched for other phenotypes and to display expression coherency among their profiles. The gene groups displayed functional coherency and involved genes with known role in cancer, providing additional support to their involvement. AVAILABILITY The program is available upon request.
DOI: 10.1001/jama.291.13.1642
发表时间: 2004-04-07
影响因子: 120.7
作者:
Moore, JH;Ritchie, MD
通讯作者: Ritchie, MD
DOI: 10.1073/pnas.0307659101
发表时间: 2004-04-13
影响因子: 11.1
作者:
Sklan, EH;Lowenthal, A;Soreq, H
通讯作者: Soreq, H
DOI: 10.1001/jama.286.18.2245
发表时间: 2001-11-14
影响因子: 120.7
作者:
Martin, ER;Scott, WK;Vance, JM
通讯作者: Vance, JM