Double feature selection and cluster analyses in mining of microarray data from cotton

Double feature selection and cluster analyses in mining of microarray data from cotton
复制标题

DOI:
10.1186/1471-2164-9-295
复制
发表时间:
2008-06-20
期刊:
影响因子:
4.4
通讯作者:
Wilkins, Thea A.
Wilkins, Thea A.
中科院分区:
生物学2区
文献类型:
--
作者:
Alabady, Magdy S.;Youn, Eunseog;Wilkins, Thea A.

文献摘要

被引文献

相似文献

背景:棉纤维是一种单细胞的种子毛状体,具有重要的生物学和经济学意义。近年来,基因组学方法如基于微阵列的表达谱被用于研究纤维的生长和发育,以在分子水平上理解纤维的发育机制。所产生的大量微阵列表达数据需要复杂的数据挖掘手段,以便提取解决生物学兴趣的基本问题的新信息。接近微阵列数据挖掘的方法之一是增加分析的维度/水平的数量,例如比较来自不同基因型的独立研究。然而,增加维度也创造了一个挑战,在寻找新的方法来分析多维微阵列data.Results:挖掘独立的微阵列研究从皮马和高地(TMI)棉花使用双功能选择和聚类分析确定的物种特异性和阶段特异性的基因转录,主张有利于离散的遗传机制,支配这两个栽培种的棉花纤维形态发生的发育编程。双特征选择分析确定了最高数量的差异表达基因,区分纤维转录组的发展Pima和TMI纤维。这些结果是基于这样的发现,即花后17和24天(dpa)之间收获的纤维差异代表了两个物种之间的最大表达距离。这种强大的选择方法确定了一个子集的基因表达的主要(PCW)和次生(SCW)细胞壁生物发生在皮马纤维,表现出的表达模式,通常是在TMI在相同的发展阶段逆转。聚类分析和功能分析表明,这些基因主要在过渡阶段受到调控,过渡阶段与PCW的终止和SCW生物发生的开始重叠,这表明这些特定的基因在Pima和TMI之间纤维性状表型差异的遗传机制中起着重要作用。双特征选择分析的新应用导致了物种和阶段特异性遗传表达模式的发现,其与Pima和TMI中纤维表型差异的遗传程序在生物学上相关。这些结果有望对正在进行的改善棉花纤维性状的努力产生深远的影响。
Background: Cotton fiber is a single-celled seed trichome of major biological and economic importance. In recent years, genomic approaches such as microarray-based expression profiling were used to study fiber growth and development to understand the developmental mechanisms of fiber at the molecular level. The vast volume of microarray expression data generated requires a sophisticated means of data mining in order to extract novel information that addresses fundamental questions of biological interest. One of the ways to approach microarray data mining is to increase the number of dimensions/levels to the analysis, such as comparing independent studies from different genotypes. However, adding dimensions also creates a challenge in finding novel ways for analyzing multi-dimensional microarray data.Results: Mining of independent microarray studies from Pima and Upland (TMI) cotton using double feature selection and cluster analyses identified species-specific and stage-specific gene transcripts that argue in favor of discrete genetic mechanisms that govern developmental programming of cotton fiber morphogenesis in these two cultivated species. Double feature selection analysis identified the highest number of differentially expressed genes that distinguish the fiber transcriptomes of developing Pima and TMI fibers. These results were based on the finding that differences in fibers harvested between 17 and 24 day post-anthesis (dpa) represent the greatest expressional distance between the two species. This powerful selection method identified a subset of genes expressed during primary (PCW) and secondary (SCW) cell wall biogenesis in Pima fibers that exhibits an expression pattern that is generally reversed in TMI at the same developmental stage. Cluster and functional analyses revealed that this subset of genes are primarily regulated during the transition stage that overlaps the termination of PCW and onset of SCW biogenesis, suggesting that these particular genes play a major role in the genetic mechanism that underlies the phenotypic differences in fiber traits between Pima and TMI.Conclusion: The novel application of double feature selection analysis led to the discovery of species- and stage-specific genetic expression patterns, which are biologically relevant to the genetic programs that underlie the differences in the fiber phenotypes in Pima and TMI. These results promise to have profound impacts on the ongoing efforts to improve cotton fiber traits.