High-dimension to high-dimension screening for detecting genome-wide epigenetic and noncoding RNA regulators of gene expression

High-dimension to high-dimension screening for detecting genome-wide epigenetic and noncoding RNA regulators of gene expression
复制标题

高维到高维筛选,用于检测基因表达的全基因组表观遗传和非编码 RNA 调节因子

DOI:
10.1093/bioinformatics/btac518
复制
发表时间:
2022
期刊:
影响因子:
5.8
通讯作者:
Alkan, ed., Can
Alkan, ed., Can
中科院分区:
生物学3区
文献类型:
--
作者:
Ke, Hongjie;Ren, Zhao;Qi, Jianfei;Chen, Shuo;Tseng, George C.;Ye, Zhenyao;Ma, Tianzhou;Alkan, ed., Can

文献摘要

被引文献

相似文献

高通量技术的进步表明,基因组中有多种表观遗传修饰和非编码rna通过调节基因表达参与疾病的发病机制。表观遗传/非编码RNA和基因表达数据的高维性使得识别基因的重要调节因子具有挑战性。对每一对可能的调控基因对进行单变量检验存在严重的多重比较负担,直接应用正则化方法选择调控基因对在计算上是不可行的。在正则化之前先进行快速筛选降维比单独使用正则化方法更有效、更稳定。我们提出了一种基于稳健偏相关的新型筛选方法,用于检测全基因组基因表达的表观遗传和非编码RNA调控因子,这一问题包括高维预测因子和高维响应。与现有的筛选方法相比,我们的方法在概念上是创新的,它降低了预测和响应的维度,并在节点(调节因子或基因)和边缘(调节因子-基因对)水平上进行筛选。我们开发了数据驱动程序来确定条件集和最佳筛选阈值,并实现了快速迭代算法。对长链非编码RNA和microRNA在肾癌中的调控和多形性胶质母细胞瘤中DNA甲基化调控的模拟和应用表明了我们方法的有效性和优越性。可用性和实现本文使用的R包、相关源代码和真实数据集提供于https://github.com/kehongjie/rPCor.Supplementary information .补充数据可在bioinformaticsonline获得。
MotivationThe advancement of high-throughput technology characterizes a wide variety of epigenetic modifications and noncoding RNAs across the genome involved in disease pathogenesis via regulating gene expression. The high dimensionality of both epigenetic/noncoding RNA and gene expression data make it challenging to identify the important regulators of genes. Conducting univariate test for each possible regulator–gene pair is subject to serious multiple comparison burden, and direct application of regularization methods to select regulator–gene pairs is computationally infeasible. Applying fast screening to reduce dimension first before regularization is more efficient and stable than applying regularization methods alone.ResultsWe propose a novel screening method based on robust partial correlation to detect epigenetic and noncoding RNA regulators of gene expression over the whole genome, a problem that includes both high-dimensional predictors and high-dimensional responses. Compared to existing screening methods, our method is conceptually innovative that it reduces the dimension of both predictor and response, and screens at both node (regulators or genes) and edge (regulator–gene pairs) levels. We develop data-driven procedures to determine the conditional sets and the optimal screening threshold, and implement a fast iterative algorithm. Simulations and applications to long noncoding RNA and microRNA regulation in Kidney cancer and DNA methylation regulation in Glioblastoma Multiforme illustrate the validity and advantage of our method.Availability and implementationThe R package, related source codes and real datasets used in this article are provided at https://github.com/kehongjie/rPCor.Supplementary informationSupplementary data are available atBioinformaticsonline.