Inferring sparse structure in genotype-phenotype maps.

Inferring sparse structure in genotype-phenotype maps.
复制标题

推断基因型-表型图中的稀疏结构。

DOI:
10.1093/genetics/iyad127
复制
发表时间:
2023
期刊:
影响因子:
3.3
通讯作者:
Desai,MichaelM
Desai,MichaelM
中科院分区:
生物学2区
文献类型:
--
作者:
Petti,Samantha;Reddy,Gautam;Desai,MichaelM

文献摘要

相似文献

相关个体之间的多个表型之间的相关性可能反映了共享遗传结构的某种模式:单个遗传位点影响多个表型(称为多效性的效应),从而在表型之间产生可观察到的关系。一个自然的假设是,多效性效应反映了一组相对较小的常见“核心”细胞过程:每个遗传位点影响一个或几个核心过程,而这些核心过程反过来又决定了观察到的表型。在这里,我们提出了一种方法来推断这种结构的基因型-表型数据。我们的方法,稀疏结构发现(SSD)是基于惩罚矩阵分解,旨在确定潜在的结构,是低维的(比表型和遗传位点的核心进程少得多),位点稀疏(每个位点影响几个核心进程),和/或表型稀疏(每个表型是由几个核心进程的影响)。我们使用稀疏性作为矩阵分解的指导是由一种新的经验测试的结果,表明在最近的几个基因型-表型数据集的稀疏结构的证据。首先,我们使用合成数据表明,我们的SSD方法可以准确地恢复核心进程,如果每个基因位点影响几个核心进程,或者如果每个表型是由几个核心进程的影响。接下来,我们将该方法应用于三个数据集,包括酵母中的适应性突变,人类细胞系中的遗传毒素稳健性测定,以及从酵母杂交中鉴定的遗传位点,并评估所鉴定的核心过程的生物相容性。更一般地说,我们提出稀疏性作为指导解决潜在结构的经验基因型-表型图。
Correlation among multiple phenotypes across related individuals may reflect some pattern of shared genetic architecture: individual genetic loci affect multiple phenotypes (an effect known as pleiotropy), creating observable relationships between phenotypes. A natural hypothesis is that pleiotropic effects reflect a relatively small set of common “core” cellular processes: each genetic locus affects one or a few core processes, and these core processes in turn determine the observed phenotypes. Here, we propose a method to infer such structure in genotype–phenotype data. Our approach,sparse structure discovery(SSD) is based on a penalized matrix decomposition designed to identify latent structure that is low-dimensional (many fewer core processes than phenotypes and genetic loci), locus-sparse (each locus affects few core processes), and/or phenotype-sparse (each phenotype is influenced by few core processes). Our use of sparsity as a guide in the matrix decomposition is motivated by the results of a novel empirical test indicating evidence of sparse structure in several recent genotype–phenotype datasets. First, we use synthetic data to show that our SSD approach can accurately recover core processes if each genetic locus affects few core processes or if each phenotype is affected by few core processes. Next, we apply the method to three datasets spanning adaptive mutations in yeast, genotoxin robustness assay in human cell lines, and genetic loci identified from a yeast cross, and evaluate the biological plausibility of the core process identified. More generally, we propose sparsity as a guiding prior for resolving latent structure in empirical genotype–phenotype maps.