A fast algorithm to factorize high-dimensional Tensor Product matrices used in Genetic Models

A fast algorithm to factorize high-dimensional Tensor Product matrices used in Genetic Models
复制标题

一种用于分解遗传模型中使用的高维张量积矩阵的快速算法

DOI:
10.1093/g3journal/jkae001
复制
发表时间:
2024
期刊:
影响因子:
3.3
通讯作者:
de los Campos, Gustavo
de los Campos, Gustavo
中科院分区:
生物学2区
文献类型:
--
作者:
Lopez-Cruz, Marco;Pérez-Rodríguez, Paulino;de los Campos, Gustavo

文献摘要

相似文献

许多遗传模型(包括上位效应模型以及环境遗传模型)涉及到的协方差结构是低秩矩阵的哈达玛乘积。实现这些模型需要分解大型Hadamard乘积矩阵。现有的因式分解算法不能很好地适用于大数据,这使得其中一些模型在大样本量下不可行。在这里,基于Hadamard积和(相关的)Kronecker积的性质,我们提出了一种算法,该算法产生的近似分解比标准特征值分解快几个数量级。在本文中,我们描述了该算法,展示了如何使用它来分解大型Hadamard乘积矩阵,给出了基准,并通过对gxe项目从基因组到领域计划(n ~ 60,000)的北部测试地点的数据分析来说明该方法的使用。我们在开源的“tensorEVD”R包中实现了所提出的算法。
Many genetic models (including models for epistatic effects as well as genetic-by-environment) involve covariance structures that are Hadamard products of lower rank matrices. Implementing these models requires factorizing large Hadamard product matrices. The available algorithms for factorization do not scale well for big data, making the use of some of these models not feasible with large sample sizes. Here, based on properties of Hadamard products and (related) Kronecker products, we propose an algorithm that produces an approximate decomposition that is orders of magnitude faster than the standard eigenvalue decomposition. In this article, we describe the algorithm, show how it can be used to factorize large Hadamard product matrices, present benchmarks, and illustrate the use of the method by presenting an analysis of data from the northern testing locations of the G × E project from the Genomes to Fields Initiative (n∼ 60,000). We implemented the proposed algorithm in the open-source “tensorEVD” R package.