FiT: fiber-based tensor completion for drug repurposing

FiT: fiber-based tensor completion for drug repurposing
复制标题

DOI:
10.1145/3535508.3545527
复制
发表时间:
2022-08
期刊:
Proceedings of the 13th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics
影响因子:
--
通讯作者:
Aysegül Bumin;Anna M. Ritz;D. Slonim;Tamer Kahveci;Kejun Huang
Aysegül Bumin;Anna M. Ritz;D. Slonim;Tamer Kahveci;Kejun Huang
中科院分区:
其他
文献类型:
--
作者:
Aysegül Bumin;Anna M. Ritz;D. Slonim;Tamer Kahveci;Kejun Huang

文献摘要

相似文献

药物再利用旨在为现有药物寻找新的用途。一种药物再利用方法,称为“连接映射”,将药物的转录组学图谱与表征疾病状态的图谱联系起来。然而,通过实验评估药物暴露在特定细胞中的转录组学效应是一个昂贵的过程。广泛表征药物-细胞组合进一步受到阻碍,因为原代组织样品可能不丰富,导致药物-细胞数据库中存在许多空白。为了最好地找到与特定条件相关的药物,我们可能因此想要将给定药物对未测定的细胞类型或类型的转录组学影响归责。然而,这一步骤偏离了经典的数据补全问题,因为用于该问题的最新数据插补技术的基本瓶颈不考虑数据的独特特征。数据中的缺失值不是随机分布的,基因不是独立的实体,而是相互作用并影响彼此的转录速率。在这里,我们解决了连接图数据插补问题的第一个也是最基本的部分之一,以实现药物再利用。我们开发了一种新的方法,名为FiT(基于纤维的张量完成),以准确有效地在高度稀疏的药物细胞系数据集中估算缺失药物细胞系组合的转录值,同时利用缺失值的分布以及基因之间的相互作用。我们的研究结果表明,即使在稀疏数据集上,大约75%的数据缺失,FiT也优于现有方法,并在更短的时间内获得更准确的结果。
Drug repurposing aims to find new uses for existing drugs. One drug repurposing approach, called "Connectivity Mapping," links transcriptomic profiles of drugs to profiles characterizing disease states. However, experimentally evaluating the transcriptomic effects of drug exposure in particular cells is a costly process. Characterizing drug-cell combinations widely is further hindered because primary tissue samples may not be abundant, leading to many gaps in drug-cell databases. To best find drugs relevant for particular conditions, we may therefore want to impute the transcriptomic impact of a given drug on an unassayed cell type or types. This step deviates from classic data completion problems, however, because of the fundamental bottleneck that state of the art data imputation techniques for this problem do not consider the unique characteristics of the data. The missing values in the data are not randomly distributed, and the genes are not independent entities, but rather they interact with and affect the transcription rates of one another. Here, we address the first and one of the most fundamental parts of the connectivity map data imputation problem to enable drug repurposing. We develop a novel method, named FiT (Fiber-based Tensor Completion) to impute the transcription values for missing drug-cell line combinations in a highly sparse drug-cell line dataset accurately and efficiently, while exploiting the distribution of missing values as well as the interactions among genes. Our results demonstrate that even on a sparse dataset, where approximately 75% of the data is missing, FiT outperforms existing approaches and obtains more accurate results in a significantly shorter amount of time.