ALSBMF: Predicting lncRNA-Disease Associations by Alternating Least Squares Based on Matrix Factorization

ALSBMF: Predicting lncRNA-Disease Associations by Alternating Least Squares Based on Matrix Factorization
复制标题

ALSBMF:通过基于矩阵分解的交替最小二乘法预测 lncRNA-疾病关联

DOI:
10.1109/access.2020.2970069
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Wu Fang-Xiang
Wu Fang-Xiang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhu Wen;Huang Kaimei;Xiao Xiaofang;Liao Bo;Yao Yuhua;Wu Fang-Xiang

文献摘要

相似文献

近年来,人们越来越清楚长链非编码RNA(longnoncodingRNAs,lncRNA)能够在转录水平、翻译水平等多个水平上调控靶基因,在细胞分化、染色质重塑等重要生物学过程中发挥重要的调控作用。推断潜在的lncRNA-疾病关联对于揭示疾病背后的秘密、开发新药和优化个性化治疗至关重要。然而,验证lncRNA-疾病关联的生物学实验非常耗时且昂贵。因此,开发有效的计算模型至关重要。在这项研究中,我们提出了一种基于矩阵分解的交替最小二乘法来预测lncRNA-疾病关联的方法,称为ALSBMF。ALSBMF首先将已知的lncRNA-疾病关联矩阵分解为两个特征矩阵,然后利用疾病语义相似度、lncRNA功能相似度和已知的lncRNA-疾病关联度定义优化函数,利用最小二乘法求解两个最优特征矩阵。最后将两个最佳特征矩阵相乘以重建评分矩阵,填充原始矩阵的缺失值以预测lncRNA-疾病关联。与现有方法相比,ALSBMF具有与BPLLDA相同的优点。它不需要阴性样本,并且可以预测与新型lncRNA或新型疾病相关的关联。此外,本研究进行了留一交叉验证(LOOCV)和五折交叉验证来评估ALSBMF的预测性能。AUC分别为0.9501和0.9215,优于现有方法。此外,结肠癌、肾癌和肝癌被选为病例研究。预测的前三位结肠癌、肾癌和肝癌相关lncRNA在最新的lncRNADisease数据库和相关文献中得到验证。为了测试ALSBMF预测新的疾病相关的lncRNA和新的lncRNA相关的疾病的能力,排除所有已知的疾病和lncRNA的关联,在PubMed和dbSNP中验证预测的前五名乳腺癌、鼻咽癌癌症相关的lncRNA和前五名H19、MALAT 1 lncRNA相关的癌症。
In recent years, it has been increasingly clear that long non-coding RNAs (lncRNAs) are able to regulate their target genes at multi-levels, including transcriptional level, translational level, etc and play key regulatory roles in many important biological processes, such as cell differentiation, chromatin remodeling and more. Inferring potential lncRNA-disease associations is essential to reveal the secrets behind diseases, develop novel drugs, and optimize personalized treatments. However, biological experiments to validate lncRNA-disease associations are very time-consuming and costly. Thus, it is critical to develop effective computational models. In this study, we have proposed a method by alternating least squares based on matrix factorization to predict lncRNA-disease associations, referred to as ALSBMF. ALSBMF first decomposes the known lncRNA-disease correlation matrix into two characteristic matrices, then defines the optimization function using disease semantic similarity, lncRNA functional similarity and known lncRNA-disease associations and solves two optimal feature matrices by least squares method. The two optimal feature matrices are finally multiplied to reconstruct the scoring matrix, filling the missing values of the original matrix to predict lncRNA-disease associations. Compared to existing methods, ALSBMF has the same advantages as BPLLDA. It does not require negative samples and can predict associations related to novel lncRNAs or novel diseases. In addition, this study performs leave-one-out cross-validation (LOOCV) and five-fold cross-validation to evaluate the prediction performance of ALSBMF. The AUCs are 0.9501 and 0.9215, respectively, which are better than the existing methods. Furthermore colon cancer, kidney cancer, and liver cancer are selected as case studies. The predicted top three colon cancer, kidney cancer, and liver cancer-related lncRNAs were validated in the latest LncRNADisease database and related literature. In order to test the ability of ALSBMF to predict novel disease-associated lncRNAs and new lncRNA-associated diseases, all known associations of diseases and lncRNAs were eliminated, the predicted top five breast cancer, nasopharyngeal carcinoma cancer-related lncRNAs and top five H19, MALAT1 lncRNA-related cancers were validated in PubMed and dbSNP.