Asymmetric trichotomous partitioning overcomes dataset limitations in building machine learning models for predicting siRNA efficacy.

Asymmetric trichotomous partitioning overcomes dataset limitations in building machine learning models for predicting siRNA efficacy.
复制标题

在构建机器学习模型中,不对称的三分法分区克服了数据集限制,以预测siRNA功效。

DOI:
10.1016/j.omtn.2023.06.010
复制
发表时间:
2023-09-12
期刊:
MOLECULAR THERAPY NUCLEIC ACIDS
影响因子:
--
通讯作者:
Khvorova, Anastasia
Khvorova, Anastasia
中科院分区:
其他
文献类型:
--
作者:
Monopoli, Kathryn R.;Korkin, Dmitry;Khvorova, Anastasia

文献摘要

参考文献

相似文献

化学修饰的小干扰RNA(SiRNAs)是指导疾病基因序列特异性沉默的一种很有前途的治疗方法。识别有效沉默靶基因的化学修饰的siRNA序列仍然具有挑战性。这样的确定需要计算算法。机器学习是解决生物学问题的一种强大的预测性方法,但通常需要比大多数可用的siRNA数据集大得多的数据集。在这里,我们描述了一种将机器学习应用于小数据集(356个修改的序列)以预测siRNA有效性的框架。为了克服siRNA数据集的噪声和生物学限制,我们应用了三分、双阈值的划分方法,产生了几种分类阈值对的组合。然后,我们使用一种新的考虑类别不平衡的评估指标来测试不同阈值对随机森林机器学习模型性能的影响。我们确定了产生具有高预测能力的模型的阈值,其表现优于由相同数据生成的线性模型,该线性模型在实验评估中具有预测性。使用一种新的模型特征提取方法,我们观察到目标位点的碱基重要性和碱基偏好与我们目前对siRNA介导的沉默机制的理解一致,随机森林提供了比线性模型更高的分辨率。这一框架适用于涉及小型生物数据集的任何分类挑战,为开发高性能的寡核苷酸疗法设计算法提供了机会。Khvorova和他的同事提出了一种将有监督的机器学习模型应用于有限大小的小干扰RNA数据集的方法。有监督的机器学习模型是强大的预测器,非常适合于识别强大的小干扰RNA,但通常需要大数据集。克服这一限制扩大了机器学习在支持寡核苷酸治疗开发方面的应用可能性。
Chemically modified small interfering RNAs (siRNAs) are promising therapeutics guiding sequence-specific silencing of disease genes. Identifying chemically modified siRNA sequences that effectively silence target genes remains challenging. Such determinations necessitate computational algorithms. Machine learning is a powerful predictive approach for tackling biological problems but typically requires datasets significantly larger than most available siRNA datasets. Here, we describe a framework applying machine learning to a small dataset (356 modified sequences) for siRNA efficacy prediction. To overcome noise and biological limitations in siRNA datasets, we apply a trichotomous, two-threshold, partitioning approach, producing several combinations of classification threshold pairs. We then test the effects of different thresholds on random forest machine learning model performance using a novel evaluation metric accounting for class imbalances. We identify thresholds yielding a model with high predictive power, outperforming a linear model generated from the same data, that was predictive upon experimental evaluation. Using a novel model feature extraction method, we observe target site base importances and base preferences consistent with our current understanding of the siRNA-mediated silencing mechanism, with the random forest providing higher resolution than the linear model. This framework applies to any classification challenge involving small biological datasets, providing an opportunity to develop high-performing design algorithms for oligonucleotide therapies. Khvorova and colleagues present a method applying supervised machine learning models to limited-size small interfering RNA datasets. Supervised machine learning models are powerful predictors well suited for identifying potent small interfering RNAs but typically require large datasets. Overcoming this limitation expands application possibilities of machine learning in supporting oligonucleotide therapy development.
DOI: 10.1038/nsmb780
发表时间: 2004-07-01
影响因子: 16.8
作者:
Haley, B;Zamore, PD
通讯作者: Zamore, PD
DOI: 10.1093/nar/gky037
发表时间: 2018-03-16
影响因子: 14.9
作者:
Hassler MR;Turanov AA;Alterman JF;Haraszti RA;Coles AH;Osborn MF;Echeverria D;Nikan M;Salomon WE;Roux L;Godinho BMDC;Davis SM;Morrissey DV;Zamore PD;Karumanchi SA;Moore MJ;Aronin N;Khvorova A
通讯作者: Khvorova A
DOI: 10.1021/acsomega.7b00291
发表时间: 2017-05-31
期刊: ACS omega
影响因子: 4.1
作者:
Iribe H;Miyamoto K;Takahashi T;Kobayashi Y;Leo J;Aida M;Ui-Tei K
通讯作者: Ui-Tei K
DOI: 10.1214/aos/1013203451
发表时间: 2001-10-01
影响因子: 4.5
作者:
Friedman, JH
通讯作者: Friedman, JH
DOI: 10.1093/nar/gkl1120
发表时间: 2007
影响因子: 14.9
作者:
Katoh T;Suzuki T
通讯作者: Suzuki T