Asymmetric trichotomous partitioning overcomes dataset limitations in building machine learning models for predicting siRNA efficacy.
Asymmetric trichotomous partitioning overcomes dataset limitations in building machine learning models for predicting siRNA efficacy.
复制标题
在构建机器学习模型中,不对称的三分法分区克服了数据集限制,以预测siRNA功效。
DOI:
10.1016/j.omtn.2023.06.010
复制
发表时间:
2023-09-12
期刊:
影响因子:
--
通讯作者:
Khvorova, Anastasia
中科院分区:
文献类型:
--
作者:
Monopoli, Kathryn R.;Korkin, Dmitry;Khvorova, Anastasia
Chemically modified small interfering RNAs (siRNAs) are promising therapeutics guiding sequence-specific silencing of disease genes. Identifying chemically modified siRNA sequences that effectively silence target genes remains challenging. Such determinations necessitate computational algorithms. Machine learning is a powerful predictive approach for tackling biological problems but typically requires datasets significantly larger than most available siRNA datasets. Here, we describe a framework applying machine learning to a small dataset (356 modified sequences) for siRNA efficacy prediction. To overcome noise and biological limitations in siRNA datasets, we apply a trichotomous, two-threshold, partitioning approach, producing several combinations of classification threshold pairs. We then test the effects of different thresholds on random forest machine learning model performance using a novel evaluation metric accounting for class imbalances. We identify thresholds yielding a model with high predictive power, outperforming a linear model generated from the same data, that was predictive upon experimental evaluation. Using a novel model feature extraction method, we observe target site base importances and base preferences consistent with our current understanding of the siRNA-mediated silencing mechanism, with the random forest providing higher resolution than the linear model. This framework applies to any classification challenge involving small biological datasets, providing an opportunity to develop high-performing design algorithms for oligonucleotide therapies. Khvorova and colleagues present a method applying supervised machine learning models to limited-size small interfering RNA datasets. Supervised machine learning models are powerful predictors well suited for identifying potent small interfering RNAs but typically require large datasets. Overcoming this limitation expands application possibilities of machine learning in supporting oligonucleotide therapy development.
登录
查看更多内容
影响因子:
16.8
作者:
Haley, B;Zamore, PD
通讯作者:
Zamore, PD
影响因子:
14.9
作者:
Hassler MR;Turanov AA;Alterman JF;Haraszti RA;Coles AH;Osborn MF;Echeverria D;Nikan M;Salomon WE;Roux L;Godinho BMDC;Davis SM;Morrissey DV;Zamore PD;Karumanchi SA;Moore MJ;Aronin N;Khvorova A
通讯作者:
Khvorova A
影响因子:
4.1
作者:
Iribe H;Miyamoto K;Takahashi T;Kobayashi Y;Leo J;Aida M;Ui-Tei K
通讯作者:
Ui-Tei K
影响因子:
4.5
作者:
Friedman, JH
通讯作者:
Friedman, JH
影响因子:
14.9
作者:
Katoh T;Suzuki T
通讯作者:
Suzuki T