iLearn: an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence data

iLearn: an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence data
复制标题

iLearn:一个集成平台和元学习器,用于 DNA、RNA 和蛋白质序列数据的特征工程、机器学习分析和建模

DOI:
10.1093/bib/bbz041
复制
发表时间:
2020-05-01
影响因子:
9.5
通讯作者:
Song, Jiangning
Song, Jiangning
中科院分区:
生物学2区
文献类型:
--
作者:
Chen, Zhen;Zhao, Pei;Song, Jiangning

文献摘要

被引文献

相似文献

随着后基因组时代产生的生物序列的爆炸性增长,生物信息学和计算生物学中最具挑战性的问题之一是以高效、准确和高通量的方式计算表征序列、结构和功能。迄今为止,已经开发了一些在线网络服务器和独立工具来解决这个问题;然而,所有这些工具在有效性、用户友好性和能力方面都有其局限性和缺点。在这里,我们提出了iLearn,一个全面的和多功能的基于Python的工具包,集成了DNA,RNA和蛋白质序列的特征提取,聚类,归一化,选择,降维,预测器构建,最佳描述符/模型选择,集成学习和结果可视化的功能。iLearn专为那些只想上传数据集并从中选择所需计算功能的用户而设计,而所有必要的程序和最佳设置都由软件自动完成。iLearn包含多种DNA、RNA和蛋白质的描述符,并支持四种特征输出格式,以方便直接输出使用或与其他计算工具通信。iLearn总共包含16种不同类型的特征聚类、选择、归一化和降维算法,以及5种常用的机器学习算法,从而极大地促进了特征分析和预测器构建。iLearn通过在线网络服务器和独立工具包免费提供。
With the explosive growth of biological sequences generated in the post-genomic era, one of the most challenging problems in bioinformatics and computational biology is to computationally characterize sequences, structures and functions in an efficient, accurate and high-throughput manner. A number of online web servers and stand-alone tools have been developed to address this to date; however, all these tools have their limitations and drawbacks in terms of their effectiveness, user-friendliness and capacity. Here, we present iLearn, a comprehensive and versatile Python-based toolkit, integrating the functionality of feature extraction, clustering, normalization, selection, dimensionality reduction, predictor construction, best descriptor/model selection, ensemble learning and results visualization for DNA, RNA and protein sequences. iLearn was designed for users that only want to upload their data set and select the functions they need calculated from it, while all necessary procedures and optimal settings are completed automatically by the software. iLearn includes a variety of descriptors for DNA, RNA and proteins, and four feature output formats are supported so as to facilitate direct output usage or communication with other computational tools. In total, iLearn encompasses 16 different types of feature clustering, selection, normalization and dimensionality reduction algorithms, and five commonly used machine-learning algorithms, thereby greatly facilitating feature analysis and predictor construction. iLearn is made freely available via an online web server and a stand-alone toolkit.