A multi-view genomic data simulator.

A multi-view genomic data simulator.
复制标题

DOI:
10.1186/s12859-015-0577-1
复制
发表时间:
2015-05-12
期刊:
影响因子:
3
通讯作者:
Greco D
Greco D
中科院分区:
生物学4区
文献类型:
--
作者:
Fratello M;Serra A;Fortino V;Raiconi G;Tagliaferri R;Greco D

文献摘要

参考文献

被引文献

相似文献

OMIC技术允许分析大量不同特征的状态(例如,mRNA表达、miRNA表达、拷贝数变异、DNA甲基化等)相同的样本。这些实验的目的通常是找到一组减少的显著特征,其可用于区分测定的条件。在开发新的特征选择计算方法方面,由于缺乏用于基准测试的完全注释的生物数据集,这项任务具有挑战性。解决这个问题的一种可能方法是生成适当的合成数据集,其组成和行为是完全受控的,并且是先验已知的。在这里,我们提出了一种新的方法集中在不同的生物分子之间的相互作用网络的生成,特别是参与调控基因表达。合成数据集从具有已知参数的基于常微分方程的模型获得。我们的研究结果表明,生成的数据集很好地模仿了真实的数据的行为,流行的数据分析方法能够有选择地识别现有的相互作用。所提出的方法可用于结合真实的生物数据集的数据挖掘技术的评估。该方法的主要优点在于完全控制模拟数据,同时保持与真实的生物过程的一致性。R软件包MVBioDataSim可在http://neuronelab.unisa.it/?上免费提供给科学界p=1722。本文的在线版本(doi:10.1186/s12859-015-0577-1)包含补充材料,可供授权用户使用。
OMICs technologies allow to assay the state of a large number of different features (e.g., mRNA expression, miRNA expression, copy number variation, DNA methylation, etc.) from the same samples. The objective of these experiments is usually to find a reduced set of significant features, which can be used to differentiate the conditions assayed. In terms of development of novel feature selection computational methods, this task is challenging for the lack of fully annotated biological datasets to be used for benchmarking. A possible way to tackle this problem is generating appropriate synthetic datasets, whose composition and behaviour are fully controlled and known a priori. Here we propose a novel method centred on the generation of networks of interactions among different biological molecules, especially involved in regulating gene expression. Synthetic datasets are obtained from ordinary differential equations based models with known parameters. Our results show that the generated datasets are well mimicking the behaviour of real data, for popular data analysis methods are able to selectively identify existing interactions. The proposed method can be used in conjunction to real biological datasets in the assessment of data mining techniques. The main strength of this method consists in the full control on the simulated data while retaining coherence with the real biological processes. The R package MVBioDataSim is freely available to the scientific community at http://neuronelab.unisa.it/?p=1722. The online version of this article (doi:10.1186/s12859-015-0577-1) contains supplementary material, which is available to authorized users.
DOI: 10.1371/journal.pcbi.1002488
发表时间: 2012
影响因子: 4.3
作者:
Sun J;Gong X;Purow B;Zhao Z
通讯作者: Zhao Z
DOI: 10.1126/science.298.5594.824
发表时间: 2002-10-25
期刊: SCIENCE
影响因子: 56.9
作者:
Milo, R;Shen-Orr, S;Alon, U
通讯作者: Alon, U
DOI: 10.1371/journal.pone.0064832
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Glass K;Huttenhower C;Quackenbush J;Yuan GC
通讯作者: Yuan GC
DOI: 10.1126/science.1073374
发表时间: 2002-08-30
期刊: SCIENCE
影响因子: 56.9
作者:
Ravasz, E;Somera, AL;Barabási, AL
通讯作者: Barabási, AL
DOI: 10.1111/j.1749-6632.2008.03756.x
发表时间: 2009-01-01
期刊: CHALLENGES OF SYSTEMS BIOLOGY: COMMUNITY EFFORTS TO HARNESS BIOLOGICAL COMPLEXITY
影响因子: --
作者:
Di Camillo, Barbara;Toffolo, Gianna;Cobelli, Claudio
通讯作者: Cobelli, Claudio