Characterizing multi-omic data in systems biology.

Characterizing multi-omic data in systems biology.
复制标题

表征系统生物学中的多组学数据。

DOI:
10.1007/978-1-4614-8778-4_2
复制
发表时间:
2014
影响因子:
--
通讯作者:
Smith,ToddM
Smith,ToddM
中科院分区:
医学4区
文献类型:
--
作者:
Mason,ChristopherE;Porter,SandraG;Smith,ToddM

文献摘要

被引文献

相似文献

在今天的生物学中,研究已经转移到分析离散的生化反应和途径上的系统。这些研究依赖于将分析DNA、mRNA、非编码RNA、DNA、RNA和蛋白质相互作用的数十种实验方法的结果,以及形成表观基因组的核苷酸修饰合并到代表不同的“组学”数据(转录、表观遗传学、蛋白质组、代谢组)的全局数据集。用于收集这些数据的方法包括高通量数据生成平台,包括高含量筛查、成像、流式细胞仪、质谱仪和核酸测序。其中,下一代DNA测序平台占据主导地位,因为它们提供了一种廉价和可扩展的方法来快速询问遗传、表观遗传和转录水平的分子变化。此外,现有的和正在开发的单分子测序平台可能会使直接测量RNA和蛋白质成为可能,从而增加当前分析的特异性,并使更好地表征发生在表观基因组和表观转录组中的“表观变化”成为可能。这些多样化的数据类型给我们带来了最大的挑战:我们如何开发能够整合这些数据集的软件系统和算法,并开始支持一个更民主的模型,在这种模型中,个人可以通过生物识别设备和个人基因组测序来捕获和跟踪自己的医疗信息?这样的系统将需要提供必要的用户交互,以处理做出科学发现所需的数万亿数据点。在这里,我们描述了这些数据的起源和处理的新方法,整合这些数据的模型,以及自我报告和自我测量的基因组学和健康数据日益普遍。
In today’s biology, studies have shifted to analyzing systems over discrete biochemical reactions and pathways. These studies depend on combining the results from scores of experimental methods that analyze DNA; mRNA; noncoding RNAs, DNA, RNA, and protein interactions; and the nucleotide modifications that form the epigenome into global datasets that represent a diverse array of “omics” data (transcriptional, epigenetic, proteomic, metabolomic). The methods used to collect these data consist of high-throughput data generation platforms that include high-content screening, imaging, flow cytometry, mass spectrometry, and nucleic acid sequencing. Of these, the next-generation DNA sequencing platforms predominate because they provide an inexpensive and scalable way to quickly interrogate the molecular changes at the genetic, epigenetic, and transcriptional level. Furthermore, existing and developing single-molecule sequencing platforms will likely make direct RNA and protein measurements possible, thus increasing the specificity of current assays and making it possible to better characterize “epi-alterations” that occur in the epigenome and epitranscriptome. These diverse data types present us with the largest challenge: how do we develop software systems and algorithms that can integrate these datasets and begin to support a more democratic model where individuals can capture and track their own medical information through biometric devices and personal genome sequencing? Such systems will need to provide the necessary user interactions to work with the trillions of data points needed to make scientific discoveries. Here, we describe novel approaches in the genesis and processing of such data, models to integrate these data, and the increasing ubiquity of self-reporting and self-measured genomics and health data.