The big data revolution and human genetics.
The big data revolution and human genetics.
复制标题
大数据革命和人类遗传学。
DOI:
10.1093/hmg/ddy123
复制
发表时间:
2018
影响因子:
3.5
通讯作者:
Schork,NicholasJ
中科院分区:
文献类型:
--
作者:
Schork,NicholasJ
The concept of ‘Big Data’is now ubiquitous. Virtually all major industries, whether associated with finance, banking, marketing, retail, social media, energy or manufacturing, have embraced the analysis of big data sets hoping to obtain insights that could improve efficiency and create better products. It is no surprise then that the health care, biomedical research and, in particular, human genetics communities have also embraced big data initiatives. However, creating and analysing large-scale data sets is not particularly new in human genetics research contexts, since a single human genome contains 2Â $3.2 billion nucleotides worth of information. Just how these nucleotides are organized as genes and their associated regulatory elements, lead to particular functions, and interact, is as complicated and fascinating a big data analysis exercise as any in science. What is becoming more frequent, however, is the coupling or integration of human genetic data with other data types, essentially adding to the already very large data sets human genetic researchers are willing and eager to analyse. This issue of Human Molecular Genetics is devoted to reviews of efforts to both mine the big data inherent in human genomes and integrate that data with other data types to advance genetically oriented biomedical science and health care. Given the fundamental manner in which elements in DNA impact human physiology and mediate pathogenic processes, there are an unlimited number of settings in which human genetic research could complement, and be combined with, data from other research areas, as these reviews make clear. Virtually all of the reviews consider combining genetic data with other data to identify associations and connections between naturally occurring genetic variants possessed by individuals and phenotypes of all sorts, most notably those that may have clinical and public health utility. Telenti and colleagues consider the development and application of bioinformatics and data analysis tools to interpret variation in the human genome and show that by combining thousands of human genomes for analysis, insights into the likely functional effects of genetic variants can be found. Scheuermann and colleagues discuss methodology for leveraging what amounts to the billions of bits of information in the human genome to catalogue, characterize and subdivide the potentially trillions of cells in the human body. Fan and colleagues consider combining genetic data with imaging data, in particular neuroimaging data, in order to derive insights into human brain morphology that genetic and imaging data alone would not allow. Combining genetic data on patients in health systems with clinical information routinely collected on them is very logical given that a goal of contemporary human genetics research is to incorporate genetic information into clinical care. Unfortunately, combining genetic data with routine clinical data is fraught with difficulties given that clinical data are often ‘noisy’(eg physician hand-written notes, different ways of measuring clinical parameters entered into patient’s record, etc.). Altman and colleagues describe efforts to identify clinically meaningful relationships between genetic variants and responses to drugs used to treat various conditions; whereas Wolford and colleagues, Diao and colleagues, Ohno-Machado and colleagues and Glicksberg and colleagues all consider general integrated analysis of clinical data and genetic data with an eye towards improving health based on the insights obtained from such analyses. Pushing things even further, Huentelman and Talboom consider integrating genetic information with data …