Maximizing the reusability of gene expression data by predicting missing metadata.
Maximizing the reusability of gene expression data by predicting missing metadata.
复制标题
DOI:
10.1371/journal.pcbi.1007450
复制
发表时间:
2020-11
影响因子:
4.3
通讯作者:
Zhang J
中科院分区:
文献类型:
--
作者:
Lung PY;Zhong D;Pang X;Li Y;Zhang J
Reusability is part of the FAIR data principle, which aims to make data Findable, Accessible, Interoperable, and Reusable. One of the current efforts to increase the reusability of public genomics data has been to focus on the inclusion of quality metadata associated with the data. When necessary metadata are missing, most researchers will consider the data useless. In this study, we developed a framework to predict the missing metadata of gene expression datasets to maximize their reusability. We found that when using predicted data to conduct other analyses, it is not optimal to use all the predicted data. Instead, one should only use the subset of data, which can be predicted accurately. We proposed a new metric called Proportion of Cases Accurately Predicted (PCAP), which is optimized in our specifically-designed machine learning pipeline. The new approach performed better than pipelines using commonly used metrics such as F1-score in terms of maximizing the reusability of data with missing values. We also found that different variables might need to be predicted using different machine learning methods and/or different data processing protocols. Using differential gene expression analysis as an example, we showed that when missing variables are accurately predicted, the corresponding gene expression data can be reliably used in downstream analyses. Large volumes of gene expression data are available at public databases such as Gene Expression Omnibus (GEO) and sequence read archive (SRA). They can be reanalyzed to solve previously infeasible biological problems. However, reanalysis studies using public genomics data have been hindered by the lack of necessary metadata for the analyses. This can be addressed by predicting the metadata using the gene expression data, which can then be used in the desired reanalysis with predicted metadata. This represents a new approach to increase the reusability of public gene expression data. Our study attempts to systematically investigate how this approach should be carried out. We found that one should not use all the gene expression data with metadata predicted for downstream analyses. While using all the gene expression data maximizes the sample size, the poorly predicted expression profiles may affect the quality of the downstream analysis. One needs to strike a balance between the amount of data included in the downstream analysis and the accuracy of predicted metadata. To address this problem, we designed a new metric called Proportion of Cases Accurately Predicted (PCAP), which is optimized in our specifically-designed machine learning pipeline. Using differential gene expression analysis as an example, we showed that when missing variables are accurately predicted, the corresponding gene expression data can be reliably used in downstream analyses.
登录
查看更多内容
影响因子:
--
作者:
Collado-Torres, Leonardo;Nellore, Abhinav;Jaffe, Andrew E
通讯作者:
Jaffe, Andrew E
影响因子:
64.8
作者:
GTEx Consortium;Laboratory, Data Analysis &Coordinating Center (LDACC)—Analysis Working Group;Statistical Methods groups—Analysis Working Group;Enhancing GTEx (eGTEx) groups;NIH Common Fund;NIH/NCI;NIH/NHGRI;NIH/NIMH;NIH/NIDA;Biospecimen Collection Source Site—NDRI;Biospecimen Collection Source Site—RPCI;Biospecimen Core Resource—VARI;Brain Bank Repository—University of Miami Brain Endowment Bank;Leidos Biomedical—Project Management;ELSI Study;Genome Browser Data Integration &Visualization—EBI;Genome Browser Data Integration &Visualization—UCSC Genomics Institute, University of California Santa Cruz;Lead analysts:;Laboratory, Data Analysis &Coordinating Center (LDACC):;NIH program management:;Biospecimen collection:;Pathology:;eQTL manuscript working group:;Battle A;Brown CD;Engelhardt BE;Montgomery SB
通讯作者:
Montgomery SB
影响因子:
4.6
作者:
Dai X;Chen A;Bai Z
通讯作者:
Bai Z
影响因子:
56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者:
Lander, ES
影响因子:
82.9
作者:
Beer, DG;Kardia, SLR;Hanash, S
通讯作者:
Hanash, S