Maximizing the reusability of gene expression data by predicting missing metadata.

Maximizing the reusability of gene expression data by predicting missing metadata.
复制标题

DOI:
10.1371/journal.pcbi.1007450
复制
发表时间:
2020-11
影响因子:
4.3
通讯作者:
Zhang J
Zhang J
中科院分区:
生物学2区
文献类型:
--
作者:
Lung PY;Zhong D;Pang X;Li Y;Zhang J

文献摘要

参考文献

被引文献

相似文献

可重用性是公平数据原则的一部分,旨在使数据可查找、可访问、可互操作和可重用。当前提高公共基因组学数据可重用性的努力之一是关注包含与数据相关的质量元数据。当必要的元数据丢失时,大多数研究人员会认为这些数据毫无用处。在这项研究中,我们开发了一个框架来预测基因表达数据集缺失的元数据,以最大限度地提高其可重用性。我们发现,当使用预测数据进行其他分析时,使用所有预测数据并不是最佳选择。相反,我们应该只使用可以准确预测的数据子集。我们提出了一个名为“准确预测的案例比例”(PCAP) 的新指标,该指标在我们专门设计的机器学习流程中进行了优化。在最大化缺失值数据的可重用性方面,新方法的性能优于使用常用指标(例如 F1 分数)的管道。我们还发现,可能需要使用不同的机器学习方法和/或不同的数据处理协议来预测不同的变量。以差异基因表达分析为例,我们表明,当准确预测缺失变量时,相应的基因表达数据可以可靠地用于下游分析。基因表达综合数据库 (GEO) 和序列读取存档 (SRA) 等公共数据库提供了大量基因表达数据。可以重新分析它们以解决以前不可行的生物学问题。然而,由于缺乏分析所需的元数据,使用公共基因组学数据的再分析研究受到了阻碍。这可以通过使用基因表达数据预测元数据来解决,然后可以将其用于具有预测元数据的所需重新分析中。这代表了一种提高公共基因表达数据可重用性的新方法。我们的研究试图系统地研究如何实施这种方法。我们发现,不应使用所有带有预测元数据的基因表达数据来进行下游分析。虽然使用所有基因表达数据可以最大化样本量,但预测不佳的表达谱可能会影响下游分析的质量。人们需要在下游分析中包含的数据量和预测元数据的准确性之间取得平衡。为了解决这个问题,我们设计了一个名为“准确预测的案例比例”(PCAP) 的新指标,该指标在我们专门设计的机器学习流程中进行了优化。以差异基因表达分析为例,我们表明,当准确预测缺失变量时,相应的基因表达数据可以可靠地用于下游分析。
Reusability is part of the FAIR data principle, which aims to make data Findable, Accessible, Interoperable, and Reusable. One of the current efforts to increase the reusability of public genomics data has been to focus on the inclusion of quality metadata associated with the data. When necessary metadata are missing, most researchers will consider the data useless. In this study, we developed a framework to predict the missing metadata of gene expression datasets to maximize their reusability. We found that when using predicted data to conduct other analyses, it is not optimal to use all the predicted data. Instead, one should only use the subset of data, which can be predicted accurately. We proposed a new metric called Proportion of Cases Accurately Predicted (PCAP), which is optimized in our specifically-designed machine learning pipeline. The new approach performed better than pipelines using commonly used metrics such as F1-score in terms of maximizing the reusability of data with missing values. We also found that different variables might need to be predicted using different machine learning methods and/or different data processing protocols. Using differential gene expression analysis as an example, we showed that when missing variables are accurately predicted, the corresponding gene expression data can be reliably used in downstream analyses. Large volumes of gene expression data are available at public databases such as Gene Expression Omnibus (GEO) and sequence read archive (SRA). They can be reanalyzed to solve previously infeasible biological problems. However, reanalysis studies using public genomics data have been hindered by the lack of necessary metadata for the analyses. This can be addressed by predicting the metadata using the gene expression data, which can then be used in the desired reanalysis with predicted metadata. This represents a new approach to increase the reusability of public gene expression data. Our study attempts to systematically investigate how this approach should be carried out. We found that one should not use all the gene expression data with metadata predicted for downstream analyses. While using all the gene expression data maximizes the sample size, the poorly predicted expression profiles may affect the quality of the downstream analysis. One needs to strike a balance between the amount of data included in the downstream analysis and the accuracy of predicted metadata. To address this problem, we designed a new metric called Proportion of Cases Accurately Predicted (PCAP), which is optimized in our specifically-designed machine learning pipeline. Using differential gene expression analysis as an example, we showed that when missing variables are accurately predicted, the corresponding gene expression data can be reliably used in downstream analyses.
DOI: 10.12688/f1000research.12223.1
发表时间: 2017-01-01
期刊: F1000Research
影响因子: --
作者:
Collado-Torres, Leonardo;Nellore, Abhinav;Jaffe, Andrew E
通讯作者: Jaffe, Andrew E
遗传对人体组织基因表达的影响。
DOI: 10.1038/nature24277
发表时间: 2017-10-11
期刊: Nature
影响因子: 64.8
作者:
GTEx Consortium;Laboratory, Data Analysis &Coordinating Center (LDACC)—Analysis Working Group;Statistical Methods groups—Analysis Working Group;Enhancing GTEx (eGTEx) groups;NIH Common Fund;NIH/NCI;NIH/NHGRI;NIH/NIMH;NIH/NIDA;Biospecimen Collection Source Site—NDRI;Biospecimen Collection Source Site—RPCI;Biospecimen Core Resource—VARI;Brain Bank Repository—University of Miami Brain Endowment Bank;Leidos Biomedical—Project Management;ELSI Study;Genome Browser Data Integration &Visualization—EBI;Genome Browser Data Integration &Visualization—UCSC Genomics Institute, University of California Santa Cruz;Lead analysts:;Laboratory, Data Analysis &Coordinating Center (LDACC):;NIH program management:;Biospecimen collection:;Pathology:;eQTL manuscript working group:;Battle A;Brown CD;Engelhardt BE;Montgomery SB
通讯作者: Montgomery SB
使用 mRNA 和 miRNA 表达谱对 ER、PR 和 HER2 定义的乳腺癌亚组进行综合研究
DOI: 10.1038/srep06566
发表时间: 2014-10-23
期刊: Scientific reports
影响因子: 4.6
作者:
Dai X;Chen A;Bai Z
通讯作者: Bai Z
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1038/nm733
发表时间: 2002-08-01
期刊: NATURE MEDICINE
影响因子: 82.9
作者:
Beer, DG;Kardia, SLR;Hanash, S
通讯作者: Hanash, S