Gentle Introduction to the Statistical Foundations of False Discovery Rate in Quantitative Proteomics

Gentle Introduction to the Statistical Foundations of False Discovery Rate in Quantitative Proteomics
复制标题

DOI:
10.1021/acs.jproteome.7b00170
复制
发表时间:
2018-01-01
影响因子:
4.4
通讯作者:
Burger, Thomas
Burger, Thomas
中科院分区:
生物学2区
文献类型:
--
作者:
Burger, Thomas

文献摘要

被引文献

相似文献

从计算蛋白质组学研究的角度来看,理论统计学的词汇可能很难接受,尽管它所传达的概念对出版指南至关重要。例如,“调整的p值”、“q值”和“错误发现率”本质上是相似的概念,而“错误发现率”和“错误发现比例”不能混淆,即使“率”和“比例”在日常语言中是相关的。在蛋白质组学的跨学科背景下,这种微妙之处可能会引起误解。这篇文章的目的是提供一个易于理解的解释这四个概念(和其他一些相关的)。它们的统计基础是从很大程度上依赖直觉的角度来处理的,主要解决蛋白质定量问题,但在一定程度上也解决肽识别问题。此外,明确区分了定义单个属性的概念(即,与肽或蛋白质相关)和定义集合性质的那些(即,与肽或蛋白质的列表相关)。
The vocabulary of theoretical statistics can be difficult to embrace from the viewpoint of computational proteomics research, even though the notions it conveys are essential to publication guidelines. For example, "adjusted p-values", "q-values", and "false discovery rates" are essentially similar concepts, whereas "false discovery rate" and "false discovery proportion" must not be confused, even though "rate" and "proportion" are related in everyday language. In the interdisciplinary context of proteomics, such subtleties may cause misunderstandings. This article aims to provide an easy-to-understand explanation of these four notions (and a few other related ones). Their statistical foundations are dealt with from a perspective that largely relies on intuition, addressing mainly protein quantification but also, to some extent, peptide identification. In addition, a clear distinction is made between concepts that define an individual property (i.e., related to a peptide or a protein) and those that define a set property (i.e., related to a list of peptides or proteins).