A Bayesian network approach to feature selection in mass spectrometry data.

A Bayesian network approach to feature selection in mass spectrometry data.
复制标题

贝叶斯网络方法以质谱数据的特征选择。

DOI:
10.1186/1471-2105-11-177
复制
发表时间:
2010-04-08
期刊:
影响因子:
3
通讯作者:
Tracy ER
Tracy ER
中科院分区:
生物学4区
文献类型:
--
作者:
Kuschner KW;Malyarenko DI;Cooke WE;Cazares LH;Semmes OJ;Tracy ER

文献摘要

参考文献

被引文献

相似文献

飞行时间质谱仪(TOF-MS)通过检测血液或其他可获得的生物样本中的蛋白质生物标志物,有可能为癌症和其他严重疾病提供非侵入性的高通量筛查。不幸的是,由于测量的高度可变性、特定人群中蛋白质分布的不确定性以及使用现有统计工具提取可重复的诊断标记的困难,这一潜力迄今在很大程度上尚未实现。由于研究可能只由几十个样本和数百个变量组成,过度拟合是一个严重的复杂问题。为了克服这些困难,我们开发了一种贝叶斯归纳方法,它使用与模型无关的方法来发现光谱特征之间的关系。这种方法似乎有效地发现了网络模型,这些模型不仅确定了疾病和关键特征之间的联系,而且组织了特征之间的关系--而且还创建了一个稳定的分类器,以预测的错误率对新数据进行分类。将该方法应用于已知特征关系和典型TOF-MS变异性引入的人工数据,能够近乎完美地恢复这些关系。它也被应用于2004年白血病研究的血清数据,并在交叉验证下显示出所选特征的高度稳定性。使用隐瞒的数据对结果进行验证,显示出极好的预测能力。该方法显示了对传统技术的改进,并自然地纳入了测量不确定度。发现的特征之间的关系允许初步识别蛋白质生物标记物,这与其他癌症研究一致,后来得到了实验验证。这种方法似乎避免了生物数据的过度拟合,并在网络模型中产生了稳定的特征集。网络结构提供了有关特征之间关系的附加信息,这对指导进一步的生化分析是有用的。此外,当用于对新数据进行分类时,这些特征集比许多传统技术产生的特征集更加一致。
Time-of-flight mass spectrometry (TOF-MS) has the potential to provide non-invasive, high-throughput screening for cancers and other serious diseases via detection of protein biomarkers in blood or other accessible biologic samples. Unfortunately, this potential has largely been unrealized to date due to the high variability of measurements, uncertainties in the distribution of proteins in a given population, and the difficulty of extracting repeatable diagnostic markers using current statistical tools. With studies consisting of perhaps only dozens of samples, and possibly hundreds of variables, overfitting is a serious complication. To overcome these difficulties, we have developed a Bayesian inductive method which uses model-independent methods of discovering relationships between spectral features. This method appears to efficiently discover network models which not only identify connections between the disease and key features, but also organizes relationships between features--and furthermore creates a stable classifier that categorizes new data at predicted error rates. The method was applied to artificial data with known feature relationships and typical TOF-MS variability introduced, and was able to recover those relationships nearly perfectly. It was also applied to blood sera data from a 2004 leukemia study, and showed high stability of selected features under cross-validation. Verification of results using withheld data showed excellent predictive power. The method showed improvement over traditional techniques, and naturally incorporated measurement uncertainties. The relationships discovered between features allowed preliminary identification of a protein biomarker which was consistent with other cancer studies and later verified experimentally. This method appears to avoid overfitting in biologic data and produce stable feature sets in a network model. The network structure provides additional information about the relationships among features that is useful to guide further biochemical analysis. In addition, when used to classify new data, these feature sets are far more consistent than those produced by many traditional techniques.
DOI: 10.1038/sj.leu.2403781
发表时间: 2005-07-01
期刊: LEUKEMIA
影响因子: 11.4
作者:
Semmes, OJ;Cazares, LH;Jacobson, S
通讯作者: Jacobson, S
DOI: 10.1016/s0140-6736(05)17866-0
发表时间: 2005-02-05
期刊: LANCET
影响因子: 168.9
作者:
Michiels, S;Koscielny, S;Hill, C
通讯作者: Hill, C
DOI: 10.1371/journal.pcbi.0030129
发表时间: 2007-08
影响因子: 4.3
作者:
Needham CJ;Bradford JR;Bulpitt AJ;Westhead DR
通讯作者: Westhead DR
DOI: 10.1002/pmic.200701146
发表时间: 2008-04-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Tracy, Maureen B.;Chen, Haijian;Cooke, William E.
通讯作者: Cooke, William E.
DOI: 10.1373/clinchem.2004.037283
发表时间: 2005-01-01
期刊: CLINICAL CHEMISTRY
影响因子: 9.3
作者:
Malyarenko, DI;Cooke, WE;Manos, DM
通讯作者: Manos, DM