Constrained mixture estimation for analysis and robust classification of clinical time series.

Constrained mixture estimation for analysis and robust classification of clinical time series.
复制标题

DOI:
10.1093/bioinformatics/btp222
复制
发表时间:
2009-06-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Schliep A
Schliep A
中科院分区:
其他
文献类型:
--
作者:
Costa IG;Schönhuth A;Hafemeister C;Schliep A

文献摘要

参考文献

被引文献

相似文献

动机:基于疾病分子方面(例如基因表达谱)的个性化医疗已变得越来越流行。然而,在分析临床基因表达数据时面临多重挑战:大多数众所周知的理论问题(例如特征空间的高维与示例较少、噪声和缺失数据)都适用。在设计支持个性化诊断和治疗选择的分类程序时需要特别小心。在这里,我们特别关注多发性硬化症(MS)患者的干扰素-β(IFNβ)治疗反应的分类,这在最近引起了广泛关注。一半的患者不受 IFNβ 治疗的影响,这仍然是标准治疗。对于他们应该及时停止治疗,以减轻副作用。结果:我们建议对隐马尔可夫模型的混合物进行约束估计,作为对患者对 IFNβ 治疗的反应进行分类的方法。我们的方法的优点是它考虑了数据的时间性质及其对噪声、丢失数据和错误标记样本的鲁棒性。此外,混合估计能够在转录水平上探索患者反应亚组的存在。我们在预测准确性方面明显优于所有先前的方法,首次将其提高了 >90%。此外,我们能够识别潜在的错误标记样本,并将良好的反应者细分为两个表现出不同转录反应程序的亚组。最近关于多发性硬化症病理学的发现支持了这一点,因此可能会提出有趣的临床随访问题。可用性:该方法在 GQL 框架中实现,可从 http://www.ghmm.org/gql 获取。数据集可在 http://www.cin.ufpe.br/∼igcf/MSConst 联系方式:igcf@cin.ufpe.br 补充信息:补充数据可在生物信息学在线获取。
Motivation: Personalized medicine based on molecular aspects of diseases, such as gene expression profiling, has become increasingly popular. However, one faces multiple challenges when analyzing clinical gene expression data; most of the well-known theoretical issues such as high dimension of feature spaces versus few examples, noise and missing data apply. Special care is needed when designing classification procedures that support personalized diagnosis and choice of treatment. Here, we particularly focus on classification of interferon-β (IFNβ) treatment response in Multiple Sclerosis (MS) patients which has attracted substantial attention in the recent past. Half of the patients remain unaffected by IFNβ treatment, which is still the standard. For them the treatment should be timely ceased to mitigate the side effects. Results: We propose constrained estimation of mixtures of hidden Markov models as a methodology to classify patient response to IFNβ treatment. The advantages of our approach are that it takes the temporal nature of the data into account and its robustness with respect to noise, missing data and mislabeled samples. Moreover, mixture estimation enables to explore the presence of response sub-groups of patients on the transcriptional level. We clearly outperformed all prior approaches in terms of prediction accuracy, raising it, for the first time, >90%. Additionally, we were able to identify potentially mislabeled samples and to sub-divide the good responders into two sub-groups that exhibited different transcriptional response programs. This is supported by recent findings on MS pathology and therefore may raise interesting clinical follow-up questions. Availability: The method is implemented in the GQL framework and is available at http://www.ghmm.org/gql. Datasets are available at http://www.cin.ufpe.br/∼igcf/MSConst Contact: igcf@cin.ufpe.br Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1093/bioinformatics/btn152
发表时间: 2008-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lin TH;Kaminski N;Bar-Joseph Z
通讯作者: Bar-Joseph Z
DOI: 10.1186/1471-2105-8-s10-s3
发表时间: 2007-01-01
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Costa, Ivan G.;Krause, Roland;Schliep, Alexander
通讯作者: Schliep, Alexander
DOI: 10.1109/tcbb.2005.31
发表时间: 2005-07-01
影响因子: 4.5
作者:
Schliep, A;Costa, IG;Schönhuth, A
通讯作者: Schönhuth, A
DOI: 10.1093/bioinformatics/bth937
发表时间: 2004-08-04
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Schliep, Alexander;Steinhoff, Christine;Schoenhuth, Alexander
通讯作者: Schoenhuth, Alexander
DOI: 10.1093/nar/gkm226
发表时间: 2007-07
影响因子: 14.9
作者:
Reimand J;Kull M;Peterson H;Hansen J;Vilo J
通讯作者: Vilo J