Using electronic patient records to discover disease correlations and stratify patient cohorts.

Using electronic patient records to discover disease correlations and stratify patient cohorts.
复制标题

DOI:
10.1371/journal.pcbi.1002141
复制
发表时间:
2011-08
影响因子:
4.3
通讯作者:
Brunak S
Brunak S
中科院分区:
生物学2区
文献类型:
--
作者:
Roque FS;Jensen PB;Schmock H;Dalgaard M;Andreatta M;Hansen T;Søeby K;Bredkjær S;Juul A;Werge T;Jensen LJ;Brunak S

文献摘要

参考文献

被引文献

相似文献

电子病历仍然是一个相当未开发的,但潜在的丰富的数据源,发现疾病之间的相关性。我们描述了一个一般的方法,收集患者的表型描述的医疗记录中的系统和非队列依赖的方式。通过从这些记录中的自由文本中提取表型信息,我们证明了我们可以扩展结构化记录数据中包含的信息,并将其用于产生细粒度的患者分层和疾病共现统计。该方法使用基于国际疾病分类本体的词典,因此原则上与语言无关。作为一个用例,我们展示了如何从丹麦精神病医院的记录导致识别疾病的相关性,随后可以映射到系统生物学框架。文本挖掘和信息提取可以被看作是将隐藏在文本中的信息转换为可管理数据的挑战。我们已经使用文本挖掘自动提取临床相关的条款,从5543精神病患者的记录,并将这些映射到疾病代码的国际疾病分类本体(ICD 10)。现有的编码数据补充了挖掘的编码。对于每个患者,我们构建了相关ICD 10编码的表型谱。这使我们能够根据患者特征的相似性将他们聚类在一起。结果是基于比通常使用的主要诊断更完整的特征的患者分层。同样,我们通过寻找比预期更频繁地在患者中同时出现的疾病代码对来调查合并症。我们的高排名配对是由一位医生手动策划的,他将93名候选人标记为有趣。对于其中的一些,我们能够使用OMIM数据库找到已知与疾病相关的基因/蛋白质。疾病相关蛋白质使我们能够构建蛋白质网络,怀疑参与每种表型。两种相关疾病之间共有的蛋白质可能为疾病共病提供见解。
Electronic patient records remain a rather unexplored, but potentially rich data source for discovering correlations between diseases. We describe a general approach for gathering phenotypic descriptions of patients from medical records in a systematic and non-cohort dependent manner. By extracting phenotype information from the free-text in such records we demonstrate that we can extend the information contained in the structured record data, and use it for producing fine-grained patient stratification and disease co-occurrence statistics. The approach uses a dictionary based on the International Classification of Disease ontology and is therefore in principle language independent. As a use case we show how records from a Danish psychiatric hospital lead to the identification of disease correlations, which subsequently can be mapped to systems biology frameworks. Text mining and information extraction can be seen as the challenge of converting information hidden in text into manageable data. We have used text mining to automatically extract clinically relevant terms from 5543 psychiatric patient records and map these to disease codes in the International Classification of Disease ontology (ICD10). Mined codes were supplemented by existing coded data. For each patient we constructed a phenotypic profile of associated ICD10 codes. This allowed us to cluster patients together based on the similarity of their profiles. The result is a patient stratification based on more complete profiles than the primary diagnosis, which is typically used. Similarly we investigated comorbidities by looking for pairs of disease codes cooccuring in patients more often than expected. Our high ranking pairs were manually curated by a medical doctor who flagged 93 candidates as interesting. For a number of these we were able to find genes/proteins known to be associated with the diseases using the OMIM database. The disease-associated proteins allowed us to construct protein networks suspected to be involved in each of the phenotypes. Shared proteins between two associated diseases might provide insight to the disease comorbidity.
DOI: 10.1197/jamia.m1727
发表时间: 2005-05-01
影响因子: 6.4
作者:
Galanter, WL;Didomenico, RJ;Polikaitis, A
通讯作者: Polikaitis, A
DOI: 10.1136/bmj.c3111
发表时间: 2010-06-16
影响因子: 105.7
作者:
Greenhalgh, Trisha;Stramer, Katja;Potts, Henry W. W.
通讯作者: Potts, Henry W. W.
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1371/journal.pcbi.1000353
发表时间: 2009-04
影响因子: 4.3
作者:
Hidalgo CA;Blumm N;Barabási AL;Christakis NA
通讯作者: Christakis NA
DOI: 10.1136/bmj.328.7437.438
发表时间: 2004-02-21
影响因子: --
作者:
Eaton, WW;Mortensen, PB;Ewald, H
通讯作者: Ewald, H