课题基金 / 基金详情

Discovering and Applying Knowledge in Clinical Databases

Discovering and Applying Knowledge in Clinical Databases
发现和应用临床数据库中的知识
批准号:
6630735
负责人:
GEORGE M HRIPCSAK
金额:
$37.76万
依托单位国家:
美国
项目类别:
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-06-01 至 2006-05-31

项目摘要

项目成果

GEORGE M HRIPCSAK的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供): 随着改进的临床信息系统产品(例如,门诊系统、订单录入系统)的出现,改进的数据录入技术(例如,语音识别、文本处理技术),以及进一步采用数据交换标准,更多的机构正在生成电子病历,并且这些记录将在未来的广度、深度和编码程度上扩展。这些记录主要用于个别患者的护理,但利用这些记录进行临床研究和质量功能的工作已经滞后。主要挑战包括范围广泛的复杂数据以及缺失和不准确的数据。 我们建议继续开发和测试挖掘临床数据存储库的方法。将特别强调利用储存库中的大量信息(潜在联系和知识),并使用计算机密集技术和数据表示和处理方面的进步,以更好地解释数据库中的内容,并克服复杂、缺失和不准确的数据带来的挑战。我们假设数据挖掘技术可以应用于存储库,以生成准确的临床解释。我们进一步假设,在临床丰富的储存库中潜在的关联可以用于改进该储存库中的病例分类。 我们的目标是开发方法来准备挖掘数据;表征临床数据库中的信息;基于自然语言处理器输出的操作和信息检索技术开发相似性度量;应用最近邻技术和基于案例的推理来改进分类;开发基于统计学的方法来改进数据不完整或不准确的病例分类;并将我们的方法应用于实际的临床研究问题,并开展更多的数据挖掘研究。 哥伦比亚大学医学信息学系的研究人员处于开展这项研究的独特地位,考虑到团队的经验(数据挖掘、统计、健康数据组织、健康知识表示、自然语言处理),200万患者13年来的数据存储库的可用性,以及名为MedLEE的自然语言处理器的可用性,可以将数百万份叙述性报告转换为高度编码的临床数据。
英文摘要
DESCRIPTION (provided by applicant): With the advent of improved clinical information system products (e.g., ambulatory systems, order entry systems), improved data entry technologies (e.g., speech recognition, text processing techniques), and further adoption of data interchange standards, more institutions are generating electronic medical records, and these records will expand in breadth, depth, and degree of coding in the future. The records are used mainly for individual patient care, but exploiting the records for clinical research and quality functions has lagged behind. Major challenges include the wide range of complex data and missing and inaccurate data. We propose to continue our work to develop and test methods to mine a clinical data repository. A special emphasis will be to exploit the vast amount of information in the repository (latent associations and knowledge) and to use computer intensive techniques and advances in data representation and manipulation to better interpret what is in the database and to overcome the challenges of complex, missing, and inaccurate data. We hypothesize that data mining techniques can be applied to a repository to generate accurate clinical interpretations. We further hypothesize that associations latent in a clinically rich repository can be used to improve the classification of cases in that repository. We aim to develop methods to prepare data for mining; to characterize the information in the clinical data repository; to develop similarity measures based on manipulation of natural language processor output and on information retrieval techniques; to apply nearest neighbor technique and case-based reasoning to improve classification; to develop a statistically based method to improve classification of cases with incomplete or inaccurate data; and to apply our methods to real clinical research questions and carry out additional data mining research. The researchers in the Department of Medical Informatics at Columbia University are uniquely positioned to carry out this research, given the experience of the team (data mining, statistics, health data organization, health knowledge representation, natural language processing), the availability of a repository of 13 years of data on 2 million patients, and the availability of a natural language processor called MedLEE to convert millions of narrative reports into richly coded clinical data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Annual OHDSI Symposium
Annual OHDSI Symposium
Annual OHDSI Symposium
2019 OHDSI Symposium
海外基金