课题基金 / 基金详情

DISCOVERING AND APPLYING KNOWLEDGE IN CLINICAL DATABASES

DISCOVERING AND APPLYING KNOWLEDGE IN CLINICAL DATABASES
发现和应用临床数据库中的知识
批准号:
6031325
负责人:
GEORGE M HRIPCSAK
金额:
$43.34万
依托单位国家:
美国
项目类别:
财政年份:
2000
资助国家:
美国
项目状态:
已结题
起止时间:
2000-04-01 至 2003-03-31

项目摘要

项目成果

GEORGE M HRIPCSAK的其他基金

相似基金

相关文献

中文摘要
翻译
实时临床存储库包含大量对临床护理、研究和管理有用的详细信息。然而,在它们的原始形式中,数据很难使用,有太多的量、太多的细节、缺少值和不准确。临床医生、研究人员和管理人员需要更高层次的解释来解决他们的问题。例如,临床医生可能需要知道患者是否有足够的风险患有活动性结核病,从而需要进行呼吸道隔离。这个问题的答案可能通过胸部X光片、实验室测试、用药史、生命体征和医生笔记传播到临床存储库。将这些原始数据转化为解释(无论是否存在风险)是一项艰巨而费力的任务。这一提议的假设是,数据挖掘技术可以应用于实时临床存储库,以发现知识并生成准确的临床解释,并且这些解释可以自动化。该项目与早期的机器学习研究的不同之处在于,它强调真实的临床存储库,并使用自然语言处理来提供编码的临床数据。具体目标是:(L)选择临床领域--将选择几个具有有趣的、非琐碎的临床问题的临床领域。将选择能够或已经为回溯性队列收集黄金标准答案的问题。(2)准备用于挖掘的原始临床数据--将来自临床存储库的原始数据转换为便于数据挖掘的结构。数据将根据域的需要进行扁平化、透视、汇总和映射。叙事数据将使用MedLEE自然语言处理器进行编码。准备过程将自动进行。(3)使用数据挖掘算法发现知识--几种数据挖掘算法将应用于选定的临床领域。算法将包括决策树生成、规则发现、神经网络、最近邻、逻辑回归和复合算法(用于变量减少)。这些算法将在每个领域的训练集上进行训练,它们的预测精度将被测量并相互比较,并与专家编写的规则进行比较。还将衡量人类专家使用手动数据挖掘可视化技术(不需要明确的训练集)编写规则的性能。(4)研究数据挖掘对训练集的依赖关系--数据挖掘算法的性能取决于训练它们所使用的数据。将测量算法对噪声(不准确数据)、缺失数据和训练集大小的敏感度。(5)使用发现的知识生成实时解释--算法的输出(决策树、规则、神经网络方程或逻辑回归方程,但不是最近邻)以及必要的数据准备步骤将被编码到Arden语法医学逻辑模块中。他们将根据临床资料库运行,以验证解释可以实时自动进行。(6)传播方法和结果--将通过出版物和网站传播方法和结果,并提供工具。
英文摘要
A real-time clinical repository contains a wealth of detailed information useful for clinical care, research, and administration. In their raw form, however, the data are difficult to use there is too much volume, too much detail, missing values, and inaccuracies. Clinicians, researchers, and administrators require higher level interpretations that address their questions. For example, a clinician may need to know whether a patient is at sufficient risk for having active tuberculosis to warrant respiratory isolation. The answer to the question may be spread around the clinical repository in chest radiographs, laboratory tests, medication histories, vital signs, and physician's notes. Translating from these raw data to the interpretation (at risk or not) is a difficult and laborious task. The hypothesis of this proposal is that data mining techniques can be applied to a real-time clinical repository to discover knowledge and generate accurate clinical interpretations, and that these interpretations can be automated. The project differs from earlier machine learning studies in its emphasis on a real clinical repository and the use of natural language processing to supply coded clinical data. The specific aims are: (l) Select clinical domains--Several clinical domains with interesting, non-trivial clinical problems will be selected. Problems for which a gold standard answer can or has been assembled for a retrospective cohort will be chosen. (2) Prepare raw clinical data for mining--The raw data from a clinical repository will be transformed into a structure that facilitates data mining. The data will be flattened, pivoted, summarized, and mapped as needed for the domains. Narrative data will be coded using the MedLEE natural language processor. The preparation process will be automated. (3) Use data mining algorithms to discover knowledge- Several data mining algorithms will be applied to the selected clinical domains. Algorithms will include decision tree generation, rule discovery, neural networks, nearest neighbor, logistic regression, and composite algorithms (for variable reduction). The algorithms will be trained on a training set for each domain, and their predictive accuracy will be measured and compared to each other and to expert-written rules. The performance of human experts writing rules using manual data mining visualization techniques (which does not require an explicit training set) will also be measured. (4) Study the dependence of data mining on the training set--The performance of data mining algorithms depends on the data used the train them. The sensitivity of the algorithms to noise (inaccurate data), missing data, and training set size will be measured. (5) Use the discovered knowledge to generate real-time interpretations-- The output of the algorithms (decision tree, rules, neural network equation, or logistic regression equation, but not nearest neighbor) along with the necessary data preparation steps will be encoded in Arden Syntax Medical Logic Modules. They will be run against the clinical repository to verify that the interpretation can be automated in real time. (6) Disseminate the methods and results--The methods and results will be disseminated via publications and a Web site, and tools will be made available.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Annual OHDSI Symposium
Annual OHDSI Symposium
Annual OHDSI Symposium
2019 OHDSI Symposium
海外基金