The TOKEn project: knowledge synthesis for in silico science.

The TOKEn project: knowledge synthesis for in silico science.
复制标题

TOKEn 项目:计算机科学的知识综合。

DOI:
10.1136/amiajnl-2011-000434
复制
发表时间:
2011
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Greaves,AndrewW
Greaves,AndrewW
中科院分区:
--
文献类型:
--
作者:
Payne,PhilipRO;Borlawsky,TaraB;Lele,Omkar;James,Stephen;Greaves,AndrewW

文献摘要

相似文献

目的:涉及大规模数据集的调查研究的开展对能够支持硅发现科学的新假设的发现和测试提出了重大挑战。数据库中概念知识发现(CKDD)方法的使用为大型数据集提供了一种扩展假设发现和测试方法的潜在手段。这些方法能够高通量地生成和评估目标数据集中发现的变量复合体之间的知识锚定关系。方法:作者进行了多部分模型制定和验证过程,重点是开发一种方法和技术方法,使用CKDD来支持计算机科学的假设发现。作者开发的模型被称为翻译本体锚定知识发现引擎(TOKEn)。该模型利用一种称为建设性归纳的特定CKDD方法来识别和优先考虑与大规模和异构生物医学数据集中发现的变量之间有意义的语义关系相关的潜在假设。结果作者在nci资助的慢性淋巴细胞白血病研究联盟维护的转化研究数据库中验证了TOKEn。这些研究表明,TOKEn具有:(1)计算可处理性;(2)能够在数据收集中产生关于表型和生物分子变量之间关系的有效和潜在有用的假设。TOKEn模型代表了在大规模和多维研究数据集的背景下,对硅发现科学的知识合成有潜在的有用和系统的方法。
ObjectiveThe conduct of investigational studies that involve large-scale data sets presents significant challenges related to the discovery and testing of novel hypotheses capable of supporting in silico discovery science. The use of what are known as Conceptual Knowledge Discovery in Databases (CKDD) methods provides a potential means of scaling hypothesis discovery and testing approaches for large data sets. Such methods enable the high-throughput generation and evaluation of knowledge-anchored relationships between complexes of variables found in targeted data sets.MethodsThe authors have conducted a multipart model formulation and validation process, focusing on the development of a methodological and technical approach to using CKDD to support hypothesis discovery for in silico science. The model the authors have developed is known as the Translational Ontology-anchored Knowledge Discovery Engine (TOKEn). This model utilizes a specific CKDD approach known as Constructive Induction to identify and prioritize potential hypotheses related to the meaningful semantic relationships between variables found in large-scale and heterogeneous biomedical data sets.ResultsThe authors have verified and validated TOKEn in the context of a translational research data repository maintained by the NCI-funded Chronic Lymphocytic Leukemia Research Consortium. Such studies have shown that TOKEn is: (1) computationally tractable; and (2) able to generate valid and potentially useful hypotheses concerning relationships between phenotypic and biomolecular variables in that data collection.ConclusionsThe TOKEn model represents a potentially useful and systematic approach to knowledge synthesis for in silico discovery science in the context of large-scale and multidimensional research data sets.