课题基金 / 基金详情

Interoperation of Genome Databases and Tools

Interoperation of Genome Databases and Tools
基因组数据库和工具的互操作
批准号:
6946756
负责人:
KEI-HOI CHEUNG
金额:
$14.62万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-09-01 至 2006-08-31

项目摘要

项目成果

KEI-HOI CHEUNG的其他基金

相似基金

相关文献

中文摘要
翻译
描述:(由申请人提供)这份申请美国国立卫生研究院导师 量化研究事业奖征集对张启海博士的支持 他开始了专注于基因组相关生物信息学的教职生涯。这 申请提出了一项研究职业生涯发展计划,在该领域 生物信息学,连接计算机科学和生物学的桥梁。该计划包括两个 部分重叠阶段:(1)强调训练的教学阶段, 包括遗传学和基因组学领域的课程和实验室工作 以补充张博士在电脑科学方面的博士训练及(Ii)a 开发阶段,侧重于拟议研究的密集开发。 这两个阶段将由一个由高级官员组成的指导委员会密切监督 科学家,他们将担任导师或顾问,在生物学和 生物信息学。 人类基因组计划和基因组技术的快速发展(例如, 微阵列)产生了大量的本地、国家和国际基因组 数据库,其中许多数据库可以通过Web访问。回答出现的问题 在高级基因组研究项目中,研究人员往往需要分析大量 从多个相关数据库收集的数据量。因此, 重要的是探索(1)如何集成 灵活实用的方式和(2)如何执行大规模数据分析,如 尽可能容易和快速地完成。为此,我们提出了两个免费的 接近了。 1.数据集成或互操作困难,原因是 所涉及的句法和语义异质性。为了解决这个问题, 我们提出了一种使用可扩展标记语言(XML)的元数据驱动方法, 它结合了标准化的词汇来映射各种Web可访问的内容 将数据集转换为便于互操作性的通用格式。 2.为了方便和加快分析大量数据,我们将 我还将探索一系列计算技术,包括使用 涡轮基因组学,代表着高性能的合作 耶鲁大学计算机科学系内的计算机组。这些 技术允许(I)集成不同的软件组件(分析 工具)容易完成和(Ii)利用并行的力量 计算。 我们将在以下背景下设计、开发、测试和评估该方法 当前的数据库项目包括:1)管理数据的三元组 大规模酵母基因组分析(与Snyder教授合作)和2)Alfred存储 不同人群的基因频率数据(与基德教授合写)。我们有 确定了一些可通过Web访问的相关外部数据库以及 用户希望从Triple和Alfred访问的工具 时尚。我们将初步开发和应用我们的方法来集成这些 数据库和工具。我们将把我们的方法扩展到其他类型的基因组数据 例如微阵列数据,实验室和其他机构很快就会 大量地产生。
英文摘要
DESCRIPTION: (provided by applicant) This application for an NIH Mentored Quantitative Research Career Award requests support for Dr. Kei-Hoi Cheung as he embarks on a faculty career focused on genome-related bioinformatics. This application presents a research career development plan in the field of bioinformatics, bridging computer science and biology. The plan includes two partially overlapping phases: (1) a didactic phase that emphasizes training, including coursework and laboratory work in the area of genetics and genomics to complement Dr. Cheung's doctoral training in Computer Science and (ii) a development phase that focuses on intense development of the proposed research. These two phases will be closely supervised by a steering committee of senior scientists, who will serve as mentors or advisors, in the area of biology and bioinformatics. The human genome project and the rapid advance in genomic technology (e.g., microarrays) have produced numerous local, national, and international genome databases, many of which are Web-accessible. To answer questions that arise in advanced genome research projects, researchers often need to analyze a large amount of data that are collected from multiple related databases. Therefore, it is important to explore (1) how to integrate the databases involved in a flexible and useful fashion and (2) how to perform large-scale data analyses as easily and rapidly as possible. To this end, we propose two complimentary approaches. 1. The problem of data integration or interoperation is difficult because of the syntactic and semantic heterogeneities involved. To address this problem, we propose a metadata-driven approach using eXtensible Markup Language (XML), which incorporates standardized vocabulary to map heterogeneous Web-accessible data sets into a common format that facilitates interoperability. 2. To facilitate and speed up analysis of a large quantity of data, we will also explore a range of computational techniques including the use of Turbogenomics, which represents collaboration with the high performance computing group within the Yale department of Computer Science. These techniques allow (i) integration of heterogeneous software components (analysis tools) to be done easily and (ii) exploitation of the power of parallel computing. We will design, develop, test, and evaluate the approach in the context of current database projects including: 1) TRIPLES that manages data for large-scale yeast genome analysis (with Prof Snyder) and 2) ALFRED that stores gene frequency data on different human populations (with Prof Kidd). We have identified a number of related external Web-accessible databases as well as tools that users would like to access from TRIPLES and ALFRED in an integrated fashion. We will initially develop and apply our approach to integrate these databases and tools. We will extend our approach to other types of genomic data such as microarray data, which both laboratories and others will soon be generating in large quantities.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1107/s160053680800425x
发表时间: 2008-02-15
期刊: Acta crystallographica. Section E, Structure reports online
影响因子: --
作者: [Macmillan SN, Tanski JM, Waterman R]
通讯作者: Waterman R
DOI: 10.1186/1471-2105-5-25
发表时间: 2004-03-10
期刊: BMC bioinformatics
影响因子: 3
作者: [de Knikker R, Guo Y, Li JL, Kwan AK, Yip KY, Cheung DW, Cheung KH]
通讯作者: Cheung KH
DOI: 10.1142/9789812701626_0018
发表时间: 2005-12
期刊: Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
影响因子: --
作者: [Kevin Y. Yip;Peishen Qi;Martin H. Schultz;D. Cheung;K. Cheung]
通讯作者: Kevin Y. Yip;Peishen Qi;Martin H. Schultz;D. Cheung;K. Cheung
Identifying projected clusters from gene expression profiles.
从基因表达谱中识别预测的簇。
DOI: 10.1016/j.jbi.2004.05.002
发表时间: 2004
期刊: Journal of biomedical informatics.
影响因子: --
作者: [Yip,KevinY, Cheung,DavidW, Ng,MichaelK, Cheung,Kei-Hoi]
通讯作者: Cheung,Kei-Hoi
Bioinformatics and Biostatistics Core
  • 批准号:
    8935163
  • 项目类别:
  • 资助金额:
    $34.27万
  • 财政年份:
    2004
  • 负责人:
    KEI-HOI CHEUNG
  • 依托单位:
Interoperation of Genome Databases and Tools
  • 批准号:
    6367736
  • 项目类别:
  • 资助金额:
    $14.77万
  • 财政年份:
    2001
  • 负责人:
    KEI-HOI CHEUNG
  • 依托单位:
Interoperation of Genome Databases and Tools
  • 批准号:
    6526869
  • 项目类别:
  • 资助金额:
    $14.3万
  • 财政年份:
    2001
  • 负责人:
    KEI-HOI CHEUNG
  • 依托单位:
Interoperation of Genome Databases and Tools
  • 批准号:
    6649804
  • 项目类别:
  • 资助金额:
    $14.2万
  • 财政年份:
    2001
  • 负责人:
    KEI-HOI CHEUNG
  • 依托单位:
海外基金