课题基金 / 基金详情

ITR: Language, Learning, and Modeling Biological Sequences

ITR: Language, Learning, and Modeling Biological Sequences
ITR:语言、学习和生物序列建模
批准号:
0205456
负责人:
Aravind Joshi
金额:
$349.98万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2002
资助国家:
美国
项目状态:
已结题
起止时间:
2002-09-15 至 2008-08-31

项目摘要

项目成果

Aravind Joshi的其他基金

相似基金

相关文献

中文摘要
翻译
EIA-0205456 Joshi。语言、学习和生物序列建模最近在自然语言处理方面的重大进展,如语法和概率机器学习技术的集成,尚未被用于生物序列建模。这些新技术与生物领域高度相关,因为它们支持在几个尺度上整合序列特征,从连续项目之间的依赖关系到涉及复杂结构的依赖关系到整体序列统计。因此,要追求的主要目标是:(1)开发整合语法和概率信息的新技术,特别是整合和评估生物分子二级和三级结构中折叠预测的语法、概率和近似计数方法。(2)开发和评价概率指数模型,用于基因发现,特别是在真核人类‘无核复合体’门的病原体中发现以细胞质膜为靶标的蛋白的基因。这项研究是高度跨学科的,涉及计算机科学、生物学和语言学等学科。它将对生物序列的建模产生重大影响。它还将提供一个极好的机会,培训新的研究人员进行这种跨学科研究,从而为科学和数学教育以及人力资源开发做出贡献。拟议的研究源于2001年2月在宾夕法尼亚大学举行的具有里程碑意义的“生物数据的语言建模”讲习班上进行的许多讨论。
英文摘要
EIA-0205456Joshi. Aravind KUniversity of PennsylvaniaITR: Language, Learning, and Modeling Biological SequencesRecent significant advances in natural language processing such as the integration of grammatical and probabilistic machine-learning techniques have not been exploited for modeling biological sequences. These new techniques are highly relevant to the biological domain because they support the integration of sequence features at several scales, from dependencies between successive items through dependencies involving complex structures to overall sequence statistics. Hence, the major goals to be pursued are: (1) Development of new techniques for integrating grammatical and probabilistic information, in particular, integration and evaluation of grammatical, probabilistic, and approximate counting methods for fold prediction in secondary and tertiary structures of biomolecules. (2) Development and evaluation of probabilistic exponential models for gene finding, in particular genes for apicoplast-targeted proteins in eukaryotic human pathogens of the phylum `Apicomplexa'. This research is highly interdisciplinary, involving the disciplines of computer science, biology and linguistics. It will have a significant impact on the modeling of biological sequences. It will also provide a wonderful opportunity to train new researchers to carry out this interdisciplinary research, thus contributing to science and mathematical education and human resource development. The proposed research arose out of many discussions that took place at a landmark workshop on `Language Modeling of Biological Data' held at the University of Pennsylvania in February 2001.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CI: ADDO-EN: Significant Enhancement of the Exisitng Penn Discourse Treebank
  • 批准号:
    1059353
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2011
  • 负责人:
    Aravind Joshi
  • 依托单位:
RI: Exploiting and Exploring Discourse Connectivity: Deriving New Technology and Knowledge from the Penn Discourse Treebank
  • 批准号:
    0705671
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $95.5万
  • 财政年份:
    2007
  • 负责人:
    Aravind Joshi
  • 依托单位:
Metagrammatical Knowledge for Grammars and Corpora
  • 批准号:
    0414409
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2004
  • 负责人:
    Aravind Joshi
  • 依托单位:
CISE Research Resources: Discourse Penn Treebank and Multimodal FORM: Development of Two Richly Annotated Corpora
  • 批准号:
    0224417
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $99.78万
  • 财政年份:
    2002
  • 负责人:
    Aravind Joshi
  • 依托单位:
海外基金