课题基金 / 基金详情

SGER: Mining Metadata for Metagenomics

SGER: Mining Metadata for Metagenomics
SGER:挖掘宏基因组学元数据
批准号:
0746650
负责人:
Lynette Hirschman
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2009-02-28

项目摘要

项目成果

Lynette Hirschman的其他基金

相似基金

相关文献

中文摘要
翻译
宏基因组学和生物多样性群落都面临着一个重大挑战:获取提取样本的环境和生态背景的详细数据。如果没有这些“元数据”(关于原始生物数据的数据),原始数据的价值就会大大降低。为了获取这些不同类型的元数据,生物学家/生态学家需要工具,使他们能够调查相关领域的文献,识别文章中的关键概念和词汇,并将数据和元数据提取为适当的表示形式,以便进一步处理、查询和交换。本研究的目标是创建一个用于元数据捕获和管理的交互式工具的概念验证演示,与宏基因组学和生物多样性社区密切合作,了解他们的需求。文本挖掘社区已经取得了重大进展:最近的生物创意研讨会(2007年4月)的结果表明,文本挖掘工具可以识别运行文本中关键生物实体的提及,并将这些提及映射到相关的唯一标识符(例如,Entrez Gene标识符),准确率为80-90%。这种进步很大程度上是由成熟的生物数据库管理员和文本挖掘社区之间的密切联系所推动的。GOA、MINT、完好无损、Flybase、MGI、SGD、Wormbase等组织都有专家策展人,有一个记录在案的策展过程,他们产生了大量的专家策展数据。这些资源为根据人类策划的“黄金标准”测试集评估新工具提供了良好的测试平台。现在的问题是如何应用这一进展来创建工具,以支持管理员在一个交互式过程中提取和绘制关键信息,如基因/蛋白质标识符、地理空间信息或栖息地信息。这些工具将支持新兴社区,如宏基因组学和生物多样性社区;它们也可以放在本体构建者和作者手中,以加速本体的设计,并从源头捕获元数据,而不是依赖于出版后的专家管理。这项工作将在四个不同的领域产生影响:首先,它将支持宏基因组学和生物多样性社区,以加速元数据的捕获;其次,它将为文本挖掘社区提供新的挑战,将工具集成到一个交互式管道中,以支持真正的策展活动;此外,通过探索文献中的概念,这种交互式工具可以对创建新本体和受控词汇的能力产生重大影响;最后,这些工具可以为作者驱动的注释提供原型,以支持在源处捕获和编码元数据,作为注释的一种“自动拼写检查器”。
英文摘要
Both the metagenomics and the biodiversity communities face a major challenge: capturing detailed data on the environmental and the ecological context from which samples are drawn. Without these 'metadata' (data about the primary, biological data), the value of the primary data is greatly diminished. To capture these different types of metadata, biologists/ecologists need tools that allow them to survey the literature in the relevant areas, identify the key concepts and vocabulary in the articles, and extract data and metadata into an appropriate representation for further processing, querying and exchange. The goal of this research is to create a proof-of-concept demonstration of interactive tools for the capture and curation of metadata, working in close collaboration with the metagenomics and biodiversity communities to understand their requirements. The text mining community has demonstrated significant progress: results from the recent BioCreative workshop (April 2007) show that text mining tools can identify mentions of key biological entities in running text and map these mentions to associated unique identifiers (e.g., Entrez Gene identifiers) at 80-90% accuracy. Much of this progress has been driven by a close association between curators of mature biological databases and the text mining community. Groups such as GOA, MINT, IntAct, Flybase, MGI, SGD, Wormbase have expert curators, a documented curation process, and they produce large quantities of expert curated data. These resources have provided good testbeds for evaluating new tools against human curated "gold standard" test sets. The question now is how to apply this progress to create tools to support curators in an interactive process for extraction and mapping of critical information, such as gene/protein identifiers, geospatial information, or habitat information. These tools will support emerging communities, such as the metagenomics and biodiversity communities; they also can be put into the hands of both ontology builders to speed design of ontologies, and authors for capture of metadata at the source, rather than relying on post-publication expert curation. This work will have impact in four distinct areas: first, it will support the metagenomics and biodiversity communities, to speed capture of metadata; second, it will provide new challenges to the text mining community, to integrate tools into an interactive pipeline to support real curation activities; furthermore, such interactive tools can have major impact on the ability to create new ontologies and controlled vocabularies, through exploration of concepts in the literature; and finally, such tools can provide a prototype for author-driven annotation, to support capture and encode metadata at the source, as a kind of "automated spell-checker" for annotations.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: Utility and Usability of Text Mining for Biological Curation
  • 批准号:
    0844419
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2008
  • 负责人:
    Lynette Hirschman
  • 依托单位:
Critical Assessment of Information Extraction in Biology
  • 批准号:
    0640153
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2006
  • 负责人:
    Lynette Hirschman
  • 依托单位:
Evaluating Bioinformatics Technology
  • 批准号:
    0326404
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.1万
  • 财政年份:
    2003
  • 负责人:
    Lynette Hirschman
  • 依托单位:
BioLINK Workshop: Biological Language, Information and Knowledge
  • 批准号:
    0228162
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.43万
  • 财政年份:
    2002
  • 负责人:
    Lynette Hirschman
  • 依托单位:
国内基金
海外基金
基于Genome mining技术研究抑制表皮葡萄球菌生物膜形成的次级代谢产物
  • 批准号:
    21242003
  • 项目类别:
    专项基金项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2012
  • 负责人:
    昌军
  • 依托单位: