课题基金 / 基金详情

项目摘要

项目成果

PAUL Warren STERNBERG的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供): 一个处理生物学论文全文的信息检索和提取系统将是 发展起来的。一个原型系统已经在WormBase运行了一年多,供线虫使用 研究人员以及WormBase生物策展人,最近已在SGD的酵母中实施。这个名为Textpresso的系统将文本分成句子,并根据本体(有组织的词典)对单词和短语进行标记,并允许在已标记句子的数据库上执行查询。目前的本体论包括37类术语,如“基因”、“调节”、“方法”等。本体论可以显著加快特定生物事实的提取,例如基因与基因的相互作用,Textpresso自动执行几乎与专家管理员一样的性能来识别句子;在搜索两个唯一命名的基因和一个交互作用术语时,本体论使搜索效率提高了三倍。该系统将从三个方面进一步开发。首先,核心系统将被改进和改变,以允许扩展到多个感兴趣的领域,例如模式生物、人类疾病。将对系统和网站功能进行简单修改,包括同义词、搜索短语和区分大小写。将支持本地安装的软件包。项目组将维护Textpresso网站(www.extpresso.org)。这将包括线虫和试点系统,但软件包将可用于在本地站点安装Textpresso,例如SGD,Flybase等。第二,本体将被构建得更深一些,并为生物体和领域扩展词汇表 具体条款。第三,将实现信息提取算法。一种方法是使用Textpresso本体的类别(高级节点)来实现相似性度量,以降低关联向量空间的维度。第二种方法将是开发隐马尔可夫模型,以基于标记文本来填充事实模板的空位。提取的信息将提交给用户或专家馆长。 公开描述:研究的质量和速度取决于对已发表信息的快速获取。这个项目将为研究人员提供一个搜索引擎,通过对研究文章的完整文本进行索引,迅速为他们提供他们想要的详细技术信息。
英文摘要
DESCRIPTION (provided by applicant): An information retrieval and extraction system that processes the full text of biological papers will be developed. A prototype system has been in operation at WormBase for over a year, used by C. elegans researchers as well as WormBase biological curators, and has recently been implemented for yeast at SGD. The system, called Textpresso, separates text into sentences, and labels words and phrases according to an ontology (an organized lexicon), and allows queries to be performed on a database of labeled sentences. The current ontology comprises 37 categories of terms, such as "gene," "regulation," "method," etc. Extraction of particular biological facts, such as gene-gene interactions, can be accelerated significantly by ontologies, with Textpresso automatically performing nearly as well as expert curators to identify sentences; in searches for two uniquely named genes and an interaction term, the ontology confers a threefold increase of search efficiency. This system will be further developed in three ways. First, the core system will be refined and altered to allow expansion to multiple domains of interest, e.g., model organisms, human disease. Simple modifications to the system and website functionality will be made, including synonym, search phrases, and case-sensitivity. A software package for local installation will be supported. The project team will maintain the Textpresso site (www.textpresso.org). which will include C. elegans and pilot systems, but software package will be available for installation of Textpresso at local sites, e.g., SGD, Flybase etc. Second, the ontology will be structured somewhat more deeply and lexica expanded for organism and field specific terms. Third, algorithms for information extraction will be implemented. One approach will be the implementation of similarity measures using categories (high level nodes) of the Textpresso ontology to reduce the dimensionality of associated vector spaces. A second approach will be the development of hidden Markov models to fill slots of a fact template based on the marked-up text. Information extracted will be presented to the user or expert curator. Public Description: The quality and pace of research depends upon rapid access to published information. This project will provide researchers with a search engine that rapidly gives them detailed, technical information they want by indexing the complete text of research articles.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Curation at scale: Integrating AI into community curation
  • 批准号:
    10621338
  • 项目类别:
  • 资助金额:
    $35.59万
  • 财政年份:
    2021
  • 负责人:
    PAUL Warren STERNBERG
  • 依托单位:
Curation at scale: Integrating AI into community curation
  • 批准号:
    10344771
  • 项目类别:
  • 资助金额:
    $35.59万
  • 财政年份:
    2021
  • 负责人:
    PAUL Warren STERNBERG
  • 依托单位:
Bipartite gene expression system for C. elegans genetic and neural circuit analysis
  • 批准号:
    9437389
  • 项目类别:
  • 资助金额:
    $24.75万
  • 财政年份:
    2017
  • 负责人:
    PAUL Warren STERNBERG
  • 依托单位:
Genetics 2012: Model Organism to Human Cancer
  • 批准号:
    8319996
  • 项目类别:
  • 资助金额:
    $1.25万
  • 财政年份:
    2012
  • 负责人:
    PAUL Warren STERNBERG
  • 依托单位:
海外基金