课题基金 / 基金详情

EAGER: Cataloging Software Using a Semantic-Based Approach for Software Discovery and Characterization

EAGER: Cataloging Software Using a Semantic-Based Approach for Software Discovery and Characterization
EAGER:使用基于语义的软件发现和表征方法对软件进行编目
批准号:
1533792
负责人:
Michael Hucka
金额:
$29.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-01 至 2017-12-31

项目摘要

项目成果

Michael Hucka的其他基金

相似基金

相关文献

中文摘要
翻译
当科学家需要为一项任务找到软件时,他们通常会使用特殊的方法,例如询问同事,搜索网络或使用其他人似乎正在使用的任何方法。这种不系统的方法往往导致非生产性的努力和较差的科学成果,因为科学家无法找到最适合他们需求的软件。如果有一个全面的软件索引,它将有助于更有效地利用现有的软件资源和工具,并有助于更有效和更有效地投资于科学研究本身。此外,软件的创建和发展太快,人类无法监测;因此,索引的创建应该自动化。此外,虽然今天的社会编码运动正在将更多的软件放入开源存储库,从而提供更多的机会来自动查找和描述软件,但仍然存在一个重大障碍:缺乏可以产生人类可接受的结果的自动索引方法。EAGER项目旨在研究使用创新的计算技术来自动创建索引,科学家可以使用该索引来找到他们需要的软件。他们的测试系统(CASICS -综合自动化软件库存创建系统)将与一组初始用户进行测试,并提供给公众使用。在这个项目中,研究人员将(1)扩展现有的软件本体;(2)开发使用本体的源代码分析方法;(3)调整存储库爬虫,将方法应用于SourceForge和GitHub中的项目;(4)开发一个新的软件库。(4)实现对结果数据库的浏览和搜索接口;以及(5)增强搜索设施以经由本体使用语义相似性。他们将使用原型系统来探索代码分析方法的变体,选择最好的,并评估在存储库中发现的软件特征的推断性能。为了使自动化索引创建可行,需要改进软件发现和特征化算法,以便索引完整且组织得更有意义。本项目将探讨一个假设,即更深层次的基础知识结构,加上适当的特征提取和分类算法,可以提高分类性能相比,过去的方法。使用本体来帮助源代码分析已经在其他工作中进行了探索,但它还没有被应用到这里提出的。该项目旨在扩展源代码分析和分类的最新技术水平,其核心问题是基于本体的分类方法产生的软件描述是否可以与人类标记相媲美。为了解决这个问题,除了上述分类方法的知识价值外,研究人员还提出了独特的方法来经验性地评估分类结果,如下所示:(a)评估他们的方法的输出与GitHub和SourceForge中已经存在的分类之间的重叠-他们提议挖掘的两个软件目录,(B)开发一个测试系统,由人类法官对分类软件进行双盲评价;(c)使用一套基准搜索查询,将天文学和系统生物学用户针对通过调查确定的各种情况提出的搜索查询,将CASICS中的搜索与Google中的搜索进行比较。
英文摘要
When scientists need to find software for a task, they often use ad hoc methods such as asking their colleagues, searching the web, or using whatever others appear to be using. This unsystematic approach often results in unproductive effort and poorer scientific results because scientists cannot find the software that is best for their needs. If a comprehensive software index were available it would help enable more efficient use of existing software resources and tools and contribute to more efficient and effective investment in the scientific research itself. Further, software is created and evolves too rapidly for humans to monitor; thus creation of the index should be automated. Also, while today's social-coding movement is putting more software into open-source repositories, thus offering greater opportunities to find and characterize software automatically, a significant obstacle remains: the lack of automated indexing methods that can produce results acceptable to humans. This proposed EAGER project seeks to investigate the use of innovative computing techniques for automating the creation of an index that can be used by scientists to find the software they need. Their test system (CASICS - Comprehensive and Automated Software Inventory Creation System) will be tested with an initial set of users and made available for public use.In this project, the researchers will (1) extend an existing software ontology; (2) develop methods for source code analysis using the ontology; (3) adapt a repository crawler to apply the methods to projects in SourceForge and GitHub; (4) implement a browsing and search interface to the database of results; and (5) augment the search facility to use semantic similarity via the ontology. They will use the prototype system to explore variants of the code analysis methods, select the best, and assess the performance of inferring characteristics of software found in the repositories. For automated index creation to be feasible, software discovery and characterization algorithms need to improve, so that the index is complete and organized more meaningfully. This project will explore the hypothesis that a deeper underlying knowledge structure, coupled with appropriate feature extraction and classification algorithms, can improve classification performance compared to past approaches. The use of ontologies to assist source code analysis has been explored in other work, but it has not been applied as proposed here. The project is expected to extend the state of the art in source code analysis and categorization.The central question in this project is whether the ontology-based descriptions of software produced by the classification methods can compare with human labeling. To address this, and in addition to the intellectual merit of the classification approach described above, the researchers have also proposed unique methods for empirically evaluating the results of the classification, as follows: (a) assess the overlap between the output of their methods with the classifications already present in GitHub and SourceForge - the two software catalogs they propose to mine, (b) develop a test system to perform double-blind evaluation with human judges on the software that is classified and (c) compare the search in CASICS to Google using a set of benchmark search queries that users in astronomy and systems biology would issue for various scenarios, identified through a survey.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Conference: Eighth International Conference on Systems Biology (ICSB 2007) to be held October 1-6, 2007 in Long Beach, CA
  • 批准号:
    0741371
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.0万
  • 财政年份:
    2007
  • 负责人:
    Michael Hucka
  • 依托单位:
海外基金