EAGER: Cataloging Software Using a Semantic-Based Approach for Software Discovery and Characterization
EAGER: Cataloging Software Using a Semantic-Based Approach for Software Discovery and Characterization
批准号:
1533792
负责人:
Michael Hucka
金额:
$29.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-01 至 2017-12-31
中文摘要
当科学家需要为一项任务寻找软件时,他们通常会使用特别的方法,如询问他们的同事、搜索网络或使用其他人似乎正在使用的任何软件。由于科学家找不到最适合他们需要的软件,这种非系统化的方法往往会导致无效的努力和更差的科学结果。如果有一个全面的软件索引,它将有助于更有效地利用现有的软件资源和工具,并有助于对科学研究本身进行更有效率和更有效的投资。此外,软件的创建和发展太快,人类无法监控;因此,索引的创建应该是自动化的。此外,尽管今天的社会编码运动正在将更多的软件放入开源存储库,从而提供了更多的机会来自动发现和描述软件,但一个重大的障碍仍然存在:缺乏能够产生人类可以接受的结果的自动索引方法。这个拟议的项目旨在调查创新计算技术的使用,以自动创建索引,科学家可以使用该索引来找到他们需要的软件。在这个项目中,研究人员将(1)扩展现有的软件本体;(2)开发使用本体的源代码分析方法;(3)采用知识库爬虫将这些方法应用于SourceForge和GitHub中的项目;(4)实现结果数据库的浏览和搜索接口;以及(5)通过本体增强搜索工具以使用语义相似度。他们将使用原型系统来探索代码分析方法的变体,选择最好的,并评估在存储库中发现的软件的推断特征的性能。为了使自动索引创建变得可行,需要改进软件发现和表征算法,以使索引完整且组织得更有意义。这个项目将探索这样一个假设,即更深层次的基础知识结构,加上适当的特征提取和分类算法,可以比过去的方法提高分类性能。在其他工作中已经探索了使用本体来辅助源代码分析,但它还没有像这里建议的那样被应用。该项目有望扩展源代码分析和分类的最新水平。该项目的核心问题是,分类方法产生的基于本体的软件描述是否可以与人工标注相媲美。为了解决这一问题,除了上述分类方法的智力优点外,研究人员还提出了对分类结果进行经验性评估的独特方法,如下:(A)评估其方法的输出与他们建议挖掘的两个软件目录GitHub和SourceForge中已有的分类之间的重叠,(B)开发一个测试系统,与人类评判员对被分类的软件进行双盲评估,以及(C)使用一组基准搜索查询将CASICS中的搜索与Google进行比较,这些查询是天文学和系统生物学用户在通过调查确定的各种情况下会提出的。
英文摘要
When scientists need to find software for a task, they often use ad hoc methods such as asking their colleagues, searching the web, or using whatever others appear to be using. This unsystematic approach often results in unproductive effort and poorer scientific results because scientists cannot find the software that is best for their needs. If a comprehensive software index were available it would help enable more efficient use of existing software resources and tools and contribute to more efficient and effective investment in the scientific research itself. Further, software is created and evolves too rapidly for humans to monitor; thus creation of the index should be automated. Also, while today's social-coding movement is putting more software into open-source repositories, thus offering greater opportunities to find and characterize software automatically, a significant obstacle remains: the lack of automated indexing methods that can produce results acceptable to humans. This proposed EAGER project seeks to investigate the use of innovative computing techniques for automating the creation of an index that can be used by scientists to find the software they need. Their test system (CASICS - Comprehensive and Automated Software Inventory Creation System) will be tested with an initial set of users and made available for public use.In this project, the researchers will (1) extend an existing software ontology; (2) develop methods for source code analysis using the ontology; (3) adapt a repository crawler to apply the methods to projects in SourceForge and GitHub; (4) implement a browsing and search interface to the database of results; and (5) augment the search facility to use semantic similarity via the ontology. They will use the prototype system to explore variants of the code analysis methods, select the best, and assess the performance of inferring characteristics of software found in the repositories. For automated index creation to be feasible, software discovery and characterization algorithms need to improve, so that the index is complete and organized more meaningfully. This project will explore the hypothesis that a deeper underlying knowledge structure, coupled with appropriate feature extraction and classification algorithms, can improve classification performance compared to past approaches. The use of ontologies to assist source code analysis has been explored in other work, but it has not been applied as proposed here. The project is expected to extend the state of the art in source code analysis and categorization.The central question in this project is whether the ontology-based descriptions of software produced by the classification methods can compare with human labeling. To address this, and in addition to the intellectual merit of the classification approach described above, the researchers have also proposed unique methods for empirically evaluating the results of the classification, as follows: (a) assess the overlap between the output of their methods with the classifications already present in GitHub and SourceForge - the two software catalogs they propose to mine, (b) develop a test system to perform double-blind evaluation with human judges on the software that is classified and (c) compare the search in CASICS to Google using a set of benchmark search queries that users in astronomy and systems biology would issue for various scenarios, identified through a survey.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Conference: Eighth International Conference on Systems Biology (ICSB 2007) to be held October 1-6, 2007 in Long Beach, CA
-
批准号:0741371
-
项目类别:Standard Grant
-
资助金额:$1.0万
-
财政年份:2007
-
负责人:Michael Hucka
-
依托单位:
海外基金