Augmenting API Documentation with Insights from Stack Overflow

Augmenting API Documentation with Insights from Stack Overflow
复制标题

DOI:
10.1145/2884781.2884800
复制
发表时间:
2016-05
期刊:
2016 IEEE/ACM 38th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Christoph Treude;M. Robillard
Christoph Treude;M. Robillard
中科院分区:
其他
文献类型:
--
作者:
Christoph Treude;M. Robillard

文献摘要

被引文献

相似文献

软件开发人员需要访问不同类型的信息,这些信息通常分散在不同的文档源中,例如API文档或Stack Overflow。我们提出了一种方法来自动增加API文档与“洞察句子”从堆栈溢出-句子是相关的一个特定的API类型,并提供洞察不包含在该类型的API文档。基于1,574个句子的开发集,我们比较了两种最先进的摘要技术以及基于模式的洞察句子提取方法的性能。然后,我们提出了SISE,一种新的基于机器学习的方法,它使用句子本身,它们的格式,它们的问题,它们的答案,它们的作者以及词性标签和句子与相应的API文档的相似性作为特征。使用SISE,我们能够在开发集上实现0.64的精度和0.7的覆盖率。在与八个软件开发人员的比较研究中,我们发现SISE导致被认为添加了API文档中没有的有用信息的句子数量最多。这些结果表明,考虑到Stack Overflow上可用的Meta数据以及词性标签,可以显着改善应用于Stack Overflow数据时的无监督提取方法。
Software developers need access to different kinds of information which is often dispersed among different documentation sources, such as API documentation or Stack Overflow. We present an approach to automatically augment API documentation with "insight sentences" from Stack Overflow -- sentences that are related to a particular API type and that provide insight not contained in the API documentation of that type. Based on a development set of 1,574 sentences, we compare the performance of two state-of-the-art summarization techniques as well as a pattern-based approach for insight sentence extraction. We then present SISE, a novel machine learning based approach that uses as features the sentences themselves, their formatting, their question, their answer, and their authors as well as part-of-speech tags and the similarity of a sentence to the corresponding API documentation. With SISE, we were able to achieve a precision of 0.64 and a coverage of 0.7 on the development set. In a comparative study with eight software developers, we found that SISE resulted in the highest number of sentences that were considered to add useful information not found in the API documentation. These results indicate that taking into account the meta data available on Stack Overflow as well as part-of-speech tags can significantly improve unsupervised extraction approaches when applied to Stack Overflow data.