关于自动软件文件使用图像说明的实证调查 (An Empirical Investigation into the Use of Image Captioning for Automated Software Documentation)

from arxiv, Published in the Proceedings of the 29th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER'22), Honolulu, Hawaii, March 15-18, 2022, pp. 514-525

Existing automated techniques for software documentation typically attempt to reason between two main sources of information: code and natural language. However, this reasoning process is often complicated by the lexical gap between more abstract natural language and more structured programming languages. One potential bridge for this gap is the Graphical User Interface (GUI), as GUIs inherently encode salient information about underlying program functionality into rich, pixel-based data representations. This paper offers one of the first comprehensive empirical investigations into the connection between GUIs and functional, natural language descriptions of software. First, we collect, analyze, and open source a large dataset of functional GUI descriptions consisting of 45,998 descriptions for 10,204 screenshots from popular Android applications. The descriptions were obtained from human labelers and underwent several quality control mechanisms. To gain insight into the representational potential of GUIs, we investigate the ability of four Neural Image Captioning models to predict natural language descriptions of varying granularity when provided a screenshot as input. We evaluate these models quantitatively, using common machine translation metrics, and qualitatively through a large-scale user study. Finally, we offer learned lessons and a discussion of the potential shown by multimodal models to enhance future techniques for automated software documentation.

翻译：软件文件的现有自动化技术通常试图在两种主要信息来源:代码和自然语言之间加以解释。然而,这种推理过程往往由于比较抽象的自然语言和结构化的编程语言之间的词典差距而变得复杂。这种差距的一个潜在桥梁是图形用户界面(GUI),因为图形用户界面内在地将关于基本程序功能的突出信息编码成丰富的像素数据表。本文是对图形用户界面与软件的功能性自然语言描述之间的联系的首次全面经验调查之一。首先,我们收集、分析和公开源收集大量功能性图形界面描述数据集,其中包括10,204个通用用户应用程序的截图的45,998个描述。这些描述是从人类标签上获得的,并经历了若干质量控制机制。为了深入了解图形用户界面的代表性潜力,我们调查了四个神经图像显示模型在提供截图时预测不同颗粒性的自然语言描述的能力。我们用通用机器翻译指标定量评估这些模型,并通过大规模用户研究定性评估这些模型。最后,我们提供了通过多式联运模型展示的自动软件的潜力,以加强未来的技术。

相关内容

Automator

关注 5

Automator是苹果公司为他们的Mac OS X系统开发的一款软件。 只要通过点击拖拽鼠标等操作就可以将一系列动作组合成一个工作流，从而帮助你自动的（可重复的）完成一些复杂的工作。Automator还能横跨很多不同种类的程序，包括：查找器、Safari网络浏览器、iCal、地址簿或者其他的一些程序。它还能和一些第三方的程序一起工作，如微软的Office、Adobe公司的Photoshop或者Pixelmator等。

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日