FoveaTer:图像分类改造变异器 (FoveaTer: Foveated Transformer for Image Classification)

Many animals and humans process the visual field with a varying spatial resolution (foveated vision) and use peripheral processing to make eye movements and point the fovea to acquire high-resolution information about objects of interest. This architecture results in computationally efficient rapid scene exploration. Recent progress in vision Transformers has brought about new alternatives to the traditionally convolution-reliant computer vision systems. However, these models do not explicitly model the foveated properties of the visual system nor the interaction between eye movements and the classification task. We propose foveated Transformer (FoveaTer) model, which uses pooling regions and saccadic movements to perform object classification tasks using a vision Transformer architecture. Our proposed model pools the image features using squared pooling regions, an approximation to the biologically-inspired foveated architecture, and uses the pooled features as an input to a Transformer Network. It decides on the following fixation location based on the attention assigned by the Transformer to various locations from previous and present fixations. The model uses a confidence threshold to stop scene exploration, allowing to dynamically allocate more fixation/computational resources to more challenging images. We construct an ensemble model using our proposed model and unfoveated model, achieving an accuracy 1.36% below the unfoveated model with 22% computational savings. Finally, we demonstrate our model's robustness against adversarial attacks, where it outperforms the unfoveated model.

翻译：许多动物和人类以不同的空间分辨率(节能视觉)处理视觉场,并使用边缘处理来进行视觉运动,并用边缘处理来显示视觉运动,以显示有关对象的高分辨率信息。这一结构导致计算高效的快速场景勘探。视觉变异器最近的进展为传统上以革命为依存的计算机视觉系统带来了新的替代物。然而,这些模型没有明确模拟视觉系统的顶部特性,或眼运动与分类任务之间的互动。我们提议了变异变异变异变异器(FoveaTer)模型,该模型利用视觉变异器结构来进行视觉运动和分解,以获得高分辨率的物体分类任务。我们提议的模型将图像特征集中到平方集合区域,接近生物激发的织变异结构,并将集合的特性用作对传统变异器网络的投入。它根据变异器对以前和现在的模型不同地点的注意,决定了以下固定位置。模型使用信任阈值来停止现场探索,从而能够以动态的方式将更多的非固定/剖变模型资源分配到更具挑战性的图像中。我们提出的模型,最后用一个模型来显示我们下面的模型。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/