M2A: 关注正确视频行动识别运动 (M2A: Motion Aware Attention for Accurate Video Action Recognition)

Advancements in attention mechanisms have led to significant performance improvements in a variety of areas in machine learning due to its ability to enable the dynamic modeling of temporal sequences. A particular area in computer vision that is likely to benefit greatly from the incorporation of attention mechanisms in video action recognition. However, much of the current research's focus on attention mechanisms have been on spatial and temporal attention, which are unable to take advantage of the inherent motion found in videos. Motivated by this, we develop a new attention mechanism called Motion Aware Attention (M2A) that explicitly incorporates motion characteristics. More specifically, M2A extracts motion information between consecutive frames and utilizes attention to focus on the motion patterns found across frames to accurately recognize actions in videos. The proposed M2A mechanism is simple to implement and can be easily incorporated into any neural network backbone architecture. We show that incorporating motion mechanisms with attention mechanisms using the proposed M2A mechanism can lead to a +15% to +26% improvement in top-1 accuracy across different backbone architectures, with only a small increase in computational complexity. We further compared the performance of M2A with other state-of-the-art motion and attention mechanisms on the Something-Something V1 video action recognition benchmark. Experimental results showed that M2A can lead to further improvements when combined with other temporal mechanisms and that it outperforms other motion-only or attention-only mechanisms by as much as +60% in top-1 accuracy for specific classes in the benchmark.

翻译：关注机制的提高使机器学习领域各个领域的绩效有了显著的改善,这是因为机器学习能够使时间序列的动态建模成为动态模型。计算机视觉中的一个特定领域可能因将关注机制纳入视频行动识别而大有益处。然而,目前研究对关注机制的关注大多集中在空间和时间上,这些关注机制无法利用视频中发现的内在动作。为此,我们开发了一个新的关注机制,称为 " 感知关注运动 " (M2A),明确纳入运动特点。更具体地说,M2A提取连续框架之间的运动信息,利用对跨框架所发现运动模式的注意,以准确识别视频中的行动。拟议的M2A机制易于实施,并可很容易地纳入任何神经网络主干结构。我们表明,将运动机制与关注机制相结合,利用拟议的M2A机制,可以导致在不同主干结构中将头一级关注率提高15%至26%,而计算复杂性仅略有增加。我们进一步将M2A的绩效与其他P1级和跨框架的移动模式的动作模式加以进一步对比,同时将其他的动态实验性运动和动态前期机制视为其他的升级机制。

相关内容

注意力机制

关注 120

Attention机制最早是在视觉图像领域提出来的，但是真正火起来应该算是google mind团队的这篇论文《Recurrent Models of Visual Attention》[14]，他们在RNN模型上使用了attention机制来进行图像分类。随后，Bahdanau等人在论文《Neural Machine Translation by Jointly Learning to Align and Translate》 [1]中，使用类似attention的机制在机器翻译任务上将翻译和对齐同时进行，他们的工作算是是第一个提出attention机制应用到NLP领域中。接着类似的基于attention机制的RNN模型扩展开始应用到各种NLP任务中。最近，如何在CNN中使用attention机制也成为了大家的研究热点。下图表示了attention研究进展的大概趋势。

计算机科学课程与视频课件合集，Computer Science courses with video lectures

专知会员服务

37+阅读 · 2022年1月24日

【CVPR2021】动态度量学习

专知会员服务

41+阅读 · 2021年3月30日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

【CVPR2020】用于细粒度动作识别的多模式域自适应，Multi-Modal Domain Adaptation for Fine-Grained Action Recognition

专知会员服务

78+阅读 · 2020年2月25日